There seems to be two factors at play here that are creating almost nonsensical results.
1. The results, when filtered through the replicability guidelines rendered a clear verdict - 61 to 39 against.
2. The question of "how closely did the findings resemble the original study?" flips the findings. Moderately similar findings are the majority - 58 to 42 in the other direction.
How can you have a study that has "virtually identical" findings that doesn't replicate the original?
Original results that were exactly on the borderline of statistical significance, with new results that were barely on the "fail" side of that line, but both sets of results are within the margin of error of each other?
(I am not a frequentist, so my bias may be showing through)
Comments
There seems to be two factors at play here that are creating almost nonsensical results.
1. The results, when filtered through the replicability guidelines rendered a clear verdict - 61 to 39 against.
2. The question of "how closely did the findings resemble the original study?" flips the findings. Moderately similar findings are the majority - 58 to 42 in the other direction.
How can you have a study that has "virtually identical" findings that doesn't replicate the original?
Original results that were exactly on the borderline of statistical significance, with new results that were barely on the "fail" side of that line, but both sets of results are within the margin of error of each other?
(I am not a frequentist, so my bias may be showing through)