Replication Research
25 Research Replication Examples Across Science, Medicine, and Social Science
See 25 research replication examples across laboratory science, clinical medicine, and social science, including successful replications, mixed results, and failures. Use the comparison criteria to judge what each result actually shows.

What these research replication examples show
Replication is not a courtroom verdict. It is a stress test: if researchers repeat a claim with new data, new analysts, new instruments, or a tighter design, does the original result still look credible?
The 25 examples below show three broad outcomes. Some findings held up through independent confirmation. Some collapsed after better measurement or randomization. Many landed in the middle: the effect existed, but was smaller, narrower, more context-dependent, or harder to measure than the original paper implied.
A few terms matter before the examples:
Direct replication repeats the original method as closely as possible, usually with a new sample.
Conceptual replication tests the same underlying idea with a different method, measure, population, or setting.
Independent confirmation asks whether separate teams, instruments, or datasets converge on the same conclusion.
Reproducibility of analysis means another researcher can obtain the same result from the original data, code, and documented analysis.
A failed replication does not automatically mean fraud. It can reflect low statistical power, different populations, changed protocols, measurement error, poor documentation, hidden exclusions, implementation differences, or a real effect that appears only under narrower conditions.

# | Example | Discipline | Replication type | Result | Main lesson |
|---|---|---|---|---|---|
1 | OPERA faster-than-light neutrinos | Physics | Independent checks | Failed | Extraordinary results need instrument-level scrutiny |
2 | BICEP2 primordial gravitational waves | Cosmology | Independent data and reanalysis | Qualified/failed as first claimed | Alternative explanations can mimic signals |
3 | Arsenic-based life | Microbiology | Direct and methodological critique | Failed/unsupported | Materials and contamination controls matter |
4 | LK-99 superconductivity | Materials science | Independent lab attempts | Failed | Viral claims are not reproducible evidence |
5 | Ambient-pressure room-temperature superconductivity | Materials physics | Independent non-replication and data scrutiny | Failed/retracted | Replication and data integrity checks reinforce each other |
6 | Higgs boson | Particle physics | Independent confirmation | Successful | Converging detectors are stronger than one analysis |
7 | LIGO gravitational waves | Physics | Multi-site detection and later events | Successful | Instrumental replication can occur across observatories |
8 | Reproducibility Project: Cancer Biology | Preclinical biology | Multi-study replication | Mixed | Missing protocols make replication hard to interpret |
9 | Hormone therapy and heart protection | Medicine | Stronger-design replication | Reversed/qualified | Randomization can overturn observational expectations |
10 | Beta-carotene and cancer prevention | Medicine | Randomized trials | Failed; safety concern | Nutrient associations do not prove supplement benefit |
11 | WHI dietary intervention | Medicine | Large multicenter trial | Mixed/null for main outcomes | Diet effects need long, well-measured tests |
12 | PREDIMED Mediterranean diet | Clinical nutrition | Reanalysis after design irregularities | Broadly supported but qualified | Reproducing analysis cannot fully fix randomization flaws |
13 | ORBITA PCI for stable angina | Cardiology | Sham-controlled trial | Weaker symptom effect | Placebo-controlled procedures can change interpretation |
14 | Dexamethasone for severe COVID-19 | Clinical medicine | Platform trial and later evidence | Successful in severe disease | Treatment effects can depend sharply on disease stage |
15 | SPRINT blood pressure target | Cardiology | Related trials and analyses | Qualified support | Measurement protocol and population affect generalization |
16 | Antidepressant trial evidence | Psychiatry | Meta-analysis with unpublished trials | Qualified/weaker | Publication bias can inflate apparent effects |
17 | Stroke neuroprotection models | Translational medicine | Clinical translation attempts | Mostly failed | Animal-model replication is not clinical generalizability |
18 | Psychology reproducibility project | Psychology | Large-scale direct replication | Mixed/low replication rate | Criteria determine how “replicated” is counted |
19 | Many Labs 1 | Psychology | Multi-lab direct replication | Mixed | Effects vary by site, task, and robustness |
20 | Ego depletion RRR | Psychology | Preregistered multi-lab replication | Failed/weaker | Preregistration reduces analytic flexibility |
21 | Power posing | Psychology | Direct and conceptual replications | Weaker/disputed | Bold behavioral claims need separated outcomes |
22 | Social priming “walking elderly” | Psychology | Direct replications | Failed/disputed | Subtle cues are sensitive to procedure and bias |
23 | Facial feedback | Psychology | Preregistered multi-lab replication | Mixed | Small effects are hard to estimate consistently |
24 | Marshmallow test | Developmental psychology | Conceptual replications | Qualified | Background context changes interpretation |
25 | Ultimatum game across cultures | Behavioral economics | Cross-cultural replication | Successful with variation | Replication can reveal boundary conditions |
8 replication examples from laboratory and physical science
1. Faster-than-light neutrinos and the OPERA follow-up
Original research question: Could neutrinos sent from CERN to the OPERA detector arrive faster than light?
Replication design: The claim prompted independent timing checks, follow-up measurements, and scrutiny of the experimental hardware rather than a simple rerun of the statistical analysis.
Outcome: Confidence in the faster-than-light result collapsed after instrumentation problems were identified. The episode became a model case for why extraordinary claims need independent measurement and engineering review.
Main limitation: This was not a clean direct replication in the ordinary sense. It was an error-tracing process around a complex instrument chain.
2. BICEP2’s claimed primordial gravitational-wave signal
Original research question: Did BICEP2 detect polarization in the cosmic microwave background consistent with primordial gravitational waves from the early universe?
Replication design: Other teams compared the signal against independent observations of galactic dust polarization and alternative foreground models.
Outcome: The original interpretation weakened. The signal could not be cleanly attributed to primordial gravitational waves once dust was accounted for.
Main limitation: The replication problem was partly observational and partly interpretive. The instrument saw a pattern, but the dispute concerned what produced it.
3. The arsenic-based life claim
Original research question: Could a bacterium use arsenic in place of phosphorus in biomolecules such as DNA?
Replication design: Independent researchers attempted to grow and analyze the organism under stricter chemical controls, while others criticized the original methods.
Outcome: The stronger claim did not hold. Later work did not support arsenic replacing phosphorus in the central way initially proposed.
Main limitation: The case turned on difficult analytical chemistry, contamination control, and definitions of “use” versus “tolerate.”
4. LK-99 room-temperature superconductivity
Original research question: Was LK-99 a room-temperature, ambient-pressure superconductor?
Replication design: Laboratories around the world attempted to synthesize the material, test magnetic behavior, measure resistance, and compare samples.
Outcome: The early claim did not survive independent testing. Observed effects were better explained by ordinary material properties or impurities than by superconductivity.
Main limitation: Rapid public attention outpaced slow materials characterization. Some failed attempts may not have reproduced the exact synthesis, but the overall evidence did not support the headline claim.
5. Retracted ambient-pressure room-temperature superconductivity claim
Original research question: Had researchers produced superconductivity near room temperature under ambient pressure?
Replication design: Independent groups tried to reproduce the reported behavior, while the published data and analysis faced closer scrutiny.
Outcome: Confidence fell sharply after non-replication and concerns about the data record. The claim was later withdrawn from the formal literature.
Main limitation: Non-replication alone does not identify the exact source of failure. In this case, interpretation also depended on editorial review, data inspection, and materials reproducibility.
6. Higgs boson discovery across ATLAS and CMS
Original research question: Did particle collisions at the Large Hadron Collider reveal the Higgs boson predicted by the Standard Model?
Replication design: Two large detector collaborations, ATLAS and CMS, analyzed independent data streams using different detector systems and analysis pipelines.
Outcome: The convergence strengthened confidence. This was not one lab repeating another lab’s bench protocol; it was independent confirmation through separate instruments built to test the same physical prediction.
Main limitation: Both collaborations operated within the same accelerator environment, so this was not independence in every possible sense.
7. LIGO gravitational-wave detections across observatories
Original research question: Could gravitational waves from distant astrophysical events be directly detected on Earth?
Replication design: LIGO used geographically separated observatories to detect the same passing signal with timing and waveform consistency. Later detections added event-level replication.
Outcome: Confidence grew because signals appeared across instruments, sites, and subsequent events. The replication unit was not a single experiment but a networked detection system.
Main limitation: Replication depends on sophisticated noise modeling. A gravitational-wave event cannot be repeated on demand like a lab assay.
8. Reproducibility Project: Cancer Biology
Original research question: Could influential preclinical cancer biology findings be reproduced by independent teams?
Replication design: Researchers selected published cancer biology experiments and attempted to reproduce key claims using available methods, materials, and protocols.
Outcome: Results were mixed and often hard to classify. Some findings were consistent, some were weaker, and some could not be interpreted cleanly because the original experimental details were incomplete.
Main limitation: Preclinical biology depends on cell lines, reagents, animal models, endpoint definitions, and tacit lab know-how. Poor documentation can make a replication failure ambiguous rather than decisive.

9 replication examples from medicine and clinical research
9. Hormone replacement therapy and cardiovascular protection
Original research question: Did postmenopausal hormone therapy protect against cardiovascular disease?
Replication design: Earlier observational findings were tested against randomized clinical trial evidence, most notably large women’s health trials.
Outcome: The protective interpretation weakened or reversed for many patients. Randomization exposed confounding that observational studies could not fully remove.
Main limitation: The answer depends on timing, formulation, age, baseline risk, and outcome. The replication did not make hormone therapy “bad” in every context; it narrowed the claim.
10. Beta-carotene supplementation for cancer prevention
Original research question: Could beta-carotene supplementation reduce cancer risk, especially where diets rich in fruits and vegetables appeared protective?
Replication design: Large randomized trials tested beta-carotene pills directly rather than inferring benefit from dietary patterns.
Outcome: The expected prevention benefit did not appear, and some trial results raised safety concerns in high-risk groups.
Main limitation: A supplement trial does not replicate the full dietary pattern associated with lower risk. It tests an isolated intervention.
11. The Women’s Health Initiative dietary intervention
Original research question: Would a low-fat dietary pattern reduce risks such as breast cancer, colorectal cancer, or cardiovascular disease?
Replication design: A large multicenter randomized dietary intervention tested the hypothesis more directly than prior observational associations.
Outcome: The main results were more modest than many expectations. Some outcomes did not show the predicted large protective effects.
Main limitation: Diet trials face adherence problems, long latency periods, baseline diet variation, and difficulty isolating one dietary component.
12. PREDIMED’s Mediterranean-diet trial
Original research question: Does a Mediterranean diet supplemented with olive oil or nuts reduce major cardiovascular events?
Replication design: The key replication issue was not a separate new trial but a correction and reanalysis after irregularities in randomization were identified.
Outcome: The reanalysis broadly supported the diet’s benefit but qualified confidence in the original trial design.
Main limitation: Reproducing an analysis can check whether conclusions survive adjusted handling of the data. It cannot turn a flawed randomization process into a perfectly randomized one.
13. ORBITA trial of percutaneous coronary intervention for stable angina
Original research question: How much symptom relief does percutaneous coronary intervention provide for stable angina beyond placebo and medical therapy?
Replication design: ORBITA used a sham-controlled design, comparing PCI with a procedure-like control condition.
Outcome: The symptomatic advantage was smaller than many clinicians and patients expected from unblinded experience.
Main limitation: The result applies most directly to the studied stable-angina population and follow-up period. It should not be generalized to acute coronary syndromes or different patient groups.
14. Dexamethasone for severe COVID-19
Original research question: Could dexamethasone reduce mortality in hospitalized COVID-19 patients?
Replication design: A large randomized platform trial tested dexamethasone across disease-severity groups, and later evidence assessed whether the effect held in other settings.
Outcome: The benefit was strongest in patients requiring oxygen or ventilatory support. The same treatment was not interpreted as broadly beneficial for all COVID-19 patients at all disease stages.
Main limitation: Changing viral variants, background care, vaccination status, and patient mix can alter absolute benefit.
15. SPRINT blood-pressure trial
Original research question: Does a lower systolic blood-pressure target reduce cardiovascular events in high-risk adults without diabetes?
Replication design: Subsequent analyses and related trials examined whether intensive targets generalized across subgroups, measurement protocols, and clinical environments.
Outcome: The intensive-treatment finding influenced practice, but interpretation depends on how blood pressure was measured and which patients match the trial population.
Main limitation: A target measured under one protocol may not translate directly to routine office readings.
16. Antidepressant efficacy and clinical-trial evidence
Original research question: Are antidepressants more effective than placebo for depressive symptoms?
Replication design: Meta-analytic reanalyses compared published trials with unpublished or negative trials and examined sponsor reporting patterns.
Outcome: The effect remained real for some drugs and patients, but selective publication and reporting made the apparent evidence base look stronger than the full trial record.
Main limitation: Antidepressant response varies by severity, drug class, trial duration, placebo response, and outcome scale. This is a cumulative-evidence replication problem, not one definitive rerun.
17. Preclinical stroke and neuroprotection studies
Original research question: Could neuroprotective treatments that worked in animal stroke models improve outcomes in human stroke patients?
Replication design: Many interventions moved from animal experiments into human clinical trials, effectively testing whether the biological effect translated across models and species.
Outcome: Many promising preclinical effects did not produce successful human treatments.
Main limitation: This is not just failed replication. It can reflect poor animal-model validity, timing differences, dosing, comorbidities, publication bias, and the complexity of human stroke care. For writers designing medical projects, this is a useful reminder that a good research design example must specify population, intervention, comparator, and outcome before the evidence can be judged.

[[OTIO_INLINE_PROMO:%7B%22title%22%3A%22Which%20clinical%20findings%20survived%20stronger%20designs%3F%22%2C%22description%22%3A%22Add%20the%20original%20and%20replication%20papers%20to%20Otio%2C%20then%20compare%20their%20methods%2C%20outcomes%2C%20and%20limitations%20with%20cited%20answers.%22%7D]]
8 replication examples from social science
18. The Open Science Collaboration’s psychology replication project
Original research question: Could a broad sample of published psychology findings be reproduced by independent teams?
Replication design: Many research groups attempted direct replications of published effects using planned protocols and new samples.
Outcome: The project found that many effects were weaker or did not meet common replication criteria. The exact replication rate depends on the criterion used: statistical significance, effect-size comparison, subjective assessment, or combined measures.
Main limitation: Large replication projects sample studies, not whole fields perfectly. Some original effects may be context-sensitive, while some replication attempts may miss important implementation details.
19. Many Labs 1
Original research question: Do well-known psychological effects appear when the same tasks are run across many laboratories and participant pools?
Replication design: Multiple labs administered shared protocols to test a set of established effects across sites.
Outcome: Some effects replicated clearly; others were weaker or inconsistent across settings.
Main limitation: Multi-lab replication improves generalizability checks, but it can still simplify local context, language, incentives, and participant interpretation.
20. Registered Replication Report on ego depletion
Original research question: Does exerting self-control on one task reduce performance on a later self-control task?
Replication design: A preregistered multi-lab project used a standardized protocol to test the ego-depletion effect with less analytic flexibility.
Outcome: The replicated effect was null or much weaker than the influential original literature suggested.
Main limitation: Critics argued about task choice and whether the manipulation captured the core theory. That makes the case a dispute over both effect size and construct validity.
21. Power posing
Original research question: Do expansive body postures change hormones, risk tolerance, or feelings of power?
Replication design: Later studies separated subjective feelings, behavioral outcomes, and biological measures, often using larger samples and more explicit protocols.
Outcome: Evidence for hormonal and behavioral effects weakened sharply. Some evidence for subjective feelings remained more plausible than the full original package.
Main limitation: “Power posing” bundled several outcomes. A replication can fail for one outcome while leaving a narrower claim open.
22. Social priming and the “walking elderly” effect
Original research question: Can subtle exposure to age-related words make people walk more slowly?
Replication design: Later researchers attempted to repeat the priming procedure under more controlled conditions and with closer attention to experimenter expectations.
Outcome: Replications were failed or disputed, lowering confidence in the original strong claim.
Main limitation: Social priming effects can depend on minute procedural details, but that sensitivity also makes them difficult to use as robust evidence.
23. Facial-feedback research
Original research question: Can manipulating facial muscles, such as by holding a pen in the mouth, change emotional experience or humor ratings?
Replication design: Larger preregistered replication efforts attempted standardized versions of the classic facial-feedback manipulation.
Outcome: Results were mixed. Some replications found little or no effect under the tested conditions, while later debate focused on whether features such as camera presence or task framing changed the manipulation.
Main limitation: A small embodied-cognition effect can be real but hard to estimate. Replication design must distinguish “no effect” from “wrong context for the effect.”
24. The marshmallow test and delayed gratification
Original research question: Does a child’s ability to delay gratification predict later life outcomes?
Replication design: Later conceptual replications used larger and more diverse samples, adding family background, socioeconomic conditions, and trust-related variables.
Outcome: The simple interpretation weakened. Delayed gratification still mattered, but it was entangled with background conditions and the child’s environment.
Main limitation: This is a conceptual replication, not a pure repeat. It tests whether the original interpretation generalizes once broader social variables are included.
25. Cross-cultural ultimatum-game experiments
Original research question: Do people across societies behave like narrow self-interest models predict when splitting money in an ultimatum game?
Replication design: Researchers ran similar bargaining tasks across different communities and cultural settings.
Outcome: The broad task replicated, but behavior varied substantially. People often rejected unfair offers, yet what counted as fair and acceptable differed across groups.
Main limitation: Cross-cultural replication is not meant to erase variation. Its value is showing which patterns recur and which depend on norms, markets, kinship, and local institutions. If the goal is to frame a student project around a testable social-science question, start with clear qualitative, quantitative, or mixed-methods research questions before choosing the replication design.

How to evaluate a replication before citing it
Treat a replication as evidence with a method, not as a label. The question is not “Did it replicate?” The better question is: what exactly was tested again, what changed, and what conclusion is now safer to make?
Start with the replication type.
Direct replication: Did the researchers repeat the original protocol closely?
Conceptual replication: Did they test the same theory through a different operational definition?
Analytical reproducibility: Did they re-run the original analysis from the same data and code?
Independent confirmation: Did separate teams, instruments, datasets, or settings converge on the claim?
Then compare the design features that usually decide whether the result is interpretable:
sample size and statistical power
preregistration and analysis plan
blinding and randomization
measurement instruments
inclusion and exclusion criteria
primary versus secondary outcomes
manipulation checks
implementation fidelity
data, code, materials, and protocol availability
correction, retraction, or critique history
Do not collapse every non-confirming result into “failed replication.” A replication may produce:
a true null result
an opposite result
a much smaller effect
an inconclusive estimate
a failed manipulation check
an effect only in a subgroup
a result that depends on an unplanned analysis
a result that reproduces the data analysis but not the original design claim
That distinction matters when citing evidence. A failed manipulation check says the replication may not have tested the same mechanism. A smaller effect says the claim may be directionally right but overstated. A subgroup-only finding says the general claim needs narrower wording.
For a literature review, record the evidence trail in one place: original paper, replication paper, dataset or code repository, protocol, correction notices, and major critiques. An AI research workspace such as Otio’s AI PDF reader can help keep PDFs, links, highlights, and notes together while comparing original and replication claims. The important part is not the software; it is that the reasoning remains auditable.
A practical citation rule: cite the original and the replication together when the replication materially qualifies the original claim. Write the sentence so the reader knows what changed.
Weak version: “Study X found that power posing changes behavior.”
Better version: “Study X reported behavioral and hormonal effects from expansive posture, but later replication work found weaker or disputed evidence, especially for the biological outcomes.”
That habit prevents the most common abuse of replication evidence: treating a famous original as settled fact or treating a single failed replication as the final word.
If you are building your own evidence table, use columns that force judgment rather than summary:
Field | What to record |
|---|---|
Original claim | The exact causal, correlational, or descriptive claim |
Original method | Sample, measure, intervention, and outcome |
Replication type | Direct, conceptual, analytical, or independent confirmation |
Key changes | Population, protocol, instruments, analysis, setting |
Result | Confirmed, weaker, null, opposite, mixed, or inconclusive |
Main limitation | The reason the result should not be overread |
Citation language | The sentence you would safely write in a paper |
For more help separating design choices before you interpret findings, see Otio’s guide to research methodology types and examples of quantitative research designs.
FAQ
Q: What is the difference between replication and reproducibility?
A: Replication tests whether a finding can be obtained again with new data or a new study. Reproducibility usually means other researchers can obtain the same result from the original data, code, and documented analysis.
Q: Does a failed replication mean the original study was wrong?
A: Not necessarily. The result may depend on population, protocol, measurement, statistical power, or context, so a failed replication usually lowers confidence or narrows the claim rather than resolving every possible explanation.
Q: Why are social-science findings often difficult to replicate?
A: Many social-science effects are sensitive to sampling, culture, timing, incentives, researcher expectations, and small procedural differences. Effects can also be small relative to measurement noise and publication bias.
Q: How should replication evidence be cited in a research paper?
A: Cite the original study and the relevant replication or reanalysis, then state whether the later work confirmed, weakened, qualified, or failed to reproduce the original result. Avoid describing a single replication as definitive evidence when the broader literature is mixed.
[[OTIO_FOOTER_PROMO:%7B%22title%22%3A%22Build%20your%20own%20replication%20evidence%20trail%22%2C%22description%22%3A%22Keep%20your%20original%20papers%2C%20replications%2C%20corrections%2C%20and%20critiques%20together%20in%20Otio%20while%20recording%20design%20changes%20and%20safe%20citation%20language.%22%7D]]




