Replication Research

25 Research Replication Examples Across Science, Medicine, and Social Science

See 25 research replication examples across laboratory science, clinical medicine, and social science, including successful replications, mixed results, and failures. Use the comparison criteria to judge what each result actually shows.

People in the office

What these research replication examples show

Replication is not a courtroom verdict. It is a stress test: if researchers repeat a claim with new data, new analysts, new instruments, or a tighter design, does the original result still look credible?

The 25 examples below show three broad outcomes. Some findings held up through independent confirmation. Some collapsed after better measurement or randomization. Many landed in the middle: the effect existed, but was smaller, narrower, more context-dependent, or harder to measure than the original paper implied.

A few terms matter before the examples:

  • Direct replication repeats the original method as closely as possible, usually with a new sample.

  • Conceptual replication tests the same underlying idea with a different method, measure, population, or setting.

  • Independent confirmation asks whether separate teams, instruments, or datasets converge on the same conclusion.

  • Reproducibility of analysis means another researcher can obtain the same result from the original data, code, and documented analysis.

A failed replication does not automatically mean fraud. It can reflect low statistical power, different populations, changed protocols, measurement error, poor documentation, hidden exclusions, implementation differences, or a real effect that appears only under narrower conditions.

Diagram comparing direct replication, conceptual replication, and reproducible analysis

#

Example

Discipline

Replication type

Result

Main lesson

1

OPERA faster-than-light neutrinos

Physics

Independent checks

Failed

Extraordinary results need instrument-level scrutiny

2

BICEP2 primordial gravitational waves

Cosmology

Independent data and reanalysis

Qualified/failed as first claimed

Alternative explanations can mimic signals

3

Arsenic-based life

Microbiology

Direct and methodological critique

Failed/unsupported

Materials and contamination controls matter

4

LK-99 superconductivity

Materials science

Independent lab attempts

Failed

Viral claims are not reproducible evidence

5

Ambient-pressure room-temperature superconductivity

Materials physics

Independent non-replication and data scrutiny

Failed/retracted

Replication and data integrity checks reinforce each other

6

Higgs boson

Particle physics

Independent confirmation

Successful

Converging detectors are stronger than one analysis

7

LIGO gravitational waves

Physics

Multi-site detection and later events

Successful

Instrumental replication can occur across observatories

8

Reproducibility Project: Cancer Biology

Preclinical biology

Multi-study replication

Mixed

Missing protocols make replication hard to interpret

9

Hormone therapy and heart protection

Medicine

Stronger-design replication

Reversed/qualified

Randomization can overturn observational expectations

10

Beta-carotene and cancer prevention

Medicine

Randomized trials

Failed; safety concern

Nutrient associations do not prove supplement benefit

11

WHI dietary intervention

Medicine

Large multicenter trial

Mixed/null for main outcomes

Diet effects need long, well-measured tests

12

PREDIMED Mediterranean diet

Clinical nutrition

Reanalysis after design irregularities

Broadly supported but qualified

Reproducing analysis cannot fully fix randomization flaws

13

ORBITA PCI for stable angina

Cardiology

Sham-controlled trial

Weaker symptom effect

Placebo-controlled procedures can change interpretation

14

Dexamethasone for severe COVID-19

Clinical medicine

Platform trial and later evidence

Successful in severe disease

Treatment effects can depend sharply on disease stage

15

SPRINT blood pressure target

Cardiology

Related trials and analyses

Qualified support

Measurement protocol and population affect generalization

16

Antidepressant trial evidence

Psychiatry

Meta-analysis with unpublished trials

Qualified/weaker

Publication bias can inflate apparent effects

17

Stroke neuroprotection models

Translational medicine

Clinical translation attempts

Mostly failed

Animal-model replication is not clinical generalizability

18

Psychology reproducibility project

Psychology

Large-scale direct replication

Mixed/low replication rate

Criteria determine how “replicated” is counted

19

Many Labs 1

Psychology

Multi-lab direct replication

Mixed

Effects vary by site, task, and robustness

20

Ego depletion RRR

Psychology

Preregistered multi-lab replication

Failed/weaker

Preregistration reduces analytic flexibility

21

Power posing

Psychology

Direct and conceptual replications

Weaker/disputed

Bold behavioral claims need separated outcomes

22

Social priming “walking elderly”

Psychology

Direct replications

Failed/disputed

Subtle cues are sensitive to procedure and bias

23

Facial feedback

Psychology

Preregistered multi-lab replication

Mixed

Small effects are hard to estimate consistently

24

Marshmallow test

Developmental psychology

Conceptual replications

Qualified

Background context changes interpretation

25

Ultimatum game across cultures

Behavioral economics

Cross-cultural replication

Successful with variation

Replication can reveal boundary conditions

8 replication examples from laboratory and physical science

1. Faster-than-light neutrinos and the OPERA follow-up

Original research question: Could neutrinos sent from CERN to the OPERA detector arrive faster than light?

Replication design: The claim prompted independent timing checks, follow-up measurements, and scrutiny of the experimental hardware rather than a simple rerun of the statistical analysis.

Outcome: Confidence in the faster-than-light result collapsed after instrumentation problems were identified. The episode became a model case for why extraordinary claims need independent measurement and engineering review.

Main limitation: This was not a clean direct replication in the ordinary sense. It was an error-tracing process around a complex instrument chain.

2. BICEP2’s claimed primordial gravitational-wave signal

Original research question: Did BICEP2 detect polarization in the cosmic microwave background consistent with primordial gravitational waves from the early universe?

Replication design: Other teams compared the signal against independent observations of galactic dust polarization and alternative foreground models.

Outcome: The original interpretation weakened. The signal could not be cleanly attributed to primordial gravitational waves once dust was accounted for.

Main limitation: The replication problem was partly observational and partly interpretive. The instrument saw a pattern, but the dispute concerned what produced it.

3. The arsenic-based life claim

Original research question: Could a bacterium use arsenic in place of phosphorus in biomolecules such as DNA?

Replication design: Independent researchers attempted to grow and analyze the organism under stricter chemical controls, while others criticized the original methods.

Outcome: The stronger claim did not hold. Later work did not support arsenic replacing phosphorus in the central way initially proposed.

Main limitation: The case turned on difficult analytical chemistry, contamination control, and definitions of “use” versus “tolerate.”

4. LK-99 room-temperature superconductivity

Original research question: Was LK-99 a room-temperature, ambient-pressure superconductor?

Replication design: Laboratories around the world attempted to synthesize the material, test magnetic behavior, measure resistance, and compare samples.

Outcome: The early claim did not survive independent testing. Observed effects were better explained by ordinary material properties or impurities than by superconductivity.

Main limitation: Rapid public attention outpaced slow materials characterization. Some failed attempts may not have reproduced the exact synthesis, but the overall evidence did not support the headline claim.

5. Retracted ambient-pressure room-temperature superconductivity claim

Original research question: Had researchers produced superconductivity near room temperature under ambient pressure?

Replication design: Independent groups tried to reproduce the reported behavior, while the published data and analysis faced closer scrutiny.

Outcome: Confidence fell sharply after non-replication and concerns about the data record. The claim was later withdrawn from the formal literature.

Main limitation: Non-replication alone does not identify the exact source of failure. In this case, interpretation also depended on editorial review, data inspection, and materials reproducibility.

6. Higgs boson discovery across ATLAS and CMS

Original research question: Did particle collisions at the Large Hadron Collider reveal the Higgs boson predicted by the Standard Model?

Replication design: Two large detector collaborations, ATLAS and CMS, analyzed independent data streams using different detector systems and analysis pipelines.

Outcome: The convergence strengthened confidence. This was not one lab repeating another lab’s bench protocol; it was independent confirmation through separate instruments built to test the same physical prediction.

Main limitation: Both collaborations operated within the same accelerator environment, so this was not independence in every possible sense.

7. LIGO gravitational-wave detections across observatories

Original research question: Could gravitational waves from distant astrophysical events be directly detected on Earth?

Replication design: LIGO used geographically separated observatories to detect the same passing signal with timing and waveform consistency. Later detections added event-level replication.

Outcome: Confidence grew because signals appeared across instruments, sites, and subsequent events. The replication unit was not a single experiment but a networked detection system.

Main limitation: Replication depends on sophisticated noise modeling. A gravitational-wave event cannot be repeated on demand like a lab assay.

8. Reproducibility Project: Cancer Biology

Original research question: Could influential preclinical cancer biology findings be reproduced by independent teams?

Replication design: Researchers selected published cancer biology experiments and attempted to reproduce key claims using available methods, materials, and protocols.

Outcome: Results were mixed and often hard to classify. Some findings were consistent, some were weaker, and some could not be interpreted cleanly because the original experimental details were incomplete.

Main limitation: Preclinical biology depends on cell lines, reagents, animal models, endpoint definitions, and tacit lab know-how. Poor documentation can make a replication failure ambiguous rather than decisive.

Timeline of original scientific claims and independent replication attempts

9 replication examples from medicine and clinical research

9. Hormone replacement therapy and cardiovascular protection

Original research question: Did postmenopausal hormone therapy protect against cardiovascular disease?

Replication design: Earlier observational findings were tested against randomized clinical trial evidence, most notably large women’s health trials.

Outcome: The protective interpretation weakened or reversed for many patients. Randomization exposed confounding that observational studies could not fully remove.

Main limitation: The answer depends on timing, formulation, age, baseline risk, and outcome. The replication did not make hormone therapy “bad” in every context; it narrowed the claim.

10. Beta-carotene supplementation for cancer prevention

Original research question: Could beta-carotene supplementation reduce cancer risk, especially where diets rich in fruits and vegetables appeared protective?

Replication design: Large randomized trials tested beta-carotene pills directly rather than inferring benefit from dietary patterns.

Outcome: The expected prevention benefit did not appear, and some trial results raised safety concerns in high-risk groups.

Main limitation: A supplement trial does not replicate the full dietary pattern associated with lower risk. It tests an isolated intervention.

11. The Women’s Health Initiative dietary intervention

Original research question: Would a low-fat dietary pattern reduce risks such as breast cancer, colorectal cancer, or cardiovascular disease?

Replication design: A large multicenter randomized dietary intervention tested the hypothesis more directly than prior observational associations.

Outcome: The main results were more modest than many expectations. Some outcomes did not show the predicted large protective effects.

Main limitation: Diet trials face adherence problems, long latency periods, baseline diet variation, and difficulty isolating one dietary component.

12. PREDIMED’s Mediterranean-diet trial

Original research question: Does a Mediterranean diet supplemented with olive oil or nuts reduce major cardiovascular events?

Replication design: The key replication issue was not a separate new trial but a correction and reanalysis after irregularities in randomization were identified.

Outcome: The reanalysis broadly supported the diet’s benefit but qualified confidence in the original trial design.

Main limitation: Reproducing an analysis can check whether conclusions survive adjusted handling of the data. It cannot turn a flawed randomization process into a perfectly randomized one.

13. ORBITA trial of percutaneous coronary intervention for stable angina

Original research question: How much symptom relief does percutaneous coronary intervention provide for stable angina beyond placebo and medical therapy?

Replication design: ORBITA used a sham-controlled design, comparing PCI with a procedure-like control condition.

Outcome: The symptomatic advantage was smaller than many clinicians and patients expected from unblinded experience.

Main limitation: The result applies most directly to the studied stable-angina population and follow-up period. It should not be generalized to acute coronary syndromes or different patient groups.

14. Dexamethasone for severe COVID-19

Original research question: Could dexamethasone reduce mortality in hospitalized COVID-19 patients?

Replication design: A large randomized platform trial tested dexamethasone across disease-severity groups, and later evidence assessed whether the effect held in other settings.

Outcome: The benefit was strongest in patients requiring oxygen or ventilatory support. The same treatment was not interpreted as broadly beneficial for all COVID-19 patients at all disease stages.

Main limitation: Changing viral variants, background care, vaccination status, and patient mix can alter absolute benefit.

15. SPRINT blood-pressure trial

Original research question: Does a lower systolic blood-pressure target reduce cardiovascular events in high-risk adults without diabetes?

Replication design: Subsequent analyses and related trials examined whether intensive targets generalized across subgroups, measurement protocols, and clinical environments.

Outcome: The intensive-treatment finding influenced practice, but interpretation depends on how blood pressure was measured and which patients match the trial population.

Main limitation: A target measured under one protocol may not translate directly to routine office readings.

16. Antidepressant efficacy and clinical-trial evidence

Original research question: Are antidepressants more effective than placebo for depressive symptoms?

Replication design: Meta-analytic reanalyses compared published trials with unpublished or negative trials and examined sponsor reporting patterns.

Outcome: The effect remained real for some drugs and patients, but selective publication and reporting made the apparent evidence base look stronger than the full trial record.

Main limitation: Antidepressant response varies by severity, drug class, trial duration, placebo response, and outcome scale. This is a cumulative-evidence replication problem, not one definitive rerun.

17. Preclinical stroke and neuroprotection studies

Original research question: Could neuroprotective treatments that worked in animal stroke models improve outcomes in human stroke patients?

Replication design: Many interventions moved from animal experiments into human clinical trials, effectively testing whether the biological effect translated across models and species.

Outcome: Many promising preclinical effects did not produce successful human treatments.

Main limitation: This is not just failed replication. It can reflect poor animal-model validity, timing differences, dosing, comorbidities, publication bias, and the complexity of human stroke care. For writers designing medical projects, this is a useful reminder that a good research design example must specify population, intervention, comparator, and outcome before the evidence can be judged.

Clinical evidence ladder from observation to randomized replication

[[OTIO_INLINE_PROMO:%7B%22title%22%3A%22Which%20clinical%20findings%20survived%20stronger%20designs%3F%22%2C%22description%22%3A%22Add%20the%20original%20and%20replication%20papers%20to%20Otio%2C%20then%20compare%20their%20methods%2C%20outcomes%2C%20and%20limitations%20with%20cited%20answers.%22%7D]]

8 replication examples from social science

18. The Open Science Collaboration’s psychology replication project

Original research question: Could a broad sample of published psychology findings be reproduced by independent teams?

Replication design: Many research groups attempted direct replications of published effects using planned protocols and new samples.

Outcome: The project found that many effects were weaker or did not meet common replication criteria. The exact replication rate depends on the criterion used: statistical significance, effect-size comparison, subjective assessment, or combined measures.

Main limitation: Large replication projects sample studies, not whole fields perfectly. Some original effects may be context-sensitive, while some replication attempts may miss important implementation details.

19. Many Labs 1

Original research question: Do well-known psychological effects appear when the same tasks are run across many laboratories and participant pools?

Replication design: Multiple labs administered shared protocols to test a set of established effects across sites.

Outcome: Some effects replicated clearly; others were weaker or inconsistent across settings.

Main limitation: Multi-lab replication improves generalizability checks, but it can still simplify local context, language, incentives, and participant interpretation.

20. Registered Replication Report on ego depletion

Original research question: Does exerting self-control on one task reduce performance on a later self-control task?

Replication design: A preregistered multi-lab project used a standardized protocol to test the ego-depletion effect with less analytic flexibility.

Outcome: The replicated effect was null or much weaker than the influential original literature suggested.

Main limitation: Critics argued about task choice and whether the manipulation captured the core theory. That makes the case a dispute over both effect size and construct validity.

21. Power posing

Original research question: Do expansive body postures change hormones, risk tolerance, or feelings of power?

Replication design: Later studies separated subjective feelings, behavioral outcomes, and biological measures, often using larger samples and more explicit protocols.

Outcome: Evidence for hormonal and behavioral effects weakened sharply. Some evidence for subjective feelings remained more plausible than the full original package.

Main limitation: “Power posing” bundled several outcomes. A replication can fail for one outcome while leaving a narrower claim open.

22. Social priming and the “walking elderly” effect

Original research question: Can subtle exposure to age-related words make people walk more slowly?

Replication design: Later researchers attempted to repeat the priming procedure under more controlled conditions and with closer attention to experimenter expectations.

Outcome: Replications were failed or disputed, lowering confidence in the original strong claim.

Main limitation: Social priming effects can depend on minute procedural details, but that sensitivity also makes them difficult to use as robust evidence.

23. Facial-feedback research

Original research question: Can manipulating facial muscles, such as by holding a pen in the mouth, change emotional experience or humor ratings?

Replication design: Larger preregistered replication efforts attempted standardized versions of the classic facial-feedback manipulation.

Outcome: Results were mixed. Some replications found little or no effect under the tested conditions, while later debate focused on whether features such as camera presence or task framing changed the manipulation.

Main limitation: A small embodied-cognition effect can be real but hard to estimate. Replication design must distinguish “no effect” from “wrong context for the effect.”

24. The marshmallow test and delayed gratification

Original research question: Does a child’s ability to delay gratification predict later life outcomes?

Replication design: Later conceptual replications used larger and more diverse samples, adding family background, socioeconomic conditions, and trust-related variables.

Outcome: The simple interpretation weakened. Delayed gratification still mattered, but it was entangled with background conditions and the child’s environment.

Main limitation: This is a conceptual replication, not a pure repeat. It tests whether the original interpretation generalizes once broader social variables are included.

25. Cross-cultural ultimatum-game experiments

Original research question: Do people across societies behave like narrow self-interest models predict when splitting money in an ultimatum game?

Replication design: Researchers ran similar bargaining tasks across different communities and cultural settings.

Outcome: The broad task replicated, but behavior varied substantially. People often rejected unfair offers, yet what counted as fair and acceptable differed across groups.

Main limitation: Cross-cultural replication is not meant to erase variation. Its value is showing which patterns recur and which depend on norms, markets, kinship, and local institutions. If the goal is to frame a student project around a testable social-science question, start with clear qualitative, quantitative, or mixed-methods research questions before choosing the replication design.

Multi-laboratory replication map for a social science experiment

How to evaluate a replication before citing it

Treat a replication as evidence with a method, not as a label. The question is not “Did it replicate?” The better question is: what exactly was tested again, what changed, and what conclusion is now safer to make?

Start with the replication type.

  • Direct replication: Did the researchers repeat the original protocol closely?

  • Conceptual replication: Did they test the same theory through a different operational definition?

  • Analytical reproducibility: Did they re-run the original analysis from the same data and code?

  • Independent confirmation: Did separate teams, instruments, datasets, or settings converge on the claim?

Then compare the design features that usually decide whether the result is interpretable:

  • sample size and statistical power

  • preregistration and analysis plan

  • blinding and randomization

  • measurement instruments

  • inclusion and exclusion criteria

  • primary versus secondary outcomes

  • manipulation checks

  • implementation fidelity

  • data, code, materials, and protocol availability

  • correction, retraction, or critique history

Do not collapse every non-confirming result into “failed replication.” A replication may produce:

  • a true null result

  • an opposite result

  • a much smaller effect

  • an inconclusive estimate

  • a failed manipulation check

  • an effect only in a subgroup

  • a result that depends on an unplanned analysis

  • a result that reproduces the data analysis but not the original design claim

That distinction matters when citing evidence. A failed manipulation check says the replication may not have tested the same mechanism. A smaller effect says the claim may be directionally right but overstated. A subgroup-only finding says the general claim needs narrower wording.

For a literature review, record the evidence trail in one place: original paper, replication paper, dataset or code repository, protocol, correction notices, and major critiques. An AI research workspace such as Otio’s AI PDF reader can help keep PDFs, links, highlights, and notes together while comparing original and replication claims. The important part is not the software; it is that the reasoning remains auditable.

A practical citation rule: cite the original and the replication together when the replication materially qualifies the original claim. Write the sentence so the reader knows what changed.

Weak version: “Study X found that power posing changes behavior.”

Better version: “Study X reported behavioral and hormonal effects from expansive posture, but later replication work found weaker or disputed evidence, especially for the biological outcomes.”

That habit prevents the most common abuse of replication evidence: treating a famous original as settled fact or treating a single failed replication as the final word.

If you are building your own evidence table, use columns that force judgment rather than summary:

Field

What to record

Original claim

The exact causal, correlational, or descriptive claim

Original method

Sample, measure, intervention, and outcome

Replication type

Direct, conceptual, analytical, or independent confirmation

Key changes

Population, protocol, instruments, analysis, setting

Result

Confirmed, weaker, null, opposite, mixed, or inconclusive

Main limitation

The reason the result should not be overread

Citation language

The sentence you would safely write in a paper

For more help separating design choices before you interpret findings, see Otio’s guide to research methodology types and examples of quantitative research designs.

FAQ

Q: What is the difference between replication and reproducibility?
A: Replication tests whether a finding can be obtained again with new data or a new study. Reproducibility usually means other researchers can obtain the same result from the original data, code, and documented analysis.

Q: Does a failed replication mean the original study was wrong?
A: Not necessarily. The result may depend on population, protocol, measurement, statistical power, or context, so a failed replication usually lowers confidence or narrows the claim rather than resolving every possible explanation.

Q: Why are social-science findings often difficult to replicate?
A: Many social-science effects are sensitive to sampling, culture, timing, incentives, researcher expectations, and small procedural differences. Effects can also be small relative to measurement noise and publication bias.

Q: How should replication evidence be cited in a research paper?
A: Cite the original study and the relevant replication or reanalysis, then state whether the later work confirmed, weakened, qualified, or failed to reproduce the original result. Avoid describing a single replication as definitive evidence when the broader literature is mixed.

[[OTIO_FOOTER_PROMO:%7B%22title%22%3A%22Build%20your%20own%20replication%20evidence%20trail%22%2C%22description%22%3A%22Keep%20your%20original%20papers%2C%20replications%2C%20corrections%2C%20and%20critiques%20together%20in%20Otio%20while%20recording%20design%20changes%20and%20safe%20citation%20language.%22%7D]]