sciencebriefs
13:00in productionCh. 1 · One hundred studies, run again/ 13:00 · ceiling 15 min
Neuroscience

Replication crisis

In 2015 a coordinated effort to redo 100 published psychology experiments found only 36 per cent reproduced the original significant result, effect sizes roughly halved on average — a number that turned scattered doubts into a named crisis.

The Open Science Collaboration, coordinated by Brian Nosek, attempted to replicate 100 psychology studies from three leading journals and found that only 36 per cent reproduced a statistically significant effect, with replicated effect sizes averaging about half the originals; social psychology fared worse than cognitive psychology, at 25 per cent against 50 per cent. High-profile individual cases, including power posing, ego depletion and John Bargh's social priming studies, followed the same pattern under direct replication. Investigators traced the pattern to low statistical power, a survey finding 94 per cent of psychologists had used at least one questionable research practice, and a publishing system that rewards positive, surprising results, prompting reforms such as preregistration and registered reports whose long-term effect is still being assessed.

Chapters & takeaways6
  1. 0:08
    One hundred studies, run again

    The Open Science Collaboration attempted to redo 100 published psychology experiments and found only 36 per cent reproduced the original significant result.

  2. 2:10
    Household names that didn't hold

    Power posing, ego depletion and a famous priming study on elderly stereotypes all failed under direct replication attempts.

  3. 4:20
    Underpowered from the start

    The average psychology study runs at roughly a third of the statistical power researchers recommend, making chance findings look like real effects.

  4. 6:30
    A survey that named the practice

    Ninety-four per cent of surveyed psychologists admitted to at least one questionable research practice, from selective reporting to peeking at data early.

  5. 8:40
    Beyond one field

    Similar reproducibility problems, with their own numbers, have since been documented in medicine, economics and public health research.

  6. 10:50
    Reform, but no settled verdict

    Preregistration and registered reports are now common, but later large-scale replication projects still found roughly half of tested findings did not hold up.

Worth your time?

Yes. Study the whole thing.

4.5/ 5
What works
  • the 2015 Reproducibility Project gives the whole subject one concrete, well-designed number to anchor around
  • naming specific failed cases, power posing, ego depletion, Bargh's priming study, Daryl Bem's ESP claims, makes an abstract statistic feel consequential
  • the survey finding on questionable research practices identifies a mechanism, not just an outcome
What does not
  • the material doesn't fully explain why findings that did replicate did so consistently across very different contexts, an observation that undercuts one popular excuse for non-replication
  • the crisis is shown extending well beyond psychology, into medicine and economics, without a unified account of why the problem is worse in some fields than others
Study it if
  • anyone who has cited a psychology finding, from power posing to priming, without checking whether it survived a direct replication attempt
  • readers interested in how statistical power and publishing incentives, not fraud, can quietly produce an unreliable literature
  • anyone who wants a numbers-first account of a problem usually discussed in vague terms
Skip it if
  • readers wanting a single villain, since the sources point to structural incentives rather than individual misconduct as the main driver
  • anyone hoping the crisis is now resolved; later large replication projects still found roughly half of findings failing to hold up
The written brief4 min read

One hundred studies, run again

The claim under scrutiny is that a substantial share of published, peer-reviewed psychology findings do not hold up when the same experiment is run again by an independent team. The clearest test of that claim came from the Open Science Collaboration, coordinated by the psychologist Brian Nosek, which selected 100 studies from three prominent journals, the Journal of Personality and Social Psychology, the Journal of Experimental Psychology: Learning, Memory, and Cognition, and Psychological Science, and attempted to replicate each one using methods as close to the original as possible. Of the 97 original studies that had reported a statistically significant effect, only 36 per cent produced a significant result again, and where an effect did reappear, it was on average roughly half the size originally reported.

Household names that didn’t hold

The replication rate was not uniform across the field. By journal, successful replication ranged from 23 per cent for the social psychology journal to 48 per cent for the cognitive journal, and by subfield, cognitive psychology studies replicated about half the time against roughly a quarter for social psychology. Several individually famous findings mirrored this pattern under direct scrutiny: power posing, the claim that adopting an expansive body position increases confidence and risk tolerance, and ego depletion, the idea that self-control draws down a limited mental resource, both drew heavily publicised failures to replicate. John Bargh’s widely cited 1996 study, in which participants primed with words related to old age reportedly walked more slowly afterward, failed two direct replication attempts in the early 2010s, and Daryl Bem’s experiments claiming evidence for extrasensory perception failed entirely under direct replication after reanalysis found no support for the original result.

Underpowered from the start

What the investigation into causes has established with reasonable confidence is a set of structural, rather than individual, explanations. Psychology studies run on average at a statistical power of only 33 to 36 per cent, well below the 80 per cent threshold generally recommended, with only 8 to 9 per cent of studies reaching that recommended level, meaning most published studies are poorly positioned to detect a real effect reliably even when one exists. A survey by Leslie K. John and colleagues found that 94 per cent of psychologists admitted to using at least one questionable research practice, such as failing to report all measures taken in a study, which 63 per cent admitted to, or selectively reporting only favourable results, which 46 per cent admitted to. Fewer than 20 per cent of journals in the field explicitly welcomed replication studies for publication at all, reinforcing a bias toward novel, positive findings.

A survey that named the practice

What has not held up as neatly is any single, simple fix, or a settled explanation for why some findings replicate and others do not. A large follow-up project, Many Labs 2, involving 186 researchers across 60 laboratories in 36 countries, still found that about 50 per cent of tested findings failed to replicate despite using large combined sample sizes designed to avoid the power problems of the original studies. Notably, the project found that findings which did replicate tended to do so consistently across very different research contexts, and findings that failed tended to fail just as consistently, which argues against the common defence that a result depends heavily on the specific population or setting it was first tested in. Reforms including preregistration of hypotheses and registered reports, where a journal commits to publishing based on methods rather than results, are now widespread, but their long-term effect on the overall replication rate is still being tracked rather than confirmed.

Beyond one field

The consequences reach beyond academic psychology because so much applied advice, in management, education, self-help and public policy, has drawn directly on the original, unreplicated findings. Power posing was popularised in a widely viewed talk before its central claim came under serious replication doubt, and similar patterns have played out with other findings that made the leap from a single study to broad public advice more quickly than the evidence justified. The problem is also not confined to psychology: a 2016 survey found over 70 per cent of researchers across fields had tried and failed to reproduce another scientist’s results, medical research has shown comparably serious reproducibility problems, and even economics, a very different discipline, has documented data-sharing compliance rates well under half. The stakes, in other words, are about how much confidence any reader should place in a single striking study before it has been checked.

Reform, but no settled verdict

Yes, and it is a genuinely useful corrective to read even for people who feel already sceptical of psychology headlines. What makes it worth the time is the specificity: this is not a vague complaint about science being untrustworthy, but a traceable account with an actual replication rate, a named survey of research practices, and individual case studies that show how the pattern plays out for a claim once treated as established fact. It is also, somewhat unusually for a science story, one where the field’s own response, in preregistration, registered reports and large coordinated replication projects, is part of the story rather than an afterthought, which makes it as much a case study in scientific self-correction as it is a cautionary tale.

Same field · Neuroscience4 of 45
Up next in Science

Ribosome

· 13:00

The ribosome is a machine made mostly of RNA that reads genetic messages and builds proteins one amino acid at a time, and this brief traces how that was worked out and what still counts as settled.

13:00