One hundred studies, run again
The claim under scrutiny is that a substantial share of published, peer-reviewed psychology findings do not hold up when the same experiment is run again by an independent team. The clearest test of that claim came from the Open Science Collaboration, coordinated by the psychologist Brian Nosek, which selected 100 studies from three prominent journals, the Journal of Personality and Social Psychology, the Journal of Experimental Psychology: Learning, Memory, and Cognition, and Psychological Science, and attempted to replicate each one using methods as close to the original as possible. Of the 97 original studies that had reported a statistically significant effect, only 36 per cent produced a significant result again, and where an effect did reappear, it was on average roughly half the size originally reported.
Household names that didn’t hold
The replication rate was not uniform across the field. By journal, successful replication ranged from 23 per cent for the social psychology journal to 48 per cent for the cognitive journal, and by subfield, cognitive psychology studies replicated about half the time against roughly a quarter for social psychology. Several individually famous findings mirrored this pattern under direct scrutiny: power posing, the claim that adopting an expansive body position increases confidence and risk tolerance, and ego depletion, the idea that self-control draws down a limited mental resource, both drew heavily publicised failures to replicate. John Bargh’s widely cited 1996 study, in which participants primed with words related to old age reportedly walked more slowly afterward, failed two direct replication attempts in the early 2010s, and Daryl Bem’s experiments claiming evidence for extrasensory perception failed entirely under direct replication after reanalysis found no support for the original result.
Underpowered from the start
What the investigation into causes has established with reasonable confidence is a set of structural, rather than individual, explanations. Psychology studies run on average at a statistical power of only 33 to 36 per cent, well below the 80 per cent threshold generally recommended, with only 8 to 9 per cent of studies reaching that recommended level, meaning most published studies are poorly positioned to detect a real effect reliably even when one exists. A survey by Leslie K. John and colleagues found that 94 per cent of psychologists admitted to using at least one questionable research practice, such as failing to report all measures taken in a study, which 63 per cent admitted to, or selectively reporting only favourable results, which 46 per cent admitted to. Fewer than 20 per cent of journals in the field explicitly welcomed replication studies for publication at all, reinforcing a bias toward novel, positive findings.
A survey that named the practice
What has not held up as neatly is any single, simple fix, or a settled explanation for why some findings replicate and others do not. A large follow-up project, Many Labs 2, involving 186 researchers across 60 laboratories in 36 countries, still found that about 50 per cent of tested findings failed to replicate despite using large combined sample sizes designed to avoid the power problems of the original studies. Notably, the project found that findings which did replicate tended to do so consistently across very different research contexts, and findings that failed tended to fail just as consistently, which argues against the common defence that a result depends heavily on the specific population or setting it was first tested in. Reforms including preregistration of hypotheses and registered reports, where a journal commits to publishing based on methods rather than results, are now widespread, but their long-term effect on the overall replication rate is still being tracked rather than confirmed.
Beyond one field
The consequences reach beyond academic psychology because so much applied advice, in management, education, self-help and public policy, has drawn directly on the original, unreplicated findings. Power posing was popularised in a widely viewed talk before its central claim came under serious replication doubt, and similar patterns have played out with other findings that made the leap from a single study to broad public advice more quickly than the evidence justified. The problem is also not confined to psychology: a 2016 survey found over 70 per cent of researchers across fields had tried and failed to reproduce another scientist’s results, medical research has shown comparably serious reproducibility problems, and even economics, a very different discipline, has documented data-sharing compliance rates well under half. The stakes, in other words, are about how much confidence any reader should place in a single striking study before it has been checked.
Reform, but no settled verdict
Yes, and it is a genuinely useful corrective to read even for people who feel already sceptical of psychology headlines. What makes it worth the time is the specificity: this is not a vague complaint about science being untrustworthy, but a traceable account with an actual replication rate, a named survey of research practices, and individual case studies that show how the pattern plays out for a claim once treated as established fact. It is also, somewhat unusually for a science story, one where the field’s own response, in preregistration, registered reports and large coordinated replication projects, is part of the story rather than an afterthought, which makes it as much a case study in scientific self-correction as it is a cautionary tale.