A question first asked of a chimpanzee
The claim under test is that understanding another person’s mind, specifically recognising that someone else can hold a belief that is false, is not automatic but develops at a particular point and can fail to develop typically in some people. The idea traces to a 1978 paper by David Premack and Guy Woodruff asking whether a chimpanzee has a theory of mind, which established the research agenda later applied to human children. Wimmer and Perner formalised the human version of the test in 1983, designing a scenario in which a character leaves a scene, an object’s location is changed in their absence, and the child being tested is asked to predict where the returning character will look for it. A child who says the character will look in the object’s new location, based on the child’s own knowledge rather than the character’s, fails to demonstrate this understanding.
A marble that moves while Sally isn’t looking
The best-known implementation of this idea is the Sally-Anne test, developed by Simon Baron-Cohen, Alan Leslie and Uta Frith in 1985 building directly on Wimmer and Perner’s 1983 design, with the underlying concept originally suggested by the philosopher Daniel Dennett. Using two dolls, Sally and Anne, the procedure has Sally place a marble in her basket and then leave the scene; while she is gone, Anne moves the marble into her own box. When Sally returns, the child is asked where she will look for it. Passing the test means saying Sally will look in the basket, where she left it, rather than the box, where the child knows it actually is. Leslie and Frith later replicated the procedure in 1988 using human actors instead of dolls, with comparable results.
Two dolls, one test, a clear number
The 1985 study, involving 61 children in total, produced results specific enough to compare directly across groups: 23 of 27 typically developing children, or 85 per cent, answered correctly, against 12 of 14 children with Down syndrome, or 86 per cent, and only 4 of 20 autistic children, 20 per cent. That gap between the autistic group and the other two, who performed almost identically to each other, became the empirical basis for treating theory-of-mind understanding as an area where autism specifically diverges from other forms of developmental difference. Separately, the developmental timeline has held up reasonably well: most typically developing children begin passing the standard false-belief task around age four, and children younger than that consistently tend to answer based on their own knowledge rather than the absent character’s belief.
Passing around age four
What has not held up as cleanly is the interpretation that autistic children categorically lack a theory of mind. A 2019 review by Gernsbacher and Yergeau argued that this claim is empirically questionable, pointing to failed replications of the original effect and generally small effect sizes across meta-analyses of later studies, and noting that some autistic individuals do pass false-belief tasks despite the theory predicting they should not. A 2001 study by Ruffman, Garnham and Rideout found that autistic and learning-disabled children answered the Sally-Anne test about equally well, but that their eye-gaze patterns while doing so differed, raising the possibility that any difference lies more in social attention than in the underlying cognitive capacity being tested. Separately, the discovery that simple reflexive robots, with nothing resembling genuine cognition, can be built to pass versions of the false-belief task has raised a more basic question about whether the paradigm measures what it claims to.
A deficit theory that hasn’t aged cleanly
Beyond child psychology, the stakes of this research touch directly on how autism is understood, diagnosed and discussed. Treating a poor score on a false-belief task as evidence of a specific cognitive deficit shaped decades of autism research and public description of the condition, and the pushback against that framing, including the idea of a double empathy problem in which autistic and non-autistic people find each other equally hard to read, has real consequences for how autistic people are characterised and accommodated. The developmental timing side of the research matters too, for anyone assessing when a young child can reasonably be expected to reason about what someone else knows or believes, relevant to everything from legal questions about children as witnesses to how educators design activities requiring children to take another person’s perspective.
What a passing score actually proves
Yes, particularly for anyone who has encountered the Sally-Anne test only as a factoid rather than as a specific, numerically documented study that has since been argued over. The original 1985 result is a clean piece of developmental psychology, with real comparative numbers across three groups of children, and it is genuinely informative to see how much weight was placed on that single study for decades afterward. What makes it worth the fuller read, rather than the one-line summary, is watching later researchers take the same test apart methodically, questioning what a correct or incorrect answer actually reveals, and finding that the honest picture is considerably less tidy than the original 20-per-cent-versus-85-per-cent contrast suggested on its own.