What a designed experiment is for
Two separate claims sit behind the terms design of experiments and statistical power, and they answer different questions. Design of experiments is the set of procedures — deciding what to manipulate, what to hold constant, and how to assign subjects to conditions — that lets a researcher attribute an observed change to a treatment rather than to some other source of variation. Its core tools are randomisation, which spreads unknown confounding factors evenly across groups; replication, which repeats measurements to separate true effects from noise; and blocking, which groups similar experimental units together before treatments are assigned. Statistical power then asks a narrower, later question of any given study: given its sample size and the size of effect it is looking for, what is the probability it would detect that effect if the effect were real?
From Peirce’s weights to Fisher’s fields
The lineage of experimental design runs back further than is often assumed. Charles Sanders Peirce set out randomisation-based statistical inference in the late 1870s and, around 1885, ran one of the earliest randomised, blinded, repeated-measures experiments, having volunteers discriminate between small weight differences. Ronald Fisher then formalised the field for a wider audience through his agricultural work, publishing The Arrangement of Field Experiments in 1926 and The Design of Experiments in 1935, establishing randomisation, replication and blocking as the standard toolkit still taught today. Later contributors — Bose and Kishen, Plackett and Burman, and the widely used 1950 textbook by Cox and Cochran — refined the efficiency of these designs without displacing Fisher’s basic logic.
What power actually measures
Power itself is defined precisely: it equals one minus the probability of a Type II error, the chance of failing to detect an effect that genuinely exists. It rises with a larger sample, with a bigger true effect, and with a looser significance threshold, and each of those relationships is direct rather than approximate — tighten the significance level to guard against false positives, for instance, and power falls unless the sample grows to compensate. This is why power calculations, done before a study begins, ask for an estimate of the effect size a researcher expects and a chosen significance level, and return the minimum sample needed to have a stated chance of detecting that effect if it is there.
Why 80% is a convention, not a law
The convention of targeting 80% power, corresponding to a Type II error rate of 0.2 against a significance level of 0.05, is presented plainly as a convention rather than a rule — the article is explicit that there are no formal standards for power, and that the 80% figure simply encodes an accepted four-to-one trade-off between missing a real effect and wrongly claiming one. A rule of thumb for a simple two-sample comparison gives the sample size needed for that convention as roughly sixteen times the population variance divided by the square of the difference being tested, which illustrates how quickly required sample sizes grow as the effect being chased gets smaller.
The arithmetic behind the replication crisis
Where this becomes more than a classroom exercise is in the field-by-field power estimates the discussion assembles: median power reported at around 18% in economics, 10% in political science, 36% in psychology, and 15% in ecology and evolutionary biology, all well below the conventional 80% target. The consequence, spelled out directly, is that when many underpowered studies are run across a field, the published findings that clear a significance threshold are more likely to be false positives than genuine effects, because a low-power study that does report a significant result is disproportionately likely to have done so by chance. This is offered as a substantial part of the mechanism behind the broader replication crisis, alongside practices such as adjusting analyses until a result crosses the 0.05 threshold.
Worth learning even without the maths
This is worth understanding even for a reader who has no intention of running a study, because it supplies a concrete reason to be sceptical of a small, striking finding beyond a vague sense that ‘more data would be nice’. Knowing that a discipline’s median power sits near 15% or 20% changes how much weight a single significant result in that discipline deserves, and knowing why — the arithmetic linking sample size, effect size and detection probability — is more useful than simply being told to distrust small studies. The material is honest about its own limits too, noting real disagreement over whether power calculated after the fact tells you anything at all, which is a more careful position than most popular accounts of the replication crisis take.