sciencebriefs
13:00in productionCh. 1 · What a designed experiment is for/ 13:00 · ceiling 15 min
Applied science

Design of experiments

A well-designed experiment and an underpowered one can ask the identical question and still deserve opposite trust. The gap between them is arithmetic, and across whole fields the arithmetic has not been going well.

Design of experiments supplies the logic — randomisation, replication, blocking — that lets a study claim a change caused an effect rather than merely accompanied it. Statistical power supplies the arithmetic that says whether a given study, given its size and the effect it is chasing, had a real chance of finding anything at all. Field-by-field surveys cited alongside the concept report median power well under the conventional 80% target in several disciplines.

Chapters & takeaways6
  1. 0:08
    What a designed experiment is for

    Randomisation, replication and blocking exist to let a result be attributed to a treatment rather than to chance or confounding.

  2. 2:10
    From Peirce's weights to Fisher's fields

    The formal logic runs from Peirce's 1880s randomised trials through Fisher's 1935 book to the textbooks still used today.

  3. 4:20
    What power actually measures

    Power is the probability a study will detect an effect that is really there, and it rises with sample size, effect size and a looser significance threshold.

  4. 6:30
    Why 80% is a convention, not a law

    The commonly used power target reflects an accepted trade-off between two error rates, not a statistical requirement.

  5. 8:40
    The arithmetic behind the replication crisis

    Surveyed fields report median power far below the 80% convention, which the analysis links directly to a glut of false positives.

  6. 10:50
    Worth learning even without the maths

    The concept explains, in plain terms, why some published results were unlikely to survive being checked.

Worth your time?

Yes. Study the whole thing.

4/ 5
What works
  • connects a specific historical lineage, from Peirce through Fisher, to present-day practice
  • explains power as a direct function of sample size, effect size and significance level
  • cites field-by-field power surveys rather than asserting the problem in the abstract
What does not
  • does not resolve the dispute over whether post-hoc power analysis is meaningful
  • leaves exactly how funders and journals should act on low power only loosely addressed
Study it if
  • anyone who reads about a single small study and wonders whether to believe it
  • students or early researchers designing their first experiment
  • readers trying to understand why the replication crisis happened
Skip it if
  • readers wanting a working formula rather than the concepts behind it
  • anyone looking for a single villain rather than a systemic arithmetic problem
The written brief4 min read

What a designed experiment is for

Two separate claims sit behind the terms design of experiments and statistical power, and they answer different questions. Design of experiments is the set of procedures — deciding what to manipulate, what to hold constant, and how to assign subjects to conditions — that lets a researcher attribute an observed change to a treatment rather than to some other source of variation. Its core tools are randomisation, which spreads unknown confounding factors evenly across groups; replication, which repeats measurements to separate true effects from noise; and blocking, which groups similar experimental units together before treatments are assigned. Statistical power then asks a narrower, later question of any given study: given its sample size and the size of effect it is looking for, what is the probability it would detect that effect if the effect were real?

From Peirce’s weights to Fisher’s fields

The lineage of experimental design runs back further than is often assumed. Charles Sanders Peirce set out randomisation-based statistical inference in the late 1870s and, around 1885, ran one of the earliest randomised, blinded, repeated-measures experiments, having volunteers discriminate between small weight differences. Ronald Fisher then formalised the field for a wider audience through his agricultural work, publishing The Arrangement of Field Experiments in 1926 and The Design of Experiments in 1935, establishing randomisation, replication and blocking as the standard toolkit still taught today. Later contributors — Bose and Kishen, Plackett and Burman, and the widely used 1950 textbook by Cox and Cochran — refined the efficiency of these designs without displacing Fisher’s basic logic.

What power actually measures

Power itself is defined precisely: it equals one minus the probability of a Type II error, the chance of failing to detect an effect that genuinely exists. It rises with a larger sample, with a bigger true effect, and with a looser significance threshold, and each of those relationships is direct rather than approximate — tighten the significance level to guard against false positives, for instance, and power falls unless the sample grows to compensate. This is why power calculations, done before a study begins, ask for an estimate of the effect size a researcher expects and a chosen significance level, and return the minimum sample needed to have a stated chance of detecting that effect if it is there.

Why 80% is a convention, not a law

The convention of targeting 80% power, corresponding to a Type II error rate of 0.2 against a significance level of 0.05, is presented plainly as a convention rather than a rule — the article is explicit that there are no formal standards for power, and that the 80% figure simply encodes an accepted four-to-one trade-off between missing a real effect and wrongly claiming one. A rule of thumb for a simple two-sample comparison gives the sample size needed for that convention as roughly sixteen times the population variance divided by the square of the difference being tested, which illustrates how quickly required sample sizes grow as the effect being chased gets smaller.

The arithmetic behind the replication crisis

Where this becomes more than a classroom exercise is in the field-by-field power estimates the discussion assembles: median power reported at around 18% in economics, 10% in political science, 36% in psychology, and 15% in ecology and evolutionary biology, all well below the conventional 80% target. The consequence, spelled out directly, is that when many underpowered studies are run across a field, the published findings that clear a significance threshold are more likely to be false positives than genuine effects, because a low-power study that does report a significant result is disproportionately likely to have done so by chance. This is offered as a substantial part of the mechanism behind the broader replication crisis, alongside practices such as adjusting analyses until a result crosses the 0.05 threshold.

Worth learning even without the maths

This is worth understanding even for a reader who has no intention of running a study, because it supplies a concrete reason to be sceptical of a small, striking finding beyond a vague sense that ‘more data would be nice’. Knowing that a discipline’s median power sits near 15% or 20% changes how much weight a single significant result in that discipline deserves, and knowing why — the arithmetic linking sample size, effect size and detection probability — is more useful than simply being told to distrust small studies. The material is honest about its own limits too, noting real disagreement over whether power calculated after the fact tells you anything at all, which is a more careful position than most popular accounts of the replication crisis take.

Same field · Applied science4 of 28
Up next in Science

Hyperthermophile

· 13:00

An organism identified in Yellowstone hot springs in 1965 pushed back what biologists thought life could survive, and a heat-tolerant enzyme from that same discovery went on to make modern DNA testing possible, which this brief follows from bacterium to lab bench.

13:00