sciencebriefs
13:00in productionCh. 1 · From one reaction to millions at once/ 13:00 · ceiling 15 min
Genetics

Massively parallel sequencing

Reading DNA one lane at a time cost the Human Genome Project years and roughly a hundred million dollars. Spreading millions of short reads across one flow cell instead brought that down toward a few hundred dollars, trading length and certainty for volume.

Next-generation, or massively parallel, sequencing replaced Sanger sequencing's one-reaction-at-a-time approach with millions of short DNA fragments read simultaneously off a single flow cell. Platforms from 454, Illumina, Ion Torrent and others drove sequencing costs down from roughly a hundred million dollars for the Human Genome Project's draft to a few hundred dollars per genome within two decades, while newer long-read platforms from Pacific Biosciences and Oxford Nanopore trade some accuracy for reads tens of thousands of bases long.

Chapters & takeaways6
  1. 0:08
    From one reaction to millions at once

    Massively parallel sequencing reads huge numbers of short DNA fragments simultaneously off a single flow cell instead of one reaction at a time.

  2. 2:10
    A genome project measured in years and millions

    The Human Genome Project's draft, finished in 2003, took a large budget and years of Sanger-based work to reach roughly 92% accuracy.

  3. 4:20
    A market that consolidated fast

    454 launched the first commercial platform in 2003, but Illumina's technology came to generate the large majority of the world's sequencing data within a decade.

  4. 6:30
    What shorter, faster reads give up

    Second-generation platforms trade read length and, in some cases, homopolymer accuracy for speed and volume, which complicates repetitive or structurally complex DNA.

  5. 8:40
    The long-read answer to that trade-off

    Pacific Biosciences and Oxford Nanopore read tens of thousands of bases at a stretch, accepting a lower per-base accuracy in exchange.

  6. 10:50
    Worth knowing before trusting a headline genome figure

    Any claim about sequencing speed or cost is meaningless without knowing which platform and which trade-offs produced it.

Worth your time?

Yes. Study the whole thing.

4/ 5
What works
  • traces the cost curve from the Human Genome Project's roughly $100 million draft to platforms nearing $200 a genome
  • names the specific technical trade-off, read length against throughput, that separates second- and third-generation platforms
  • is specific about which platform achieved which milestone rather than crediting 'NGS' in the abstract
What does not
  • does not settle which current platform is best for a given use, since the market keeps shifting
  • leaves the underlying chemistry of each sequencing method only briefly described
Study it if
  • anyone who has heard the '$1,000 genome' phrase and wants to know what it actually means
  • readers curious how genomic medicine became affordable enough for routine use
  • people who want the underlying trade-offs behind competing sequencing platforms
Skip it if
  • readers wanting a single simple sequencing method rather than a landscape of competing platforms
  • anyone after the chemistry detail rather than the historical and cost trajectory
The written brief3 min read

From one reaction to millions at once

The claim is that DNA sequencing can be made dramatically cheaper and faster by changing the unit of work from a single reaction to millions of them running side by side. Where Sanger sequencing reads one set of DNA fragments through electrophoresis at a time, massively parallel sequencing attaches spatially separated, clonally amplified DNA templates, or in some cases single molecules, across a flow cell and reads all of them simultaneously, generating anywhere from roughly a million to tens of billions of short reads, each between about fifty and four hundred bases, in a single instrument run. That shift from sequential to parallel reading is the entire basis for the throughput and cost gains that followed.

A genome project measured in years and millions

The scale of the earlier, one-reaction-at-a-time era is the necessary baseline. The Human Genome Project’s draft sequence, built on Sanger-based methods, was completed in 2003 at around 92% accuracy, following large-scale sequencing trials that the NIH ran around 1990 at an estimated cost of about seventy-five cents per base. Sequencing costs during that period were reported falling from roughly one hundred million dollars in 2001 to about ten thousand dollars by 2011 as the technology improved, and it was not until 2022 that the final, most difficult eight percent of the genome was completed, producing the fully finished reference sequence of 3.1 billion base pairs.

A market that consolidated fast

Massively parallel platforms arrived and reordered the market quickly. Technology of this kind first emerged in the 1990s and became commercially available from around 2005, with 454 Life Sciences launching the GS20, the first commercial next-generation sequencer, in 2003, producing reads of several hundred bases at high accuracy but modest total output per run. Illumina’s Genome Analyzer followed in 2006 with shorter individual reads but far higher total throughput, and the platform’s later instruments, including the HiSeq X Ten in 2014 and NovaSeq in 2017, drove the field toward the long-sought ‘$1,000 genome’ and beyond, with Illumina reported to generate more than ninety percent of global sequencing data by the early 2020s as competing platforms such as 454 and SOLiD were discontinued.

What shorter, faster reads give up

This gain in throughput did not come free. Second-generation platforms generally produce reads shorter than the several-hundred-to-thousand-base reads Sanger sequencing could manage, which makes reconstructing repetitive or structurally complex regions of a genome harder, since short overlapping fragments give a computer program less unambiguous information to piece back together. Specific chemistries carried their own weaknesses too — semiconductor-based platforms, for instance, struggled particularly with runs of the same repeated base. These are documented trade-offs rather than fatal flaws, but they explain why ‘sequencing a genome’ does not mean the same thing, in terms of completeness or reliability, on every platform.

The long-read answer to that trade-off

Long-read, third-generation platforms answer that specific weakness rather than the cost question. Pacific Biosciences’ real-time single-molecule method and Oxford Nanopore’s technology can each produce individual reads tens of thousands of bases long, with Nanopore reads reported reaching well over a million bases in some cases, which resolves repetitive stretches that defeat short-read assembly. The cost is a materially lower single-read accuracy than second-generation platforms typically achieve, though consensus accuracy across multiple overlapping long reads can still be very high. Nanopore’s hardware has also become genuinely portable, described as palm-sized, with results streamed in real time rather than produced only at the end of a run.

Worth knowing before trusting a headline genome figure

This is worth understanding less for any single fact than for the habit of scepticism it supports: a stated cost or turnaround time for ‘sequencing a genome’ means very little without knowing which platform produced it and what it traded away to do so. The story is one of competing engineering choices rather than a single decisive breakthrough, and readers who want to follow genomic medicine, agricultural genomics or outbreak surveillance as they are reported in the news will get more out of that context than out of any individual number. It rewards patience with a landscape of platforms rather than a tidy single narrative, but the reward is a much better filter for judging future sequencing claims.

Same field · Genetics4 of 57
Up next in Science

Nicolaus Copernicus

· 10:54

Copernicus didn’t prove the Earth moves—he made it impossible to ignore that it must.

10:54