An open question about proteins
By the early 1940s it was known that proteins were built from amino acids, but it was not settled whether a given protein had one fixed sequence of them, repeated identically in every molecule, or something closer to a variable mixture without a single exact chemical identity. Frederick Sanger, working in Cambridge’s biochemistry department from 1943, set out to answer this using insulin, a small and relatively well-characterised hormone that was also available in reasonably pure and plentiful form because it was already manufactured for diabetes treatment. The claim his work eventually supported was specific and, at the time, not obvious: that insulin consists of exactly 51 amino acids arranged in a single, invariant order, split across two separate chains held together by chemical bonds.
Sanger’s reagent and the fingerprint method
Sanger’s method combined several separate techniques into one procedure. He first labelled the amino group at the very end of a protein chain using a reagent he developed, 1-fluoro-2,4-dinitrobenzene, now known as Sanger’s reagent, which stuck permanently to that end and let him identify which amino acid sat there. He then broke the protein into smaller, overlapping fragments using partial hydrolysis with hydrochloric acid or the enzyme trypsin, and separated those fragments using two-dimensional paper chromatography combined with electrophoresis, producing a distinctive pattern of spots for each fragment that he called a fingerprint, visualised by staining with ninhydrin. Piecing together the overlapping fragments let him reconstruct the full order of amino acids along each chain.
Reading the two chains
The work proceeded in stages that hold up as a clear, sequential record. Sanger completed the sequence of insulin’s shorter B chain, 30 amino acids long, by 1951, and the A chain, 21 amino acids long, the following year, with collaborators including Hans Tuppy and E.O.P. Thompson contributing to the published results across 1951 to 1953. He then worked out how the two chains connect, establishing by 1955 that two disulfide bonds link the A and B chains to each other and a third bonds within the A chain alone. The completed structure, a defined sequence of 51 amino acids with a fixed pattern of disulfide bonds, held up as insulin’s genuine chemical structure and has not needed revision since.
What the sequence actually proved
What Sanger’s sequencing could not establish was insulin’s three-dimensional shape: knowing the order in which amino acids are strung together says nothing directly about how that chain folds up in space, which is what ultimately determines a protein’s biological activity. That question was answered separately, by X-ray crystallography, when Dorothy Hodgkin’s group worked out insulin’s three-dimensional structure in 1969, nearly two decades after Sanger’s sequence was complete. Sanger’s method was also slow and laborious by later standards, requiring years of manual fragment separation and reassembly for a molecule of only 51 amino acids; it would not have scaled easily to larger proteins without the faster techniques that followed, some of which Sanger himself later developed.
What sequencing could not yet do
The immediate significance went beyond insulin itself. By demonstrating that a protein has one specific, determinable sequence, Sanger’s result underpinned the idea that genetic information specifies proteins precisely, a link central to what became molecular biology’s account of how genes work. His 1958 Nobel Prize in Chemistry recognised exactly this: not insulin’s biology, which was already well studied via Banting and Best’s original 1921 isolation of the hormone, but the demonstration that protein sequencing was possible at all, plus a method other researchers could apply to different proteins. Insulin’s status as the first fully sequenced protein also made it a natural test case later, when it became the first protein both chemically synthesised and produced through recombinant DNA technology.
A method, not just a molecule
This is worth an hour for anyone who wants to understand what came before genetic sequencing existed, since Sanger’s insulin work predates DNA sequencing by over two decades and used entirely different, more laborious chemical methods to answer a related question: what exactly is this molecule made of, in what order. The reward is seeing a genuinely patient piece of experimental chemistry, spread across a full decade of Sanger’s career, settle a real uncertainty about whether proteins even have fixed identities. It also sets up a satisfying throughline, since the same fingerprinting instincts Sanger developed on insulin’s amino acids he later adapted, first to RNA and then to DNA, winning a second Nobel Prize in 1980 for the sequencing method still used as a byword for the technique today.