Discovery
| Parameter | Value |
|---|---|
| Dates | 1990 (launch) → April 2003 (final publication) |
| Coordinator | Francis Collins (NIH) |
| Consortium | 20 centers, 6 countries (USA, UK, Japan, France, Germany, China) |
| Total cost | ~$2.7 billion (1991 dollars) |
| Genome size | ~2.85 billion nucleotides (finished sequence, 2003–2004) |
| Protein-coding genes | ~20,000–25,000 (vs 100,000 estimated) |
| Coverage | >99% of euchromatic regions |
Technical Explanation
1. BAC-by-BAC method (public HGP) — The genome is cut into 100–200 kb fragments inserted into bacterial artificial chromosomes (BAC). Each BAC is physically mapped to a chromosome, then subcloned into 1–2 kb fragments sequenced by the Sanger method (chain termination with fluorescent ddNTPs). Reads of ~700 bp are assembled by overlap.
2. Shotgun method (Celera, Craig Venter) — No prior mapping: genomic DNA is randomly fragmented, massively sequenced, then assembled by algorithms (overlapping reads, scaffolds). Faster but more sensitive to repetitive regions (~45% of the human genome consists of repeated sequences).
3. Assembly and annotation — Assembly requires supercomputers: aligning millions of reads (an NP-hard problem in theory, solved by heuristics — BLAST, Phrap, ARACHNE). Annotation identifies open reading frames (ORFs), promoters, and alternative splicing events.
4. Gap filling (2003–2022) — The 2003 version contained ~340 gaps in heterochromatin and centromeres. The Telomere-to-Telomere (T2T) consortium filled them in 2022 using long-read sequencing (Oxford Nanopore, PacBio HiFi): continuous runs of 10–100 kb that span repetitive regions.
Why It Worked
The HGP strategy — map then sequence — was slower but produced a high-quality assembly essential for a reference genome. Competition with Celera accelerated the project by two years. The open-access data policy (Bermuda Principles, 1996: publication within 24 hours) established a foundational precedent for open science in biology.
When the HGP revealed that humans have only ~20,000 genes (compared to ~13,000 for Drosophila and ~20,000 for the nematode C. elegans), the scientific community was stunned: human complexity does not stem from gene count but from regulation — alternative splicing, non-coding RNAs, and epigenetic modifications. The "book of life" turned out to be more of a grammar book than a dictionary.
Causal Chain
DNA structure (Watson-Crick, 1953) → Sanger sequencing (1977) → PCR (Mullis, 1983) → HGP launch (1990) → Bermuda Principles open access (1996) → Celera accelerates competition (1998) → Draft published (Nature/Science, 2001) → Final version (2003) → GWAS, precision medicine → $200 sequencing → T2T completed (2022)
Anecdote
When the HGP revealed that humans possess only ~20,000 genes (compared with ~13,000 for the fruit fly and ~20,000 for the roundworm C. elegans), the community was stunned: human complexity does not come from gene count but from gene regulation (alternative splicing, non-coding RNAs, epigenetic modifications). The "book of life" turned out to be more of a grammar book than a dictionary.
Sources
References verified during the August 2026 fact-checking audit: these are the pages
against which this bulletin's claims were checked.
