HarmonyFidelisHarmonyFidelis
Login
NewsMajor ProjectsActorsAcademy
Back to ecosystems
Reading our DNA. Building a shared reference.

LIFE SCIENCES / GENOMICS

Human Genome Project

Reading our DNA. Building a shared reference.

Completed project1990 – 2003
View roadmapSources

AI-generated conceptual illustration — not experimental data.

How it worked

Map → Sequence → Assemble → Check & share
AI-generated conceptual illustration — not experimental data.
  1. 01Map

    Extract DNA and construct large-fragment clone libraries; Map overlapping BAC clones to genomic regions

  2. 02Sequence

    Sequence smaller fragments from selected clones

  3. 03Assemble

    Assemble overlapping reads, then order clone sequences

  4. 04Check & share

    Check inconsistencies, finish resolvable regions and release the reference

Read the evidence and caveats

Purpose

Completed in 2003, the Human Genome Project made a shared DNA reference available for studying human biology. It did not read every region or capture every person's genetic variation. Later T2T and pangenome projects addressed different parts of those limitations; their results are not achievements of the completed programme.

20 sequencing centres[S1]

6 countries[S1]

1990–2003 Programme period[S6]

DNA Open reference[S1]

AI drafted the English reference and translated it into seven languages. Source checking and automated structural and numerical checks were performed. No human or independent expert validation is claimed. Editorial method

Status, results & limitations

What completion did not mean

Euchromatin contains most genes, but is not the whole genome. Highly repetitive heterochromatic regions, including centromeric regions, remained difficult to resolve. A high percentage of euchromatic coverage therefore cannot be presented as the same percentage of all human DNA.

Read the evidence and caveats

People & organisations

20 centres in the United States, United Kingdom, France, Germany, Japan and China.
Participating countries. Geographic context; no contribution shares are represented. 20 centres in the United States, United Kingdom, France, Germany, Japan and China. Natural Earth. [S1][S6]

NHGRI · U.S. Department of Energy · Wellcome Trust

United States: programme coordination. United Kingdom: funding and support.

[S1][S6]

Major sequencing contributors

Sanger · Baylor · Washington University · Whitehead/MIT · DOE Joint Genome Institute

[S1]

Francis Collins

US programme leadership from 1993.

[S1]

John Sulston

Sanger leadership and, with Robert Waterston, adoption of the Bermuda data-sharing principles in 1996.

[S1]

Roadmap & legacy

Milestones are ordered; spacing is not a time scale.

Programme

1990 – 2003

  1. 1990Launch1990 — Launch[S6]
  2. 2001Draft publication2001 — Draft publication[S2]
  3. 2003Programme completed2003 — Programme completed[S1]

Technical publications

Technical follow-up

  1. 2004Finishing paper2004 — Finishing paper[S3]

Independent extensions

Separate research programmes

  1. 2022T2T2022 — T2T[S4]
  2. 2023Pangenome2023 — Pangenome[S5]

Resources & evidence

Purpose: a common coordinate system for biology

A, C, G and T are the four letters used to represent the bases of DNA.

The programme sought to determine the order of DNA bases and provide maps and tools for finding and studying genes. A reference sequence gives researchers shared coordinates: observations from different laboratories can be located in the same genomic regions rather than described with incompatible local maps.

Its value was infrastructural. Producing a reusable sequence enabled subsequent questions about biological function, evolution and disease; it did not itself explain the function of every base or constitute a treatment.

[S1][S2][S6]

Method: map, sequence, assemble and finish

The public consortium used a hierarchical approach. Large DNA fragments, often maintained as bacterial artificial chromosomes (BACs), were mapped and selected to cover the genome. Each selected clone was subdivided, sequenced using Sanger chemistry, and computationally assembled from overlapping reads. The clone assemblies were then ordered into a larger reference.

BACs maintain large DNA fragments stably in bacteria, keeping mapped pieces available for sequencing. In Sanger sequencing, copies are stopped at specific bases and separated by length so that the base order can be reconstructed.

The map constrained where fragments belonged, reducing ambiguity. Repeated stretches could nevertheless look almost identical, while cloning biases left some sequences poorly represented. Finishing therefore required targeted additional reads and checks of overlaps and assembly consistency, not simply more automated sequencing. An assembled sequence is a reconstruction from measurements, not a photograph of a chromosome.

[S2][S7]

People, organisation and resources

The International Human Genome Sequencing Consortium brought together 20 centres in the United States, United Kingdom, France, Germany, Japan and China. NHGRI and the Department of Energy led the US effort; the Wellcome Trust supported the UK effort. Biologists, engineers and computational specialists shared a distributed production task.

The 2003 completion announcement compares an initial US estimate of USD 3 billion with a reported USD 2.7 billion, both in fiscal-year 1991 dollars. These are official programme figures, not an independently audited worldwide expenditure total. NHGRI's later account explicitly notes the difficulty of precise accounting across international funders. The sequence was released openly for reuse.

[S1][S6]

Chronology and demonstrated sequence results

The programme ran from 1990 to 2003. A working draft was announced in June 2000, followed by the public consortium's research paper in February 2001. On 14 April 2003, the consortium announced completion and made an essentially finished reference available. This programme milestone must be distinguished from the later technical publication.

In October 2004, the finishing paper reported Build 35: 2.85 billion nucleotides, 341 gaps, approximately 99% of euchromatin and an estimated error rate of about 1 event per 100,000 bases. Those are the paper's assembly-specific results, not exact measurements to backdate to the April 2003 announcement. That announcement described 99% of gene-containing regions with 99.99% accuracy. Coverage and accuracy measure different things.

[S1][S2][S3]

What completion did not mean

Human → cell → nucleus → condensed chromosome → DNA. Not to scale.
Human → cell → nucleus → condensed chromosome → DNA. Not to scale. AI-generated conceptual illustration — not experimental data.

Euchromatin contains most genes, but is not the whole genome. Highly repetitive heterochromatic regions, including centromeric regions, remained difficult to resolve. A high percentage of euchromatic coverage therefore cannot be presented as the same percentage of all human DNA.

The reference was also a mosaic assembled from several donors, not both chromosome copies of one representative person and not a catalogue of humanity's variation. Sequence completeness, correctness of assembly, population representation and biological interpretation are separate requirements. Success against the programme's finishing standard did not remove these limitations.

[S3][S4][S6]

T2T: a later advance in sequence continuity

Nurk and colleagues' paper, released on 31 March 2022, described T2T-CHM13, an assembly of 3.055 billion base pairs with no sequence gaps for the autosomes and chromosome X; chromosome Y was absent. PacBio HiFi and ultra-long Oxford Nanopore reads helped resolve repetitive regions that earlier approaches could not. The source cell line was nearly homozygous, simplifying the separation of chromosome copies.

The paper still identifies model sequences for some ribosomal DNA arrays whose internal order was not fully resolved. Thus “gapless” must not be interpreted as direct, unambiguous determination of every repeated copy. T2T was a later consortium's achievement, not a result delivered by the Human Genome Project in 2003, and one assembly cannot represent all human diversity.

[S4]

Pangenome: a later advance in representation

On 10 May 2023, Liao and colleagues published the first HPRC draft: 47 phased diploid assemblies, corresponding to 94 haplotypes. Phasing separates the chromosome sequences inherited from each parent. Aligning multiple assemblies into variation graphs provides alternative sequence paths instead of forcing everyone onto a single linear reference.

The paper reports 119 million additional base pairs of euchromatic polymorphic sequence relative to GRCh38. Its evaluated workflows reduced small-variant discovery errors by 34% and increased structural variants detected per haplotype by 104%. These are comparisons within the study, not universal improvements or evidence of a corresponding clinical benefit. The draft was neither a census of all variation nor error-free; difficult repeats and representation of additional populations remained challenges. GRCh38 is a later reference, not the unchanged 2003 sequence.

[S5]

Legacy, ethical questions and continuing work

A durable outcome is a publicly reusable reference and a culture of rapid sequence sharing, supported by the Bermuda principles. The resource made gene annotation and comparative genomic analysis more precise. Later T2T and pangenome results illustrate how subsequent programmes could test and extend that foundation.

The programme also supported research into ethical, legal and social implications: privacy, consent and potential genetic discrimination matter alongside technical performance. Better sequence data do not automatically establish causality for a disease or deliver a therapy; interpretation and clinical evaluation require additional evidence. The Human Genome Project remains completed. Future updates to this historical record should track documented reinterpretations and downstream results without assigning new deadlines or later achievements to the closed programme.

[S1][S3][S5][S6]

Sources

  1. [S1] International Consortium Completes Human Genome Project

    Consulted on 2026-09-10

  2. [S2] Initial sequencing and analysis of the human genome

    Consulted on 2026-09-10

  3. [S3] Finishing the euchromatic sequence of the human genome

    Consulted on 2026-09-10

  4. [S4] The complete sequence of a human genome

    Consulted on 2026-09-10

  5. [S5] A draft human pangenome reference

    Consulted on 2026-09-10

  6. [S6] Human Genome Project

    Consulted on 2026-09-10

  7. [S7] DNA sequencing with chain-terminating inhibitors

    Consulted on 2026-09-10

Editorial method

Relevant passages in original research papers and official accounts were read; dates, denominators, assembly versions and attribution were compared. All eight texts were checked against the same English reference, with automated checks of structure, citations and numerical consistency. Raw sequencing data were not reanalysed, experimental results were not replicated, and clinical benefits were not independently evaluated. This is a historical account with selected later milestones, not an exhaustive review of genomics through 2026.

AI drafted the English reference and translated it into seven languages. Source checking and automated structural and numerical checks were performed. No human or independent expert validation is claimed.

Revision 2 · 2026-09-10

Presentation update: illustrated dashboard, participant roles and sourced milestones added. The detailed scientific text is retained.