Several hundred human proteins are suspected of regulating our genes even though the DNA sequence they recognize is unknown. The Codebook consortium obtains a reproducible motif for 177 of the 332 candidate proteins: about 130 motifs are new. The analysis also maps 113,577 direct, conserved sites, including 82,760 for Codebook proteins and 30,817 for controls. For anyone interpreting genetic variants, the dictionary has just gained a chapter.
Source: nature.com
In plain terms
Our genes are not permanently switched on: proteins called “transcription factors” bind to short DNA words near genes to decide which ones are expressed and when. We know the word recognized by most of these proteins—but for several hundred of them, identified only because they resemble regulators, the word remained unknown. A large international collaboration tested 332 of these mysterious proteins with five complementary techniques and more than 4,000 experiments. Half of them yielded their word, making it possible to predict where they bind in the genome and how a single-letter DNA change can break their foothold. This is neither a drug nor a clinical test: it is a fundamental building block for reading the genome, especially when a sequence variant lands directly on one of these words.
Technical brief — Discovery
| Parameter | Value |
|---|---|
| Publication date | 5 August 2026 |
| Journal | Nature (peer-reviewed article) |
| Title | An expanded codebook of human transcription factor DNA-binding specificity |
| DOI | 10.1038/s41586-026-10798-9 |
| Corresponding authors | Bart Deplancke, Ivan V. Kulakovskiy and Timothy R. Hughes—Codebook–GRECO-BIT collaboration |
| Proteins analysed | 393 (332 putative transcription factors + 61 controls) |
| Experimental volume | 4,804 experiments (4,377 on Codebook proteins and 427 on controls) |
| Assays used | HT-SELEX · GHT-SELEX · SMiLE-seq · PBM arrays · ChIP–seq (HEK293 cells) |
| Motifs obtained | 177/332 putative factors (53%) + 58/61 controls |
| Net contribution | ≈ 130 novel motifs → human lexicon expanded by at least 15% (new total: 1,421 characterized TFs) |
| Conserved sites mapped | 113,577 CTOP sites: 82,760 Codebook + 30,817 controls (1,394,530 bases) |
| Data repository | CisBP build 3.1, accession Jolma,2026a; humantfs.ccbr.utoronto.ca |
| Status | Peer-reviewed, project conducted over ~6 years |
Technical explanation
What a motif physically is. A transcription factor (TF) recognizes DNA through a DNA-binding domain (DBD) whose surface fits the groove of the double helix. Specific contacts—hydrogen bonds and van der Waals interactions—between certain amino acids and the edges of base pairs read in the major groove determine sequence preference. This preference is modelled as a position weight matrix (PWM): the relative probability of A, C, G and T at every position in the site. The PWM is not the protein; it is the statistical summary of what the protein tolerates at each letter.
Why only half yield a motif—and what that means. Among 332 putative factors, 177 (53%) produce a reproducible motif. The C2H2 zinc-finger class dominates (121/180, or 67%); by contrast, among 49 proteins without a recognized DNA-binding domain, only 6 (12%) produce a motif—and all six ultimately contain a plausible binding region. The study does not turn every experimental failure into proof of absence: re-curation classifies 83 of the 155 failures as probable true negatives. The others may, for example, require a partner, a DNA modification or a post-translational modification that was absent from the assays.
How the methods work—what each assay proves. The study's strength lies in combining measurements that fail and succeed for different reasons, making a motif retained by two platforms unlikely to be an artefact.
| Assay | What it measures | What it demonstrates |
|---|---|---|
| HT-SELEX | Iterative selection on random DNA | Intrinsic preference outside genomic context |
| GHT-SELEX | HT-SELEX on fragmented genomic DNA | Preference in the context of real genomic sequences |
| SMiLE-seq | Binding of single TFs or dimers in microfluidics | Motifs of complexes, with low bias |
| PBM | Arrays of defined double-stranded probes | Exhaustive scan of short k-mers |
| ChIP–seq | Sites occupied in vivo (HEK293) | Where the protein actually binds in a cell |
The success criterion is explicitly cross-platform: a motif is retained only if it is similar across ≥ 2 platforms and predicts the data from the others. This convergence, rather than any isolated measurement, provides the reliability.
- From motifs to probably important sites: the “TOP” logic. A motif alone overpredicts because the genome is saturated with chance matches. The authors define TOP sites (triple optimized overlap) as the intersection of three lines of evidence: a ChIP–seq peak (in vivo occupancy), a GHT-SELEX signal (direct in vitro binding) and a PWM match—each with its own threshold. The result is base-resolution sites bound directly in the assays. By comparing motif conservation (phyloP, an alignment of 241 Zoonomia mammals) with that of the surrounding sequence, the team isolates 113,577 conserved CTOP sites—82,760 for Codebook proteins and 30,817 for controls—covering about 1.39 million bases. This enrichment strongly supports their functional importance without demonstrating the individual function of every site.
Why It Worked
The most telling quantitative comparison is this: grouping the 177 Codebook motifs with the already known motifs of 1,211 human TFs—7.8× more proteins—produces only 4.75× more distinct motif families, and 92 to 135 of those families, depending on the method, contain only Codebook motifs. In other words, these long-overlooked proteins provide a disproportionate share of novelty: at least 15% more words in a lexicon thought to be nearly complete. The result remains robust when the motif database, similarity metric and clustering procedure are changed.
Measured predictive utility: in the study's HEK293 data, the Codebook PWMs distinguish binding sites from random sequences with a median AUROC of 0.71; the value is slightly lower in other cell types. Across 2,260 highly disruptive variants, the experimentally observed allelic preference agrees with the PWM prediction in 74% of cases.
One claim–evidence gap must be kept in mind: “expanding the lexicon by 15%” is firmly established and robust. The study does not, however, claim that all 177 proteins are active regulators in any given physiological context—ChIP–seq was performed in a single cell line (HEK293), and tissue-specific activity was not tested. The limitations are explicitly acknowledged: only five assays, one cell line and nearly half of the putative factors without a motif, many of which are probably not sequence binders.
Causal Chain
2018 census of human TFs → finding: several hundred “putative” TFs without a known motif → Codebook project (5 assays, nearly 6 years) → 4,804 experiments on 393 proteins → 177 reproducible motifs (including ~130 novel ones) → ChIP–seq ∩ GHT-SELEX ∩ PWM intersection → 113,577 conserved CTOP sites (82,760 Codebook + 30,817 controls) → prediction of variant effects (74% allelic concordance across 2,260 highly disruptive cases) → enrichment in disease SNPs (OR 1.35) and GTEx eQTLs (OR 1.34) → curated total of 1,421 human TFs with characterized sequence specificity, combining Codebook successes with 33 recent external contributions.
Anecdote
The name “Codebook” directly describes the project's purpose: each transcription factor recognizes a sequence “word” in the genome. The C2H2 zinc-finger class dominates the results, with motifs obtained for 121 of the 180 candidate proteins in this family. Yet the dictionary remains incomplete: a motif describes a binding preference, not by itself the biological role of a protein in every cell.
Legacy and Current Data
The reference census of human TFs (Lambert et al., Cell 2018) left several hundred “putative” proteins without a known specificity. The 177 Codebook successes, supplemented by 33 TFs found in recent external datasets, bring the curated total to 1,421 human TFs with characterized sequence specificity; representative PWMs appear in Supplementary Data 1. Codebook data are deposited in CisBP build 3.1 under accession Jolma,2026a. For a non-coding variant, a motif can predict a binding disruption, within the limits measured here: a median AUROC of 0.71 in HEK293 and 74% allelic concordance among highly disruptive cases. A motif alone is not enough to declare a site functional or a consequence clinically meaningful.
The researcher's view — open questions
(Interpretation, not results of the study.) The decisive extensions are: (1) tissue portability—repeat ChIP–seq in several cell lines to distinguish universal sites from context-dependent ones, the only way to determine what the 177 TFs regulate and where; (2) targeted disruptive mutation—abolish a specificity residue in a zinc finger and verify the loss of the expected motif, providing causal evidence that the residue carries the readout; and (3) the 155 silent factors—a subset may require an obligate partner, a DNA modification such as methylation or a post-translational modification absent in vitro; retesting them under those conditions would distinguish “true negatives” from “methodological false negatives.”
Sources
References verified during the fact-checking audit (accessed 6 August 2026).
- Codebook–GRECO-BIT Collaboration (Hughes, T. R. et al.). An expanded codebook of human transcription factor DNA-binding specificity. Nature, published 5 August 2026. DOI: 10.1038/s41586-026-10798-9. (Peer-reviewed article. Data: CisBP build 3.1, accession
Jolma,2026a.)
Background references
(Context distinct from the study's primary source.)
- Lambert, S. A. et al. The human transcription factors. Cell 172, 650–665 (2018).
- Stormo, G. D. & Zhao, Y. Determining the specificity of protein–DNA interactions. Nat. Rev. Genet. 11, 751–760 (2010).
- Najafabadi, H. S. et al. C2H2 zinc finger proteins greatly expand the human regulatory lexicon. Nat. Biotechnol. 33, 555–562 (2015).
Confidence statement. Solidly established: motif characterization for 177/332 putative TFs, the addition of about 130 motifs (at least 15% more lexicon), the 113,577 conserved CTOP sites—82,760 Codebook and 30,817 controls—and the measured predictive value (median AUROC 0.71 in HEK293; 74% allelic concordance across 2,260 highly disruptive cases). Uncertain or not tested here: the tissue-specific regulatory activity of these TFs, the individual function of every CTOP site and the cause of every motif failure. Technical checks: DOI resolved, exact experiment volume, CTOP breakdown and composition of the 1,421 total corrected. Deduplication verdict: CONTROLLED BACKLOG RECOVERY—the DOI key already exists as A005 in the registry, but no corresponding bulletin exists in the database; this recovery restores the historical content without creating a new angle or registry entry.
