HarmonyFidelisHarmonyFidelis
Login
NewsMajor ProjectsActorsAcademy

Codebook: 130 missing words in the language that switches our genes on

By combining five assays and 4,804 experiments on 332 poorly characterized transcription factors and 61 controls, Codebook obtains motifs for 177 of the 332 candidates and expands the known human lexicon by at least 15%.

Source: Nature

Codebook: 130 missing words in the language that switches our genes on

Several hundred human proteins are suspected of regulating our genes even though the DNA sequence they recognize is unknown. The Codebook consortium obtains a reproducible motif for 177 of the 332 candidate proteins: about 130 motifs are new. The analysis also maps 113,577 direct, conserved sites, including 82,760 for Codebook proteins and 30,817 for controls. For anyone interpreting genetic variants, the dictionary has just gained a chapter.

Source: nature.com

In plain terms

Our genes are not permanently switched on: proteins called “transcription factors” bind to short DNA words near genes to decide which ones are expressed and when. We know the word recognized by most of these proteins—but for several hundred of them, identified only because they resemble regulators, the word remained unknown. A large international collaboration tested 332 of these mysterious proteins with five complementary techniques and more than 4,000 experiments. Half of them yielded their word, making it possible to predict where they bind in the genome and how a single-letter DNA change can break their foothold. This is neither a drug nor a clinical test: it is a fundamental building block for reading the genome, especially when a sequence variant lands directly on one of these words.

Technical brief — Discovery

ParameterValue
Publication date5 August 2026
JournalNature (peer-reviewed article)
TitleAn expanded codebook of human transcription factor DNA-binding specificity
DOI10.1038/s41586-026-10798-9
Corresponding authorsBart Deplancke, Ivan V. Kulakovskiy and Timothy R. Hughes—Codebook–GRECO-BIT collaboration
Proteins analysed393 (332 putative transcription factors + 61 controls)
Experimental volume4,804 experiments (4,377 on Codebook proteins and 427 on controls)
Assays usedHT-SELEX · GHT-SELEX · SMiLE-seq · PBM arrays · ChIP–seq (HEK293 cells)
Motifs obtained177/332 putative factors (53%) + 58/61 controls
Net contribution≈ 130 novel motifs → human lexicon expanded by at least 15% (new total: 1,421 characterized TFs)
Conserved sites mapped113,577 CTOP sites: 82,760 Codebook + 30,817 controls (1,394,530 bases)
Data repositoryCisBP build 3.1, accession Jolma,2026a; humantfs.ccbr.utoronto.ca
StatusPeer-reviewed, project conducted over ~6 years

Technical explanation

  1. What a motif physically is. A transcription factor (TF) recognizes DNA through a DNA-binding domain (DBD) whose surface fits the groove of the double helix. Specific contacts—hydrogen bonds and van der Waals interactions—between certain amino acids and the edges of base pairs read in the major groove determine sequence preference. This preference is modelled as a position weight matrix (PWM): the relative probability of A, C, G and T at every position in the site. The PWM is not the protein; it is the statistical summary of what the protein tolerates at each letter.

  2. Why only half yield a motif—and what that means. Among 332 putative factors, 177 (53%) produce a reproducible motif. The C2H2 zinc-finger class dominates (121/180, or 67%); by contrast, among 49 proteins without a recognized DNA-binding domain, only 6 (12%) produce a motif—and all six ultimately contain a plausible binding region. The study does not turn every experimental failure into proof of absence: re-curation classifies 83 of the 155 failures as probable true negatives. The others may, for example, require a partner, a DNA modification or a post-translational modification that was absent from the assays.

  3. How the methods work—what each assay proves. The study's strength lies in combining measurements that fail and succeed for different reasons, making a motif retained by two platforms unlikely to be an artefact.

AssayWhat it measuresWhat it demonstrates
HT-SELEXIterative selection on random DNAIntrinsic preference outside genomic context
GHT-SELEXHT-SELEX on fragmented genomic DNAPreference in the context of real genomic sequences
SMiLE-seqBinding of single TFs or dimers in microfluidicsMotifs of complexes, with low bias
PBMArrays of defined double-stranded probesExhaustive scan of short k-mers
ChIP–seqSites occupied in vivo (HEK293)Where the protein actually binds in a cell

The success criterion is explicitly cross-platform: a motif is retained only if it is similar across ≥ 2 platforms and predicts the data from the others. This convergence, rather than any isolated measurement, provides the reliability.

  1. From motifs to probably important sites: the “TOP” logic. A motif alone overpredicts because the genome is saturated with chance matches. The authors define TOP sites (triple optimized overlap) as the intersection of three lines of evidence: a ChIP–seq peak (in vivo occupancy), a GHT-SELEX signal (direct in vitro binding) and a PWM match—each with its own threshold. The result is base-resolution sites bound directly in the assays. By comparing motif conservation (phyloP, an alignment of 241 Zoonomia mammals) with that of the surrounding sequence, the team isolates 113,577 conserved CTOP sites—82,760 for Codebook proteins and 30,817 for controls—covering about 1.39 million bases. This enrichment strongly supports their functional importance without demonstrating the individual function of every site.

Why It Worked

The most telling quantitative comparison is this: grouping the 177 Codebook motifs with the already known motifs of 1,211 human TFs—7.8× more proteins—produces only 4.75× more distinct motif families, and 92 to 135 of those families, depending on the method, contain only Codebook motifs. In other words, these long-overlooked proteins provide a disproportionate share of novelty: at least 15% more words in a lexicon thought to be nearly complete. The result remains robust when the motif database, similarity metric and clustering procedure are changed.

Measured predictive utility: in the study's HEK293 data, the Codebook PWMs distinguish binding sites from random sequences with a median AUROC of 0.71; the value is slightly lower in other cell types. Across 2,260 highly disruptive variants, the experimentally observed allelic preference agrees with the PWM prediction in 74% of cases.

One claim–evidence gap must be kept in mind: “expanding the lexicon by 15%” is firmly established and robust. The study does not, however, claim that all 177 proteins are active regulators in any given physiological context—ChIP–seq was performed in a single cell line (HEK293), and tissue-specific activity was not tested. The limitations are explicitly acknowledged: only five assays, one cell line and nearly half of the putative factors without a motif, many of which are probably not sequence binders.

Causal Chain

2018 census of human TFs → finding: several hundred “putative” TFs without a known motif → Codebook project (5 assays, nearly 6 years) → 4,804 experiments on 393 proteins → 177 reproducible motifs (including ~130 novel ones) → ChIP–seq ∩ GHT-SELEX ∩ PWM intersection → 113,577 conserved CTOP sites (82,760 Codebook + 30,817 controls) → prediction of variant effects (74% allelic concordance across 2,260 highly disruptive cases) → enrichment in disease SNPs (OR 1.35) and GTEx eQTLs (OR 1.34) → curated total of 1,421 human TFs with characterized sequence specificity, combining Codebook successes with 33 recent external contributions.

Anecdote

The name “Codebook” directly describes the project's purpose: each transcription factor recognizes a sequence “word” in the genome. The C2H2 zinc-finger class dominates the results, with motifs obtained for 121 of the 180 candidate proteins in this family. Yet the dictionary remains incomplete: a motif describes a binding preference, not by itself the biological role of a protein in every cell.

Legacy and Current Data

The reference census of human TFs (Lambert et al., Cell 2018) left several hundred “putative” proteins without a known specificity. The 177 Codebook successes, supplemented by 33 TFs found in recent external datasets, bring the curated total to 1,421 human TFs with characterized sequence specificity; representative PWMs appear in Supplementary Data 1. Codebook data are deposited in CisBP build 3.1 under accession Jolma,2026a. For a non-coding variant, a motif can predict a binding disruption, within the limits measured here: a median AUROC of 0.71 in HEK293 and 74% allelic concordance among highly disruptive cases. A motif alone is not enough to declare a site functional or a consequence clinically meaningful.

The researcher's view — open questions

(Interpretation, not results of the study.) The decisive extensions are: (1) tissue portability—repeat ChIP–seq in several cell lines to distinguish universal sites from context-dependent ones, the only way to determine what the 177 TFs regulate and where; (2) targeted disruptive mutation—abolish a specificity residue in a zinc finger and verify the loss of the expected motif, providing causal evidence that the residue carries the readout; and (3) the 155 silent factors—a subset may require an obligate partner, a DNA modification such as methylation or a post-translational modification absent in vitro; retesting them under those conditions would distinguish “true negatives” from “methodological false negatives.”

“TOP” site = intersection of three independent lines of evidence ChIP–seq occupancy in vivo GHT-SELEX binding in vitro PWM motif sequence match TOP base resolution + evolutionary conservation (phyloP) → probably important CTOP site

Sources

References verified during the fact-checking audit (accessed 6 August 2026).

  • Codebook–GRECO-BIT Collaboration (Hughes, T. R. et al.). An expanded codebook of human transcription factor DNA-binding specificity. Nature, published 5 August 2026. DOI: 10.1038/s41586-026-10798-9. (Peer-reviewed article. Data: CisBP build 3.1, accession Jolma,2026a.)

Background references

(Context distinct from the study's primary source.)

  • Lambert, S. A. et al. The human transcription factors. Cell 172, 650–665 (2018).
  • Stormo, G. D. & Zhao, Y. Determining the specificity of protein–DNA interactions. Nat. Rev. Genet. 11, 751–760 (2010).
  • Najafabadi, H. S. et al. C2H2 zinc finger proteins greatly expand the human regulatory lexicon. Nat. Biotechnol. 33, 555–562 (2015).

Confidence statement. Solidly established: motif characterization for 177/332 putative TFs, the addition of about 130 motifs (at least 15% more lexicon), the 113,577 conserved CTOP sites—82,760 Codebook and 30,817 controls—and the measured predictive value (median AUROC 0.71 in HEK293; 74% allelic concordance across 2,260 highly disruptive cases). Uncertain or not tested here: the tissue-specific regulatory activity of these TFs, the individual function of every CTOP site and the cause of every motif failure. Technical checks: DOI resolved, exact experiment volume, CTOP breakdown and composition of the 1,421 total corrected. Deduplication verdict: CONTROLLED BACKLOG RECOVERY—the DOI key already exists as A005 in the registry, but no corresponding bulletin exists in the database; this recovery restores the historical content without creating a new angle or registry entry.