About

Mycobrowser, the reference annotation resource for Mycobacterium, is no longer maintained, and a large share of MTBC genes are still labelled hypothetical protein. This site takes over that role and re-annotates the MTBC gene set, with two distinctive choices.

Continuously updated. This live version serves 3974 MTBC genes (1794 requalified, 1950 family-assigned, 230 still unknown) plus 200 microproteins. Counts on this page are read live from the database; the accompanying manuscript describes the 3,906-gene submission set, so its figures may differ. Download the per-gene changelog.

Anchored on the ancestral genome

Annotation starts from MTBC0, the imputed ancestral genome of the most recent common ancestor of the MTBC (Harrison et al. 2024), rather than from H37Rv. Each protein shown here is the ancestral MTBC0 sequence; H37Rv remains as the historical anchor (its Rv locus tag and legacy product).

A traceable, graded pipeline

Every gene goes through: the MTBC0 PGAP re-annotation, a Pfam domain scan (hmmscan --cut_ga), the ESM Atlas protein-language-model signal (used only as an exploratory indicator), and a literature check. Each fiche carries an explicit verdict (Resolved, Family assigned, Still unknown), a confidence level, and a Sources section that cites the provenance of every field.

Reading a conjunction of criteria: correlated, not independent

A single evidence layer here (dark verdict, no Pfam domain, undetected by mass spectrometry, few TA sites) is informative on its own, but they are not statistically independent of each other: all four correlate strongly with gene length. Measured across the atlas (2026-07-31, signalled by a companion project's characterisation of Rv2438A): between CDS under 150 nt and over 1200 nt, the dark rate is 38× higher (38% vs 1%), the no-domain rate 17× higher (50% vs 3%), and the undetected-in-MS rate higher (62% vs 7%). Reading “dark AND short AND no domain AND undetected” as four independent confirmations therefore overstates the evidence: it is closer to counting the same underlying fact (a short ORF gives every method less material to work with) up to four times. This does not invalidate any individual layer — it is a caution about how to read their conjunction, most relevant for very short genes — each gene page flags essentiality calls resting on very few Himar1 TA sites for the same reason.

The codon-position test: a confirmatory tool, not a general screen

A reading frame under purifying selection accumulates substitutions preferentially at the third codon position. This statistic is implemented as a reusable, standalone module (codon_phase.py, released with a companion project's characterisation of Rv2438A) and can in principle be run on any gene, but it is not recommended as a general screen without its measured limits in view. On five H37Rv genes whose protein is directly attested by mass spectrometry, it recovers only two of five as coding: specific (no false positive observed across these controls) but insensitive, so a negative result carries a likelihood ratio of only about 1.7. The reason is biological, not a shortfall of data: the statistic tracks purifying constraint rather than translation as such — the third-position fraction correlates with a locus's variant density (Spearman ρ = −0.534), not with gene length (ρ = +0.159, p = 0.23) — so a gene that is genuinely translated but only weakly constrained stays invisible to it, and more sequencing depth would not change that. Below a density of 0.55 variants per nucleotide the test recovers most coding genes; above it, only a minority — a condition met by roughly a third of essential genes on this atlas, and by very few non-essential ones. The tool enforces this itself: above the threshold it reports a negative verdict as uninformative rather than as an absence of coding, and it refuses to run without a positive control of comparable length in the same query.

What the atlas changed vs Mycobrowser

A per-gene changelog (one row per gene) records, for every gene, the Mycobrowser legacy annotation, the atlas result (verdict, revised function, EC), and what is new: the novelty class, whether the atlas is ahead of Mycobrowser, the EC status, the evidence that provided the handle, and the annotation route (literature-curated vs orthology/structure). Highlights, in the companion analysis (the 3,906-gene submission set): 869 of the 1013 Mycobrowser “conserved hypotheticals” gained a functional handle and 304 EC numbers were updated. In the live resource today, 230 of the 3974 genes remain honestly unknown.

⇩ Download the per-gene changelog (CSV) · also queryable via the JSON API.

Anchored in the literature — including under other names

Every gene carries an In the literature section. Publications are found not only under the H37Rv locus tag, but also under the gene name and under the identifiers the gene bears in related species (M. bovis, M. marinum, M. smegmatis, M. leprae, M. abscessus). This matters: a gene may be well characterised under the name of its ortholog and completely invisible to a search on its Rv tag. Each hit is verified against the abstract text, so a stated absence of literature is a verified absence — itself an informative statement about an annotation with no primary study behind it.

Two consequences are made explicit on the fiche. A gene can be dark yet heavily studied (Rv2660c has no known molecular function, but is a component of the H56 vaccine candidate in clinical trials), which separates “dark because nobody looked” from “dark despite being studied”. And a gene can have its biology documented mostly outside M. tuberculosis, because mycobacterial genetics is largely done in M. smegmatis; the fiche says so, and warns that such findings do not transfer automatically to M. tuberculosis.

Code and data availability

The annotation pipeline and the derived per-gene records are openly available at https://github.com/cguyeux/mtbc-gene-atlas. The pipeline and application code are released under the MIT licence; the annotation records and derived data under CC-BY 4.0. Redistributed third-party annotations retain the licences of their sources (UniProt and STRING under CC-BY 4.0; Pfam; eggNOG), listed in the Sources section of every fiche.

Cite this resource

Christophe Guyeux. From hypothetical to functional: a continuously updated, structure- and population-scale re-annotation of the Mycobacterium tuberculosis complex gene set, anchored on the MTBC0 ancestral genome. FEMTO-ST Institute, CNRS UMR 6174, Université Marie et Louis Pasteur (Université de Franche-Comté), Besançon, France. Archived on Zenodo: doi:10.5281/zenodo.20815246 (concept DOI, always resolves to the latest release) or doi:10.5281/zenodo.20815247 (the exact snapshot the manuscript was written against).

Citing a stable link

Because this atlas is updated continuously, cite one of two persistent identifiers rather than a hosting URL. For the living resource, use the stable domain https://mtbc.gclab.fr, which is author-controlled and resolves to the current instance independently of the hosting provider, so the service can be migrated without breaking published links. For a fixed, quotable state, cite the Zenodo concept DOI above (latest release) or a version DOI (one specific release). Provider-specific endpoints are deliberately not advertised: they change when the deployment changes.

FEMTO-ST Institute, CNRS UMR 6174, Université Marie et Louis Pasteur (Université de Franche-Comté), Besançon, France. Content under CC-BY 4.0.