ProtNLM2 Evaluation
Bottom line: ProtNLM2 is Google DeepMind's sequence-to-text model that
writes protein names, GO terms and function paragraphs straight from sequence,
and UniProt ships its output on unreviewed entries. We assessed it the way a
curator would, protein by protein, using the COR/CNN/LSP/UNC/NPI/PLI/REP
categories of de Crécy-Lagard et al. 2025 (PMID:40703034). As of the review
snapshot (2026-09-27, commit c7551cb3db), the assessment covers 242 prediction
targets and 288 GO-term assessments, plus separate reviews of the narrative
function text.
Correct novel predictions outnumber correct but already-known ones, because the
purposive cohorts favour proteins without existing annotation, and contradicted
predictions are a small minority. Narrative text behaves differently from GO
terms: paralog confusion is common in the function paragraphs and nearly absent
from the GO claims.
The recurring failure is not a wrong family but a claim pitched above what the
sequence supports: a catalytic activity where the deposited sequence lacks the
catalytic region (wheat patatin A0A3B6GK97, Arabidopsis LRR F4JLB7), or a
localization that cannot exist in the organism (neuron projection in wheat
F6LAX4). The most common single verdict is uncertain, which is the honest
result for TrEMBL proteins with no direct characterization.
Function prediction evaluation index
Browse all ProtNLM predictions — filter prediction sets, narrative reviews, and individual GO claims by species, cohort, and assessment. Fly records include reviewed empty GO outputs.
Where the numbers are. The dated counts, deduplicated across overlapping
cohorts, are in the cross-cohort results
and benchmark-summary.json. The
browser gives the same breakdowns as facet counts over the current reviews:
all GO claims, one cohort at a time
(ARGO-50,
HORSE40,
FLY41,
POMBE20,
NEUROSPORA20,
MOD_EVOLUTION20), or only the
contradicted claims.
Narrative records can hold several judgments each and are counted separately from
GO terms.
The cross-species cohorts examine biological support, substrate and paralog
specificity, domain completeness, and transfer of annotations across species. The
original ARGO-ProtNLM-50 subset samples 14 taxonomic groups; its findings and
illustrative case studies are below.
Interactive prediction evaluation table — filterable/sortable assessments, rationales, and links to all 50 protein reviews.
Independent adjudication, both directions. Focused OpenScientist investigations evaluate both uncertain predictions and predictions disputed by the review. Their integration of sequence, structure, comparative biology, and literature provides substantial evidence for adjudication. Examples include the missing kinase domain in ARATH/F4JLB7 and experimental autophagosome localization in the human ortholog of GADMO/A0A8C5FPT8. See the OpenScientist investigation report for the individual investigations and their findings.
ARGO-ProtNLM-50 key findings
-
Useful additional annotations can follow from established biology and mapping gaps. Supported family transfers include chloroplastic EF4 localization for ARTAN/A0A2U1PS28 and a laterality role for MACFA/A0A2K5UJ34. The InterPro2GO coverage analysis identifies absent mappings, mappings on unassigned superfamily entries, and unintegrated Pfam signatures as routes by which plausible functions can be absent from GOA. A mapping gap alone does not establish the predicted function: nuclear localization for the 74-residue CAEEL/A0A061AL94 record remains uncertain. Matrix organization for 9PRIM/A0A8C9H4D2 is supported by an experimentally grounded mouse OLFML2A annotation, providing a basis for ortholog transfer beyond localization alone.
-
Exact matches often lack specificity. Most predictions classified as EXACT in the GOA comparison are LSP rather than CNN: the predicted term is already present, but a more informative, supported annotation is also available.
-
Catalytic-domain completeness matters. Shared domains or family membership can support an inference while missing catalytic regions contradict a specific activity. The wheat patatin and Arabidopsis LRR protein below illustrate why the deposited sequence needs to be examined.
-
Cross-kingdom errors occur. Neuronal cell body and neuron projection are incompatible with wheat F6LAX4, and its protein antigen binding prediction transfers an animal immune term from mammalian PP2A literature. Its other predictions require separate assessment: heterodimerization is supported by the conserved PP2A core complex, though less informatively than the existing scaffold annotations.
-
Core functions transfer more readily than regulatory context. JMJ22-related sequence and experimental evidence support epigenetic regulation for WHEAT/A0A3B6RKV1, while its four specific light, hormone, and germination predictions remain uncertain. For COLLI/A0A2I0M3K7, TRUB2 family membership does not establish the predicted tRNA substrate.
-
Biological assessment adds information beyond ontology overlap. Sigma-factor activity supports transcription initiation even when an is_a/part_of comparison reports NO_OVERLAP. Conversely, matching an existing annotation does not override target-specific contrary evidence.
ARGO-ProtNLM-50 GO results
Assessment categories follow de Crécy-Lagard et al. 2025 (PMID:40703034), with the project's GO prediction guidelines. About half of the ARGO-50 GO claims are supported, most of the rest are uncertain, and a minority are contradicted. The ARGO-50 claims view gives the category counts; the dated counts are in the cross-cohort results. The mean assessment score is an ordinal summary, not a calibrated estimate of model accuracy: the stratified sample is small, and many proteins lack direct experimental characterization.
The closure-based GOA comparison classifies each prediction's overlap with existing annotation (EXACT, MORE_SPECIFIC, LESS_SPECIFIC, NO_OVERLAP, NOT_IN_GOA). Overlap categories describe the comparison dataset, while assessment categories express biological judgments, and the two diverge in both directions: an exact match can still be less precise than another supported annotation on the same protein, and a prediction with no ontology overlap can still be a sound biological inference across GO aspects.
The contradicted ARGO-50 claims show three patterns. Some are impossible for the organism, such as neuronal localizations and protein antigen binding for the wheat PP2A scaffold F6LAX4. Most assert an intrinsic activity that the deposited sequence or domain architecture cannot support, for example kinase activity for ARATH/F4JLB7, lipase activity for WHEAT/A0A3B6GK97, and dephosphorylation by the auxilin pseudophosphatase region of DANRE/dnajc6. A few transfer a function from the wrong paralog subfamily, such as pollen maturation for the rice BURP protein ORYSI/B8BAB0. These are biological incompatibilities; frequency bias and training-data contamination are not established as their causes, and no error mechanism is inferred where the optional error_type field is unset.
Illustrative case studies
These five proteins, included in ARGO-50, illustrate annotation transfer, catalytic-domain checks, taxonomic constraints, and limitations of ontology-based evaluation. Each has a full AIGR gene review and a separate ProtNLM prediction assessment.
Catalytic-region check: A0A3B6GK97 (wheat patatin)
ProtNLM2 predicts lipase activity and lipid catabolic process for WHEAT/A0A3B6GK97. Existing IBA annotations include more specific lipase activities, making the process prediction look like a straightforward extension of known biology. However, the reproducible alignment and motif analysis show that the deposited 302-residue sequence lacks the patatin catalytic-serine region. Lipase activity is NPI; lipid catabolic process is UNC. The sequence may reflect an incomplete gene model or an inactive protein; neither a corrected full-length product nor a noncatalytic role in lipid catabolism is established. This case shows why annotation overlap alone cannot settle correctness.
Phmmer transfer: A0A3B6RKV1 (wheat JmjC)
ProtNLM2 predicts five plant biology terms for WHEAT/A0A3B6RKV1: gibberellin signaling, photomorphogenesis, seed germination, epigenetic regulation, and red-light response. The corroboration records identify Arabidopsis JMJ22 (Q67XX3; phmmer score 689.2), which has experimental annotations for all five terms. This provides a concrete basis for examining transfer from a characterized relative. Epigenetic regulation is COR, supported by the JMJ22 relationship and histone-arginine-demethylation evidence. The four specific regulatory/process predictions are UNC because their Arabidopsis experimental contexts do not establish the same roles in wheat. The phmmer hit is corroborating evidence; it does not reveal how the model generated its predictions or why PAINT omitted a term.
False positive: F4JLB7 (Arabidopsis LRR protein)
ProtNLM2 predicts kinase activity and phosphorylation for ARATH/F4JLB7. The exploratory analysis identifies a weak phmmer corroboration hit to mouse LRRK2 (score 33), a multidomain protein containing LRR and kinase regions. Sharing an LRR region does not establish kinase catalysis. The focused OpenScientist investigation integrates LRR domain assignments, catalytic-motif analysis, and predicted structure to refute a kinase domain in F4JLB7. Kinase activity is NPI; phosphorylation is UNC, since participation through another protein remains possible. The RIC7 name in the database record does not establish that this LRR protein is the CRIB-domain ROP effector described in RIC7 literature.
The RIC7 locus-identity report examines the distinction between AT4G28560/F4JLB7/EXLRR12 and the adjacent CRIB-domain RIC7 locus AT4G28556/Q1G3K8, with sequence checks, the JBrowse view, TAIR record comparisons, and an independent Cornell dissertation finding.
Cross-kingdom error: F6LAX4 (wheat PP2A scaffold)
ProtNLM2 predicts neuron projection and neuronal cell body for WHEAT/F6LAX4. Wheat has no neurons, so both localizations are NPI. Protein antigen binding is also NPI: binding of the viral small-t antigen inhibitor to human PP2A A scaffolds does not establish antigen-recognition activity, and a focused report traced the term to text transfer from mammalian PP2A literature. The remaining predictions have different evidential standing: protein heterodimerization is LSP, supported by the PP2A A-C core complex but less informative than the existing PP2A scaffold annotations; chromosome segregation and centromeric localization are UNC. This example separates clear taxonomic errors from plausible but unverified transfers.
Ontology gap: Q9KZ33 (S. coelicolor sigma factor)
STRCO/Q9KZ33 has an IBA annotation for sigma factor activity; ProtNLM2 predicts DNA-templated transcription initiation. The closure-based GOA comparison classifies the prediction as NO_OVERLAP. The molecular activity and biological process are nevertheless functionally connected: an ECF sigma factor supports promoter recognition and transcription initiation. The prediction is COR, adding a supported process annotation absent from the cached record. This illustrates a limitation of evaluating biological agreement solely through is_a/part_of paths across GO aspects.
What is ProtNLM2?
ProtNLM2 is a transformer-based sequence-to-sequence model developed by Google DeepMind with UniProt, trained on 240 million protein entries from UniProt release 2023_04. It generates protein names, GO terms, subcellular locations, keywords, and function comments from amino acid sequence. UniProt describes the current model as trained entirely on sequence. See the UniProt ProtNLM documentation.
Predictions are post-processed by the Evidencer, which applies exclusion criteria including GO taxon constraints and seeks corroboration through string matches, phmmer sequence similarity (bit score greater than 25), and TM-align structural similarity. This corroboration can explain the biological source of support for a prediction, but it is separate from the neural model's generation of that prediction. The exploratory XML dataset and public release differ in coverage; the data provenance describes the source versions used here.
Expanded cohorts and family curation
The cross-benchmark family curation covers all 282 selected protein records and paired references, with 211 structured family reviews and individual assessments for 18 inputs without exact-record PANTHER assignments. The gene-to-family index distinguishes direct sequence assignments from verified canonical gene context.
The human and model-organism challenge set selects twenty additional prediction-bearing genes across seven species, with priorities for evolutionary analysis of substrate specificity, catalytic divergence and complex participation. The cohort complements the fly and pombe selections with substrate-specificity, catalytic-divergence and complex-participation cases.
The Neurospora cohort covers every GO/function-bearing entry in the published species subset plus selected localization cases; GO predictions, function paragraphs and localization claims are reviewed separately against a full source census.
The pombe cohort includes every GO/function-bearing entry among the fission-yeast accessions identified in the original export. These currently reviewed/Swiss-Prot entries are API-accessible despite being absent from the published pilot accession list.
The fly cohort includes every Drosophila melanogaster gene with GO or function-text predictions in the published species subset, plus a separate tier of location/keyword-only records. Its census preserves exact sequences, current FlyBase identifiers and original prediction provenance.
The horse cohort contains 40 selected horse genes with released functional predictions and paired human–horse reviews. See the review findings and evidence gaps. The horse-first benchmark design includes the prediction census and mammalian evidence leads.
ARGO-ProtNLM-50 benchmark design
ARGO-ProtNLM-50 was constructed after the ProtNLM release, rather than specified in advance as a benchmark for the model. ProtNLM predictions were released for a partly arbitrary set of proteins, mostly unreviewed/TrEMBL entries. We then selected 50 proteins from that available set to sample different species and kinds of annotations. The resulting benchmark is an exploratory, stratified sample, not a prospective test set or a random sample of protein space.
The selection covers:
- 14 taxonomic groups, including mammals, plants, bacteria, fish, insects, and fungi.
- Four prediction categories: rich, partial, GO-only, and name-only.
- Multiple corroboration methods: string match, phmmer, and TM-align.
- Five case studies from exploratory analysis, described above.
All 50 proteins have AIGR gene reviews and prediction-review YAMLs; nine prediction lists are empty. The benchmark CSV records selection metadata. Each protein's *-protnlm-predictions-review.yaml preserves the prediction and source-method metadata alongside its assessment, rationale, and supporting sources.
Overlap with existing AIGR reviews
The exploratory comparison against the 1,334-review AIGR collection found eight proteins in the ProtNLM2 dataset, all unreviewed/TrEMBL entries: C5AXM3, O94267, Q09490, Q21303, Q86WA8, Q9BZE2, Q9UNW9, and Q9XUS3. This is the comparison set used in the exploratory analysis, rather than a count of the expanding AIGR collection; ARGO-50 provides a broader dedicated evaluation sample.
Evidence standards
Prediction sidecars are checked with just validate-predictions, including publication titles, source excerpts, local paths, and assessment scores. The CI artifact prediction-evidence-validation records those checks.
The function-prediction review skill defines the review criteria. Assessments integrate primary literature, sequence and domain evidence, structural analyses, experimentally grounded curated annotations, and focused OpenScientist investigations. These investigations synthesize multiple lines of evidence and carry substantial weight when their findings address the prediction. Reviews cite the relevant analyses and their limitations, distinguishing computational inference from experimental validation. A well-supported family transfer can establish a reasonable function or localization inference without a new experiment on every target; the rationale identifies the characterized relative, the target's family evidence, and the limits of transfer.
Each prediction is assessed at the specificity of its actual GO term. Extracellular localization does not establish matrix organization, and a catalytic fold does not establish a substrate. Conversely, absence of intrinsic catalytic activity does not exclude participation in the corresponding biological process through a regulatory complex. Missing evidence leads to uncertainty unless there is contrary evidence. Broad but true annotations are not biological errors.
COR and CNN distinguish absence versus presence of an equivalent annotation in the target's cached GOA/UniProt records, after biological support has been established. LSP requires an existing, supported, more specific annotation. These labels do not establish whether an example was in the model's training data. The assessment uses the available evidence, including studies published after the prediction release; it is not a time-restricted prospective benchmark.
References
- UniProt ProtNLM help page — model and Evidencer documentation.
- ProtNLM2 accession list — public release coverage.
- de Crécy-Lagard et al. 2025 (PMID:40703034) — assessment categories.
Files and methods
| Resource | Role |
|---|---|
| Benchmark CSV | Sampling metadata for all 50 proteins |
| GOA overlap comparison | Closure-based overlap comparison for the ARGO-50 predictions |
| Prediction evaluation table | Current GO assessments across all 50 ARGO-50 records |
| UniProt ProtNLM documentation | Prediction method and corroboration pipeline |
| REST API fetch pipeline | Retrieval of raw prediction and corroboration records |
| Exploratory notebook | Dataset exploration |
| Benchmark notebook | Benchmark overlap analysis |
| Slide deck (Marp source: protnlm_evaluation_slides.md) — AI generated | Exploratory presentation; assessment totals and case judgments on this page reflect the current reviews |
| OpenScientist investigation report | Focused investigations integrating multiple lines of evidence to inform prediction assessments |
| InterPro2GO coverage analysis | Domain-to-GO mapping coverage across the benchmark |
| Data history | XML/API source provenance |