ProtNLM2 Evaluation

IN_PROGRESS EVALUATIONML_PREDICTIONS

Collection: Function prediction evaluation

Species: human, HORSE, 9PRIM, ABRPR, AEDAE, AQUCT, ARAHY, ARATH, ARTAN, ASPOR, BALMU, BORPE, BOVIN, CAEEL, CALMI, CANLF, CHRVO, COLLI, COTJA, CUCME, DANRE, DEIRA, DROME, DROPS, DROVI, GADMO, GIBF5, JUGRE, MACFA, MAIZE, MYTGA, NEUCR, ORYSI, ORYSJ, PANPA, PARTE, PHATC, RABIT, SCHPO, SOYBN, STRCO, TAKRU, TOBAC, TRIV3, WHEAT, XANCP, XENLA, XENNA, XENTR, mouse, rat, worm

Genes: 9PRIM/A0A8C9H4D2 ABRPR/A0A8B8L1Z3 AEDAE/A0A6I8TLE4 AQUCT/A0A2G9RZF1 ARAHY/A0A444Z7V7 ARATH/AT4G38370 ARATH/DRS1 ARATH/F4JLB7 ARATH/FTSH12 ARTAN/A0A2U1PS28 ASPOR/Q2U1U6 BALMU/A0A8B8WEG2 BORPE/Q7VZI5 BOVIN/E1BL04 CAEEL/A0A061AL94 CALMI/A0A4W3GVU1 CANLF/A0A8I3PI07 CHRVO/Q7NUH2 COLLI/A0A2I0M3K7 COTJA/A0A8C2TBA7 CUCME/A0A1S3BTE3 DANRE/A0A8M9QG43 DANRE/dcxr DANRE/hes6 DEIRA/Q9RSY6 DROME/Ank2 DROME/CG10359 DROME/CG13096 DROME/CG13494 DROME/CG14662 DROME/CG17404 DROME/CG18507 DROME/CG18547 DROME/CG30287 DROME/CG30288 DROME/CG31099 DROME/CG31606 DROME/CG32086 DROME/CG32354 DROME/CG32706 DROME/CG33090 DROME/CG33116 DROME/CG33453 DROME/CG34117 DROME/CG34171 DROME/CG3515 DROME/CG3631 DROME/CG42331 DROME/CG42404 DROME/CG43124 DROME/CG43742 DROME/CG45100 DROME/CG46301 DROME/CG4793 DROME/CG5565 DROME/CG5611 DROME/CG6830 DROME/CG6836 DROME/CG8353 DROME/CG8745 DROME/CG8841 DROME/CG8915 DROME/CSN5 DROME/Cables1 DROME/Ctns DROME/CycA DROME/Dic4 DROME/Fbxo42 DROME/Gfat1 DROME/Gpdh3 DROME/Hmt-1 DROME/Hn DROME/IKKepsilon DROME/Lcp3 DROME/MESK2 DROME/Mrm2 DROME/Mst27D DROME/NTPase DROME/Nepl19 DROME/NtR DROME/P58IPK DROME/Pcf11 DROME/Pde4 DROME/Pld DROME/Rhp DROME/RluA-1 DROME/Serinc DROME/Synj DROME/Tango5 DROME/TyrRS DROME/alpha-Man-Ia DROME/amon DROME/awd DROME/betaTub97EF DROME/cdm DROME/dati DROME/ftz-f1 DROME/glo DROME/hbt DROME/jumu DROME/loqs DROME/metro DROME/msk DROME/orb DROME/qkr58E-1 DROME/rgn DROME/scaf DROME/sns DROME/ttv DROPS/A0A6I8W8A2 DROVI/B4MAQ2 GADMO/A0A8C5FPT8 GIBF5/S0EDH7 HORSE/AFAP1L2 HORSE/ALDH7A1 HORSE/ALG5 HORSE/BCAT2 HORSE/CACNB3 HORSE/CAPSL HORSE/CC2D2A HORSE/CDK7 HORSE/CH25H HORSE/CTDSP2 HORSE/CXCR3 HORSE/DARS2 HORSE/DNMT3A HORSE/DNMT3L HORSE/DUOX1 HORSE/DYNLT2B HORSE/EFR3A HORSE/GEMIN5 HORSE/GHSR HORSE/GPAM HORSE/HSPA4 HORSE/HSPD1 HORSE/IRAK3 HORSE/KRIT1 HORSE/MAP2K2 HORSE/MTMR9 HORSE/MYL10 HORSE/OLFML2A HORSE/OMA1 HORSE/PEA15 HORSE/PPP4R4 HORSE/PTPRN2 HORSE/SHLD2 HORSE/SIRT5 HORSE/TRAF2 HORSE/USP8 HORSE/VAPA HORSE/WDPCP HORSE/WEE1 HORSE/ZDHHC23 JUGRE/A0A2I4G8T1 MACFA/A0A2K5UJ34 MAIZE/A0A804UIX9 MYTGA/A0A8B6BFL6 MYTGA/A0A8B6GS20 NEUCR/NCU01245 NEUCR/NCU01540 NEUCR/NCU02539 NEUCR/NCU03033 NEUCR/NCU04302 NEUCR/NCU04637 NEUCR/NCU04937 NEUCR/NCU06005 NEUCR/NCU06296 NEUCR/NCU07379 NEUCR/NCU08595 NEUCR/NCU08990 NEUCR/NCU09721 NEUCR/NCU09880 NEUCR/NCU11362 NEUCR/NCU12035 NEUCR/glt-1 NEUCR/kal-1 NEUCR/mek-1 NEUCR/vtc-4 ORYSI/B8BAB0 ORYSJ/Q6YYC5 PANPA/A0A2R9CAF4 PARTE/A0BFB4 PHATC/B7FXQ8 RABIT/G1TUN6 SCHPO/SPAC25B8.09 SCHPO/asr1 SCHPO/cao1 SCHPO/cbh1 SCHPO/cem1 SCHPO/cis4 SCHPO/crt10 SCHPO/gpi16 SCHPO/lsm6 SCHPO/mre11 SCHPO/nip7 SCHPO/pta1 SCHPO/rfc3 SCHPO/rpa49 SCHPO/rpo41 SCHPO/rrp36 SCHPO/sec59 SCHPO/sen15 SCHPO/spo2 SCHPO/spt16 SCHPO/sus1 SCHPO/sws2 SCHPO/trm402 SCHPO/uch2 SCHPO/ulp2 SCHPO/vas2 SCHPO/wss1 SCHPO/yml6 SOYBN/C6T1A2 STRCO/Q9KZ33 STRCO/Q9L243 TAKRU/A0A674PKV4 TOBAC/A0A1S3Y076 TRIV3/A2FPI7 WHEAT/A0A3B6GK97 WHEAT/A0A3B6NKR6 WHEAT/A0A3B6RKV1 WHEAT/F6LAX4 XANCP/Q8P365 XENLA/uap1.S XENNA/D3VIU4 XENTR/A0A8J0SCI2 XENTR/A0A8J1IYX6 XENTR/F6WPT1 human/ACAD9 human/AFAP1L2 human/ALDH7A1 human/ALG5 human/BCAT2 human/CACNB3 human/CAPSL human/CC2D2A human/CDK7 human/CH25H human/CTDSP2 human/CXCR3 human/DARS2 human/DNMT3A human/DNMT3L human/DTD1 human/DUOX1 human/DYNLT2B human/EFR3A human/GEMIN5 human/GHSR human/GPAM human/HSPA4 human/HSPD1 human/IRAK3 human/KRIT1 human/MAP2K2 human/MTMR9 human/MYL10 human/NARF human/OLFML2A human/OMA1 human/PEA15 human/PPP4R4 human/PTPRN2 human/RHOJ human/SHLD2 human/SIRT5 human/TRAF2 human/UBE2F human/USP8 human/VAPA human/WDPCP human/WEE1 human/ZDHHC23 mouse/Sdhaf2 mouse/Spcs2 mouse/Vmn2r73 rat/Mtmr12 rat/Pnkd rat/Ptk7 worm/C28G1.2 worm/dpm-1 worm/wdr-23

Warnings (1)

ProtNLM2 Evaluation

Bottom line: ProtNLM2 is Google DeepMind's sequence-to-text model that
writes protein names, GO terms and function paragraphs straight from sequence,
and UniProt ships its output on unreviewed entries. We assessed it the way a
curator would, protein by protein, using the COR/CNN/LSP/UNC/NPI/PLI/REP
categories of de Crécy-Lagard et al. 2025 (PMID:40703034). As of the review
snapshot (2026-09-27, commit c7551cb3db), the assessment covers 242 prediction
targets and 288 GO-term assessments, plus separate reviews of the narrative
function text.
Correct novel predictions outnumber correct but already-known ones, because the
purposive cohorts favour proteins without existing annotation, and contradicted
predictions are a small minority. Narrative text behaves differently from GO
terms: paralog confusion is common in the function paragraphs and nearly absent
from the GO claims.

The recurring failure is not a wrong family but a claim pitched above what the
sequence supports: a catalytic activity where the deposited sequence lacks the
catalytic region (wheat patatin A0A3B6GK97, Arabidopsis LRR F4JLB7), or a
localization that cannot exist in the organism (neuron projection in wheat
F6LAX4). The most common single verdict is uncertain, which is the honest
result for TrEMBL proteins with no direct characterization.

Function prediction evaluation index

Browse all ProtNLM predictions — filter prediction sets, narrative reviews, and individual GO claims by species, cohort, and assessment. Fly records include reviewed empty GO outputs.

Where the numbers are. The dated counts, deduplicated across overlapping
cohorts, are in the cross-cohort results
and benchmark-summary.json. The
browser gives the same breakdowns as facet counts over the current reviews:
all GO claims, one cohort at a time
(ARGO-50,
HORSE40,
FLY41,
POMBE20,
NEUROSPORA20,
MOD_EVOLUTION20), or only the
contradicted claims.
Narrative records can hold several judgments each and are counted separately from
GO terms.

The cross-species cohorts examine biological support, substrate and paralog
specificity, domain completeness, and transfer of annotations across species. The
original ARGO-ProtNLM-50 subset samples 14 taxonomic groups; its findings and
illustrative case studies are below.

Interactive prediction evaluation table — filterable/sortable assessments, rationales, and links to all 50 protein reviews.

Independent adjudication, both directions. Focused OpenScientist investigations evaluate both uncertain predictions and predictions disputed by the review. Their integration of sequence, structure, comparative biology, and literature provides substantial evidence for adjudication. Examples include the missing kinase domain in ARATH/F4JLB7 and experimental autophagosome localization in the human ortholog of GADMO/A0A8C5FPT8. See the OpenScientist investigation report for the individual investigations and their findings.

ARGO-ProtNLM-50 key findings

  1. Useful additional annotations can follow from established biology and mapping gaps. Supported family transfers include chloroplastic EF4 localization for ARTAN/A0A2U1PS28 and a laterality role for MACFA/A0A2K5UJ34. The InterPro2GO coverage analysis identifies absent mappings, mappings on unassigned superfamily entries, and unintegrated Pfam signatures as routes by which plausible functions can be absent from GOA. A mapping gap alone does not establish the predicted function: nuclear localization for the 74-residue CAEEL/A0A061AL94 record remains uncertain. Matrix organization for 9PRIM/A0A8C9H4D2 is supported by an experimentally grounded mouse OLFML2A annotation, providing a basis for ortholog transfer beyond localization alone.

  2. Exact matches often lack specificity. Most predictions classified as EXACT in the GOA comparison are LSP rather than CNN: the predicted term is already present, but a more informative, supported annotation is also available.

  3. Catalytic-domain completeness matters. Shared domains or family membership can support an inference while missing catalytic regions contradict a specific activity. The wheat patatin and Arabidopsis LRR protein below illustrate why the deposited sequence needs to be examined.

  4. Cross-kingdom errors occur. Neuronal cell body and neuron projection are incompatible with wheat F6LAX4, and its protein antigen binding prediction transfers an animal immune term from mammalian PP2A literature. Its other predictions require separate assessment: heterodimerization is supported by the conserved PP2A core complex, though less informatively than the existing scaffold annotations.

  5. Core functions transfer more readily than regulatory context. JMJ22-related sequence and experimental evidence support epigenetic regulation for WHEAT/A0A3B6RKV1, while its four specific light, hormone, and germination predictions remain uncertain. For COLLI/A0A2I0M3K7, TRUB2 family membership does not establish the predicted tRNA substrate.

  6. Biological assessment adds information beyond ontology overlap. Sigma-factor activity supports transcription initiation even when an is_a/part_of comparison reports NO_OVERLAP. Conversely, matching an existing annotation does not override target-specific contrary evidence.

ARGO-ProtNLM-50 GO results

Assessment categories follow de Crécy-Lagard et al. 2025 (PMID:40703034), with the project's GO prediction guidelines. About half of the ARGO-50 GO claims are supported, most of the rest are uncertain, and a minority are contradicted. The ARGO-50 claims view gives the category counts; the dated counts are in the cross-cohort results. The mean assessment score is an ordinal summary, not a calibrated estimate of model accuracy: the stratified sample is small, and many proteins lack direct experimental characterization.

The closure-based GOA comparison classifies each prediction's overlap with existing annotation (EXACT, MORE_SPECIFIC, LESS_SPECIFIC, NO_OVERLAP, NOT_IN_GOA). Overlap categories describe the comparison dataset, while assessment categories express biological judgments, and the two diverge in both directions: an exact match can still be less precise than another supported annotation on the same protein, and a prediction with no ontology overlap can still be a sound biological inference across GO aspects.

The contradicted ARGO-50 claims show three patterns. Some are impossible for the organism, such as neuronal localizations and protein antigen binding for the wheat PP2A scaffold F6LAX4. Most assert an intrinsic activity that the deposited sequence or domain architecture cannot support, for example kinase activity for ARATH/F4JLB7, lipase activity for WHEAT/A0A3B6GK97, and dephosphorylation by the auxilin pseudophosphatase region of DANRE/dnajc6. A few transfer a function from the wrong paralog subfamily, such as pollen maturation for the rice BURP protein ORYSI/B8BAB0. These are biological incompatibilities; frequency bias and training-data contamination are not established as their causes, and no error mechanism is inferred where the optional error_type field is unset.

Illustrative case studies

These five proteins, included in ARGO-50, illustrate annotation transfer, catalytic-domain checks, taxonomic constraints, and limitations of ontology-based evaluation. Each has a full AIGR gene review and a separate ProtNLM prediction assessment.

Catalytic-region check: A0A3B6GK97 (wheat patatin)

ProtNLM2 predicts lipase activity and lipid catabolic process for WHEAT/A0A3B6GK97. Existing IBA annotations include more specific lipase activities, making the process prediction look like a straightforward extension of known biology. However, the reproducible alignment and motif analysis show that the deposited 302-residue sequence lacks the patatin catalytic-serine region. Lipase activity is NPI; lipid catabolic process is UNC. The sequence may reflect an incomplete gene model or an inactive protein; neither a corrected full-length product nor a noncatalytic role in lipid catabolism is established. This case shows why annotation overlap alone cannot settle correctness.

Phmmer transfer: A0A3B6RKV1 (wheat JmjC)

ProtNLM2 predicts five plant biology terms for WHEAT/A0A3B6RKV1: gibberellin signaling, photomorphogenesis, seed germination, epigenetic regulation, and red-light response. The corroboration records identify Arabidopsis JMJ22 (Q67XX3; phmmer score 689.2), which has experimental annotations for all five terms. This provides a concrete basis for examining transfer from a characterized relative. Epigenetic regulation is COR, supported by the JMJ22 relationship and histone-arginine-demethylation evidence. The four specific regulatory/process predictions are UNC because their Arabidopsis experimental contexts do not establish the same roles in wheat. The phmmer hit is corroborating evidence; it does not reveal how the model generated its predictions or why PAINT omitted a term.

False positive: F4JLB7 (Arabidopsis LRR protein)

ProtNLM2 predicts kinase activity and phosphorylation for ARATH/F4JLB7. The exploratory analysis identifies a weak phmmer corroboration hit to mouse LRRK2 (score 33), a multidomain protein containing LRR and kinase regions. Sharing an LRR region does not establish kinase catalysis. The focused OpenScientist investigation integrates LRR domain assignments, catalytic-motif analysis, and predicted structure to refute a kinase domain in F4JLB7. Kinase activity is NPI; phosphorylation is UNC, since participation through another protein remains possible. The RIC7 name in the database record does not establish that this LRR protein is the CRIB-domain ROP effector described in RIC7 literature.

The RIC7 locus-identity report examines the distinction between AT4G28560/F4JLB7/EXLRR12 and the adjacent CRIB-domain RIC7 locus AT4G28556/Q1G3K8, with sequence checks, the JBrowse view, TAIR record comparisons, and an independent Cornell dissertation finding.

Cross-kingdom error: F6LAX4 (wheat PP2A scaffold)

ProtNLM2 predicts neuron projection and neuronal cell body for WHEAT/F6LAX4. Wheat has no neurons, so both localizations are NPI. Protein antigen binding is also NPI: binding of the viral small-t antigen inhibitor to human PP2A A scaffolds does not establish antigen-recognition activity, and a focused report traced the term to text transfer from mammalian PP2A literature. The remaining predictions have different evidential standing: protein heterodimerization is LSP, supported by the PP2A A-C core complex but less informative than the existing PP2A scaffold annotations; chromosome segregation and centromeric localization are UNC. This example separates clear taxonomic errors from plausible but unverified transfers.

Ontology gap: Q9KZ33 (S. coelicolor sigma factor)

STRCO/Q9KZ33 has an IBA annotation for sigma factor activity; ProtNLM2 predicts DNA-templated transcription initiation. The closure-based GOA comparison classifies the prediction as NO_OVERLAP. The molecular activity and biological process are nevertheless functionally connected: an ECF sigma factor supports promoter recognition and transcription initiation. The prediction is COR, adding a supported process annotation absent from the cached record. This illustrates a limitation of evaluating biological agreement solely through is_a/part_of paths across GO aspects.

What is ProtNLM2?

ProtNLM2 is a transformer-based sequence-to-sequence model developed by Google DeepMind with UniProt, trained on 240 million protein entries from UniProt release 2023_04. It generates protein names, GO terms, subcellular locations, keywords, and function comments from amino acid sequence. UniProt describes the current model as trained entirely on sequence. See the UniProt ProtNLM documentation.

Predictions are post-processed by the Evidencer, which applies exclusion criteria including GO taxon constraints and seeks corroboration through string matches, phmmer sequence similarity (bit score greater than 25), and TM-align structural similarity. This corroboration can explain the biological source of support for a prediction, but it is separate from the neural model's generation of that prediction. The exploratory XML dataset and public release differ in coverage; the data provenance describes the source versions used here.

Expanded cohorts and family curation

The cross-benchmark family curation covers all 282 selected protein records and paired references, with 211 structured family reviews and individual assessments for 18 inputs without exact-record PANTHER assignments. The gene-to-family index distinguishes direct sequence assignments from verified canonical gene context.

The human and model-organism challenge set selects twenty additional prediction-bearing genes across seven species, with priorities for evolutionary analysis of substrate specificity, catalytic divergence and complex participation. The cohort complements the fly and pombe selections with substrate-specificity, catalytic-divergence and complex-participation cases.

The Neurospora cohort covers every GO/function-bearing entry in the published species subset plus selected localization cases; GO predictions, function paragraphs and localization claims are reviewed separately against a full source census.

The pombe cohort includes every GO/function-bearing entry among the fission-yeast accessions identified in the original export. These currently reviewed/Swiss-Prot entries are API-accessible despite being absent from the published pilot accession list.

The fly cohort includes every Drosophila melanogaster gene with GO or function-text predictions in the published species subset, plus a separate tier of location/keyword-only records. Its census preserves exact sequences, current FlyBase identifiers and original prediction provenance.

The horse cohort contains 40 selected horse genes with released functional predictions and paired human–horse reviews. See the review findings and evidence gaps. The horse-first benchmark design includes the prediction census and mammalian evidence leads.

ARGO-ProtNLM-50 benchmark design

ARGO-ProtNLM-50 was constructed after the ProtNLM release, rather than specified in advance as a benchmark for the model. ProtNLM predictions were released for a partly arbitrary set of proteins, mostly unreviewed/TrEMBL entries. We then selected 50 proteins from that available set to sample different species and kinds of annotations. The resulting benchmark is an exploratory, stratified sample, not a prospective test set or a random sample of protein space.

The selection covers:

All 50 proteins have AIGR gene reviews and prediction-review YAMLs; nine prediction lists are empty. The benchmark CSV records selection metadata. Each protein's *-protnlm-predictions-review.yaml preserves the prediction and source-method metadata alongside its assessment, rationale, and supporting sources.

Overlap with existing AIGR reviews

The exploratory comparison against the 1,334-review AIGR collection found eight proteins in the ProtNLM2 dataset, all unreviewed/TrEMBL entries: C5AXM3, O94267, Q09490, Q21303, Q86WA8, Q9BZE2, Q9UNW9, and Q9XUS3. This is the comparison set used in the exploratory analysis, rather than a count of the expanding AIGR collection; ARGO-50 provides a broader dedicated evaluation sample.

Evidence standards

Prediction sidecars are checked with just validate-predictions, including publication titles, source excerpts, local paths, and assessment scores. The CI artifact prediction-evidence-validation records those checks.

The function-prediction review skill defines the review criteria. Assessments integrate primary literature, sequence and domain evidence, structural analyses, experimentally grounded curated annotations, and focused OpenScientist investigations. These investigations synthesize multiple lines of evidence and carry substantial weight when their findings address the prediction. Reviews cite the relevant analyses and their limitations, distinguishing computational inference from experimental validation. A well-supported family transfer can establish a reasonable function or localization inference without a new experiment on every target; the rationale identifies the characterized relative, the target's family evidence, and the limits of transfer.

Each prediction is assessed at the specificity of its actual GO term. Extracellular localization does not establish matrix organization, and a catalytic fold does not establish a substrate. Conversely, absence of intrinsic catalytic activity does not exclude participation in the corresponding biological process through a regulatory complex. Missing evidence leads to uncertainty unless there is contrary evidence. Broad but true annotations are not biological errors.

COR and CNN distinguish absence versus presence of an equivalent annotation in the target's cached GOA/UniProt records, after biological support has been established. LSP requires an existing, supported, more specific annotation. These labels do not establish whether an example was in the model's training data. The assessment uses the available evidence, including studies published after the prediction release; it is not a time-restricted prospective benchmark.

References

Files and methods

Resource Role
Benchmark CSV Sampling metadata for all 50 proteins
GOA overlap comparison Closure-based overlap comparison for the ARGO-50 predictions
Prediction evaluation table Current GO assessments across all 50 ARGO-50 records
UniProt ProtNLM documentation Prediction method and corroboration pipeline
REST API fetch pipeline Retrieval of raw prediction and corroboration records
Exploratory notebook Dataset exploration
Benchmark notebook Benchmark overlap analysis
Slide deck (Marp source: protnlm_evaluation_slides.md) — AI generated Exploratory presentation; assessment totals and case judgments on this page reflect the current reviews
OpenScientist investigation report Focused investigations integrating multiple lines of evidence to inform prediction assessments
InterPro2GO coverage analysis Domain-to-GO mapping coverage across the benchmark
Data history XML/API source provenance