Function prediction evaluation

Testing protein-function predictors claim by claim against reviewed genes

AI Gene Review · projects/FUNCTION_PREDICTION_EVALUATION · 2026

Bottom line

  • Aggregate benchmarks say a field is improving; a curator needs to know whether one method's predictions are safe to import.
  • We review predictions one claim at a time against agent-adjudicated gene reviews, with the COR/CNN/LSP/UNC/PLI/NPI/REP taxonomy.
  • Errors concentrate in specificity, paralogs, pseudoenzymes and organism context; in the larger benchmarks most correct predictions were already known.

One review loop, many predictors

Why claim by claim

  • A CAFA-style score averages over thousands of terms and hides which terms are wrong.
  • Curators import individual annotations; a method that is 95% redundant and 5% wrong is a net cost.
  • The same seven-category taxonomy (de Crécy-Lagard et al. 2025, PMID:40703034) makes different tools comparable within a cohort.
  • Caveat: the reference reviews are AI-assisted and of mixed maturity, not expert-signed ground truth.

Results so far

Headlines by project

Project Cohort Headline
BioReason-Pro 139 genes correctness 4.0/5, completeness 2.9/5; 23 of 955 SFT terms correct and novel
ProtNLM2 242 targets 288 GO terms: 53 COR, 83 LSP, 103 UNC, 17 NPI
DeepECTransformer 7 E. coli genes one gene per error category; paper found 3/453 correct novel
Affinage 42 + 22 + 91 genes GO layer 1/42 specific; narrative added 13 annotations
TreeGrafter 898 IEA rows 41% accepted vs 72% for PAINT/IBA

Recurring failure modes

  • Less precise: a true parent instead of the specific activity (Affinage oxidoreductase activity for GPX4).
  • Pseudoenzymes: ancestral catalysis assigned to inactive members (BioReason on Epe1, pmp20).
  • Paralogs: wrong subfamily or interchangeable summaries (DeepECTF on yciO; BioReason on sigF/sigG/sigK).
  • Organism context ignored: a mycothiol synthase predicted in E. coli, which has no mycothiol pathway (yjhQ).
  • Localization default: cytoplasm when no transmembrane segment is found.

Status and next steps

  • ✅ BioReason-Pro, Affinage, TreeGrafter and the E. coli DeepECTF set have results; ProtNLM2 cohorts are in progress.
  • ⬜ Provenance-blinded re-test of the TreeGrafter corroboration effect.
  • ⬜ Ontology-aware (ancestor-distance) scoring for Affinage.

Read more: projects/FUNCTION_PREDICTION_EVALUATION.md · shared browser app/predictions/ · each project's own page