Function Prediction Evaluation

IN_PROGRESS EVALUATIONPIPELINEFLAGSHIP

Function Prediction Evaluation

Bottom line: new protein-function predictors appear faster than curators can
judge them, and aggregate benchmarks do not say whether a given method's
predictions are safe to import. This page indexes the AI Gene Review projects
that test such predictions claim by claim against agent-adjudicated gene
reviews, scoring GO terms with the COR/CNN/LSP/UNC/PLI/NPI/REP taxonomy from de
Crécy-Lagard et al. 2025 (PMID:40703034). Across projects the errors concentrate
in specificity, paralogs, pseudoenzymes and organism context. In the largest
model benchmark most correct predictions were already known: as of the review
snapshot (2026-09-27, commit c7551cb3db), 682 of 955 BioReason-Pro SFT terms
were correct but not novel and 23 were correct and novel. ProtNLM2's purposive
cohorts invert that balance because they were selected for likely-novel targets.
Affinage's GO layer rarely reached the specific curated function, and reviewers
accepted uncorroborated TreeGrafter inferences much less often than curated
PAINT/IBA ones. Each project below has its own cohorts, methods and denominators.

Browse all predictions — a shared faceted catalog of prediction sets and GO/EC claims, including narrative reviews and assessed empty outputs. Filter by method, species, project, cohort, or assessment; share the resulting URL. Browser guide.

The project pages keep their prose to findings and methods; per-category counts
are facet counts in the browser. Useful starting views:
ARGO95 SFT claims,
ProtNLM2 GO claims,
DeepECTF claims, and the
GO-GPT three-level overlap, which is fixed at the review snapshot.

Model and agent evaluations

Project What is evaluated Explore
ProtNLM2 GO predictions across a taxonomically diverse protein benchmark, assessed for biological support, specificity, and overlap with existing annotations. Prediction reviews
BioReason-Pro and GO-GPT BioReason-Pro functional summaries and reasoning traces, its SFT GO predictions, and the separate upstream GO-GPT term predictions. SFT reviews · GO-GPT reviews · Manuscript
DeepECTransformer / E. coli Enzyme-function predictions for selected E. coli proteins, with attention to substrate specificity, paralogs, and physiological context. Prediction reviews · Recapitulation experiment
Affinage Literature-derived functional narratives, GO grounding, and retrieval of relevant publications. Pilot results · Narrative versus GO analysis · Project findings

BioReason-Pro SFT, RL narratives, and upstream GO-GPT outputs are separate
evaluation targets. The GO-GPT review includes unresolved predictions; its table
is a review workspace as well as a results browser. The DeepECTransformer table
is hosted with the BioReason comparison material and is also accessible through
the E. coli project.

Annotation-transfer and rule reviews

These projects examine the methods and mappings behind existing annotations and
provide context for evaluating additional model predictions. Transfers by
orthology, phylogeny, and family membership have their own index,
Propagation by Homology, with a
browser of all propagated annotations.

Project Focus
TreeGrafter Automated placement onto PANTHER trees and transfer of ancestral GO annotations.
InterPro2GO GO mappings attached to domain and family signatures, including specificity and propagation limits.
NCBIFam Functional-family mappings and opportunities or risks in extending GO coverage.
PAINT / IBA Curator-assessed phylogenetic function inheritance and the evidence for individual transfers.
ARBA rule reviews Reviews of UniProt's automated annotation rules and their biological scope.

Reading the evaluations

Biological correctness, annotation specificity, novelty relative to existing
annotations, and reference quality are distinct questions. The projects use
different cohorts and review procedures, so their scores should be read with
their own methods and denominators. Local AI-assisted reviews also vary in
maturity; they are not automatically independent experimental ground truth.

For the shared approach to term-level review, see the
evidence standards. For narrative
correctness and completeness, see the
BioReason evaluation rubric.

Slides

All projects in this collection

Generated from project frontmatter (collections: [FUNCTION_PREDICTION]).

ProjectMaturityTags
Affinage Evaluation ProjectMATUREPIPELINE, EVALUATION
BioReason-Pro Comparison ProjectMATUREPIPELINE, FLAGSHIP, EVALUATION
IBA Annotation Quality ProjectMATUREPIPELINE, FLAGSHIP
InterPro Mapping Review ProjectIN_PROGRESSPIPELINE
NCBIFAM / CDD → GO Contribution & Gap ProjectIN_PROGRESSPIPELINE
ProtNLM2 EvaluationIN_PROGRESSEVALUATION, ML_PREDICTIONS
TreeGrafter Inference EvaluationMATUREEVALUATION, PIPELINE
Validating E. coli ML PredictionsCOMPLETEPIPELINE, FLAGSHIP, EVALUATION