Manually reviewed all 453 DeepECTransformer EC predictions for uncharacterised E. coli proteins.
Only 3 / 453 were genuinely novel and correct.
The lasting contribution is the error taxonomy — the structure a curator needs after the aggregate score:
| COR correct novel | CNN correct but not novel |
| LSP less precise | PLI paralog-incorrect |
| NPI non-paralog-incorrect | REP frequency-biased |
| UNC uncertain |
Every rejection required synthesis — domain architecture, paralog subfamily, pathway presence, in-vitro vs in-vivo, primary literature.
Each de Crécy-Lagard verdict needed a human expert to integrate many lines of evidence.
That does not scale when methods ship monthly and prediction sets run to the thousands.
Can the synthesis step itself be partially automated?
→ AI Gene Review (AIGR): LLM curator-agents grounded in a per-gene evidence package, producing structured, traceable synthetic reviews.
AIGR is a complement to, not a replacement for, CAFA.
1 · Evidence assembly (per gene, cached & reproducible)
UniProt record · full GO annotation table (QuickGO) · InterPro architecture · cached publications (full text where available; abstract-level otherwise) · orthogonal deep-research report
2 · Curator-agent review (three phases)
ACCEPT / KEEP_AS_NON_CORE / MODIFY / REMOVE / MARK_AS_OVER_ANNOTATED / UNDECIDED + verbatim supporting quote3 · Validation — LinkML schema + best-practice checks; every quote must literally appear in a cached publication
Open source · Typer CLI · just targets · browsable at ai4curation.io/ai-gene-review
A two-stage agentic predictor (Fallahpour et al. 2026):
| Stage | What it is | Output |
|---|---|---|
| GO-GPT | Autoregressive transformer (ESM2 + organism) | GO term hierarchy (upstream input) |
| BioReason-Pro | Qwen3-4B fine-tune | <think> trace + free-text functional summary |
Provenance: the web app's GO panels are upstream GO-GPT output; the HuggingFace catalogue separately documents its structured GO section as BioReason-Pro SFT output.
COMPLETE, 48 DRAFT, 23 IN_PROGRESS, 4 INITIALIZEDcsr-1 case (n=138) and flags seven 2,000-aa truncationsGO_REF:0000002 annotations: 92/139)Correctness 4.0 / 5 · Completeness 2.9 / 5 (n=138)
| Score | Correctness | Completeness |
|---|---|---|
| 5 | 70 (51%) | 1 (1%) |
| 4 | 25 (18%) | 40 (29%) |
| 3 | 22 (16%) | 51 (37%) |
| 2 | 14 (10%) | 39 (28%) |
| 1 | 7 (5%) | 7 (5%) |
51% score 5/5 on correctness, but only one gene (Uggt1) reaches 5/5 completeness.
The failure tail is small but structurally distinctive — not random noise.
Blinded n=20 second review: correctness 80% exact / kappa 0.950; completeness 55% exact / kappa 0.744.

Mouse has the highest selected-case mean (4.7), followed by B. subtilis (4.5), rat (4.4), and human (4.3); S. pombe is lowest (3.2). The panel is deliberately uneven, so this is descriptive, not an organism-level performance estimate.
Immediately diagnostic to a reader of the narrative.
| # | Failure mode | Example |
|---|---|---|
| 1 | Pseudoenzyme blind spot | Epe1 — "JmjC demethylase" despite degenerate active site |
| 2 | Localisation defaults to cytoplasm | CpxP periplasmic → called cytoplasmic |
| 3 | Paralog indistinguishability | Fyn ≡ Src; sigF ≡ sigG ≡ sigK |
| 4 | Organism-specific biology absent | daf-16 generic FoxO, no IIS/dauer/longevity |
| 5 | Neo-functionalisation / moonlighting missed | Nmnat NAD⁺ enzyme; chaperone role lost |
| 6 | Narrative–GO disconnect | RidA: protein binding not deaminase activity |
| 7 | Cross-kingdom fold bias | aprE subtilisin → "human blood coagulation" |
| 8 | Generated UniProt-style fabrication | Slc5a1 → steroid-sulfate transporter |
The biases are architectural — they predict where the model will fail on deployment.
RAS2 (yeast, 2/5) — "a Ras-family GTPase … regulating intracellular vesicle traffic converging on the vacuole"
✗ Actually the primary activator of the cAMP/PKA pathway.
Epe1 (S. pombe, 1/5) — "a nuclear histone demethylase … JmjC oxygenase core"
✗ A pseudoenzyme (HVD not HXD); anti-silencing factor via HP1/Swi6.
TOR1 (yeast, 4/4) — "PIKK serine/threonine kinase … HEAT repeats scaffold regulatory assemblies … integrates nutrient & stress cues"
✓ Correct — the FRB + multi-domain architecture enabled pathway-level inference.
The dominant mode across 139 genes: translate InterPro domains into prose, no new biology. Where a family label or actual InterPro2GO mapping is misleading, BioReason-Pro can recapitulate and amplify it.
Adds genuine value only when multi-domain architecture is diagnostic:
TOR1 · NOTCH1 · PTEN · EGFR · spo0A · (informative family names: Uggt1, KAR2, bst1)
A method can average 4.0/5 correctness while providing little net annotation value when it mainly restates supplied domain labels.
GO-GPT run directly on 300 genes; overlap measured against three progressively stricter references:

The 3-fold gap between raw-GOA agreement (11.7%) and agent-adjudicated core-function agreement (3.9%) illustrates the difference between snapshot agreement and coverage of the local core-function reference.
955 SFT HF-catalogue terms (95 ARGO139 genes), audited against current GOA/AIGR and targeted biological review:

71.0% CNN (correct/non-novel; 631 exact GOA) · 15.9% NPI/PLI/REP · 2.5% COR · 4.6% LSP · 6.0% UNC
The 2.5% COR are known-literature gaps, not discoveries of previously unknown biology.
BioReason-Pro's narrative and its GO-term list are generated semi-independently — and can disagree:
Neither the narrative nor the term list can be trusted in isolation — a deployment protocol must evaluate both. CAFA metrics see only the term list.
(SFT-specific risk: 16% of SFT outputs fabricate fake "UniProt Summary" text for uncharacterised proteins.)
7 E. coli genes spanning all classes; AIGR reproduces the published taxonomy.
Not blinded: the project artifacts include the published expert labels/rationales.
Dataset ID: 10.5281/zenodo.20751016
| Gene | Paper | AIGR | Recovered rationale |
|---|---|---|---|
| ygfF | COR | COR | SDR family; GDH activity confirmed |
| yciO | PLI | PLI | TsaC paralog; ~10⁴× weaker activity |
| yegV | PLI | PLI | Correct sugar-kinase EC prefix; substrate unknown |
| yjhQ | NPI | NPI | Mycothiol pathway absent from E. coli |
| yrhB | NPI | NPI | QueD already encodes activity; Imm35 domain |
| yjdM | UNC | UNC | In-vitro activity, no in-vivo phenotype |
| fepE | REP | REP | No HK similarity; Wzz O-antigen regulator |
7 / 7 classifications + mechanistic rationales reproduced. This is a positive control for the schema/workflow, not a blinded accuracy estimate.
A separate literature/bioinformatics-assisted run excluded the de Crécy-Lagard paper and published rationales.
| Gene | Expert | Withheld run | Interpretation |
|---|---|---|---|
| fepE | REP | REP | Frequency-bias smell test recovered |
| yciO | PLI | PLI | Paralog-overannotation recovered |
| yjhQ / yrhB | NPI | NPI | Pathway-context failures recovered |
| yegV / ygfF | PLI / COR | UNC | Conservative misses |
| yjdM | UNC | NPI | Too harsh on in-vitro vs in-vivo boundary |
4 / 7 exact labels. Good enough to triage suspicious sequence-AI predictions; not a substitute for expert boundary judgments.
| Tier | What | Scales? | Grades narrative? |
|---|---|---|---|
| 1 · Aggregate (CAFA |
GOA temporal holdout | ✓ 10⁴ proteins | ✗ |
| 2 · Expert / agentic review (AIGR) | Per-gene synthesis + taxonomy | partially automated | ✓ |
| 3 · Prospective experiment | Assays, genetics, microscopy | ✗ no protocol | n/a |
Recommendation: report a Tier-1 score and a Tier-2 agentic biological-validity score.
BioReason-Pro mostly tells you what you already know, occasionally something correct GOA has not recorded, and assigns 15.9% of ARGO95 terms to incorrect classes in predictable, diagnosable ways.
The most valuable thing a foundation model can produce is a well-reasoned narrative — it can be reviewed, corrected, combined. Naked GO terms cannot.
Agentic Tier-2 review reads narratives, surfaces systematic failures, separates novelty from restatement — and is already useful as a triage/smell-test layer, even though expert-level nuance remains human.
Data, reviews, pipeline, schema & validator — all open:
github.com/ai4curation/ai-gene-review
Browse 139 BioReason-Pro reviews + ESR-ECOLI-DET-Mini:
ai4curation.io/ai-gene-review
de Crécy-Lagard et al. 2025 (G3, PMID:40703034) · Fallahpour et al. 2026 (bioRxiv 10.64898/2026.03.19.712954)
Talk in one line: aggregate metrics tell you the field is improving; they don't tell a database lead whether to import a given method's predictions. We propose partially-automated agentic review to answer that.