TreeGrafter Inference Evaluation

MATURE EVALUATIONPIPELINE

Collection: Propagation by homology · Function prediction evaluation

TreeGrafter Inference Evaluation

Bottom line: TreeGrafter grafts a protein that is not in a PANTHER reference
tree onto the best-matching node and copies that node's GO terms to it as IEA
annotations (GO_REF:0000118), with no curator in the loop. We took every such
annotation in the review corpus at a frozen 2026-09-06 snapshot (898 annotations
on 510 reviewed proteins) and tallied how reviewers treated them, alongside the
curated PAINT/IBA set as a contrast. Reviewers accepted 41% of TreeGrafter
annotations as-is and rejected 26% (REMOVE or MARK_AS_OVER_ANNOTATED), against
72% accepted for PAINT/IBA; molecular-function terms fared worst, with 52%
down-graded. When another pipeline reproduced the same TreeGrafter call
(GO_REF:0000120) acceptance rose to 77%, though reviewers could see that label.
In five of six down-graded cases the tree placement was sound and the inherited
term was the problem (too coarse, a sibling term, or a generic localization). The
errors cluster by family: 29 of the 63 PANTHER families with at least four
reviewed annotations had half or more of their terms down-graded.

We did this because TreeGrafter output is routinely conflated with curated
PAINT/IBA, and knowing where automated grafting over-reaches gives PANTHER and
PAINT curators concrete families to fix. About 70% of the rows come from the
Pseudomonas putida KT2440 batch, so the rates are directional; the corpus has
grown since the snapshot and the tables were deliberately not chased.

Overview

TreeGrafter (Tang et al. 2019,
doi:10.1093/bioinformatics/bty625)
is the algorithm — bundled into InterProScan — that grafts a query protein
onto the most appropriate PANTHER reference phylogenetic tree and then
propagates the GO annotations attached to the grafting node down onto the query.

The critical distinction this project keeps straight is TreeGrafter vs.
PAINT/IBA
, which are routinely conflated:

PAINT / IBA TreeGrafter
Applies to genes already in the PANTHER reference tree (well-studied "core" species) sequences not in the reference tree — grafted on
Made by curators, by hand, on the tree fully automated propagation
Evidence code IBA (Inferred from Biological aspect of Ancestor) IEA (Inferred from Electronic Annotation)
Reference GO_REF:0000033 GO_REF:0000118
Assigned by GO_Central TreeGrafter
WITH/FROM PANTHER:PTN... PANTHER:PTN...

So the TreeGrafter-based inferences are the IEA / GO_REF:0000118 /
assigned-by TreeGrafter annotations
— not the IBA ones. (We verified this
directly in the corpus GOA: every GO_REF:0000118 row is IEA / TreeGrafter
/ PANTHER.)

One more reference matters, and it changes what the GO_REF:0000118 set is.
GO_REF:0000120 is UniProt's "Combined Automated Annotation using Multiple
IEA Methods"
: per its GO reference-collection definition it "integrates
identical annotations from multiple electronic pipelines, including UniRule,
ARBA, InterPro2GO, TreeGrafter2GO (GO_REF:0000118), RHEA2GO, KeyWord2GO,
SubCellular2GO, EC2GO and EnsEMBL Compara", listing the contributing pipelines
pipe-separated in WITH/FROM. In this corpus every GO_REF:0000120 row
with a PANTHER:PTN… in WITH/FROM (891 rows) also lists at least one other
pipeline (InterPro in most cases), and no gene has the same term under both
GO_REF:0000118 and GO_REF:0000002
. So the three electronic references
partition the predictions by corroboration:

Reference What the row means
GO_REF:0000118 (TreeGrafter) a TreeGrafter prediction that no other pipeline reproduced
GO_REF:0000120 with PANTHER:PTN… a TreeGrafter prediction corroborated by ≥1 other pipeline
GO_REF:0000002 (InterPro2GO) an InterPro2GO prediction that no other pipeline (TreeGrafter included) reproduced

The headline rates below are therefore for the uncorroborated TreeGrafter
output. The corroborated set and the uncorroborated InterPro2GO set are
evaluated alongside it in Corroboration.

This project evaluates how good those automated TreeGrafter inferences are when
each is held to standard GO review criteria, using the AIGR corpus of
expert/AI-adjudicated gene reviews as the gold standard.

Method

Each genes/*/*/*-ai-review.yaml records, per existing annotation, an
evidence_type, an original_reference_id, and a reviewer action from the
AIGR action enum (ACCEPT, KEEP_AS_NON_CORE, MODIFY,
MARK_AS_OVER_ANNOTATED, REMOVE, UNDECIDED, …). We treat the AIGR
adjudication as the reference judgement and ask how the annotations with
evidence_type: IEA and original_reference_id: GO_REF:0000118 (TreeGrafter)
were treated. The PAINT/IBA set (GO_REF:0000033) is reported alongside purely
as a contrast — it is a different, curator-driven pipeline.

Reproduce with:

python3 projects/TREEGRAFTER/analyze_treegrafter.py
# or, hermetically:
uv run --with pyyaml projects/TREEGRAFTER/analyze_treegrafter.py

This writes three committed sidecars (no hard-coded numbers):

Two further scripts build on it (see the failure-modes page):
analyze_placement.py joins each annotation to its PANTHER family / subfamily
/ graft node and writes
treegrafter_family_hotspots.tsv;
classify_failure_modes.py assigns every down-graded annotation a failure
mode and writes
treegrafter_failure_modes.tsv.

Results (frozen corpus snapshot, 2026-09-06)

Snapshot. Every number on this page and its sub-page, and every
committed sidecar, was generated from the review corpus as of
2026-09-06 (branch commit 49d8cc0b, on main at b62182cc) and has
been deliberately frozen there so that the figures below are the ones
that were verified line by line in review. The corpus has grown since and
keeps growing; the tables were not chased.
As a dated floor: by the
branch's merge of main at 3246edc2 (2026-09-19) the tree already held at
least 63 more GO_REF:0000118 annotations in 28 more review files than
the tables (HETGA 29 and 9AVES 4 — two species the tables have never seen,
the first a mammalian gene set — PSEPK 28, XENLA 2; none removed), and at
least four already-counted actions had changed: PSEPK fruA
GO:0090563 and fabF GO:0005829 (ACCEPT → KEEP_AS_NON_CORE), and PSEPK
pgm GO:0006166 and GO:0008973 (KEEP_AS_NON_CORE →
MARK_AS_OVER_ANNOTATED) — the pgm pair being the first post-snapshot
changes that move rows into the down-graded population the failure-mode
analysis is built on. Later merges add more. Current figures come from
re-running analyze_treegrafter.py, not from this page; regenerate the
sidecars only together with a fresh pass over failure_mode_curated.tsv
(see Refresh the snapshot under Next steps).

At the snapshot, 4,540 review files were scanned, yielding 898
reviewed TreeGrafter annotations (GO_REF:0000118) across 510 reviewed
proteins (493 distinct gene symbols) —
at that date every GO_REF:0000118 row in the corpus GOA had a review
decision (an earlier snapshot, before the P. putida KT2440 batch was
reviewed, had 415 annotations / 202 proteins).

Reviewer action TreeGrafter (IEA) PAINT/IBA (contrast)
ACCEPT 371 41.3% 7,286 71.7%
KEEP_AS_NON_CORE 192 21.4% 1,611 15.9%
MODIFY 73 8.1% 478 4.7%
REMOVE 120 13.4% 276 2.7%
MARK_AS_OVER_ANNOTATED 113 12.6% 432 4.3%
UNDECIDED 26 2.9% 45 0.4%
NEW / PENDING 3 0.3% 36 0.4%

Headline: only ~41% of TreeGrafter inferences are accepted as-is, and
~26% are outright rejected (REMOVE + MARK_AS_OVER_ANNOTATED), with a
further ~8% needing a better term (MODIFY). The accept rate was 41.0% at the
previous 415-annotation snapshot, so it has been stable while the corpus more
than doubled. This is markedly noisier than curated PAINT/IBA on the same
corpus (72% accept, ~7% rejected) — which is
exactly what you would expect: TreeGrafter is fully automated propagation onto
sequences that were not curated into the reference tree, with no curator
checking residue-level evidence at the graft point.

By GO aspect

The summary sidecar also breaks the TreeGrafter set down by aspect (taken
from the GO ASPECT column of each gene's cached GOA). Molecular-function
propagations are the worst category by a wide margin, confirming the
hypothesis that catalysis-implying MF terms are where tree grafting
over-reaches; CC terms are rarely wrong but are mostly parked as non-core.

Aspect n ACCEPT KEEP_AS_NON_CORE down-graded (REMOVE+MODIFY+OVER)
Molecular function 238 30% 13% 52%
Biological process 318 41% 17% 38%
Cellular component 342 49% 31% 18%

Corroboration: TreeGrafter-only vs multi-method vs InterPro2GO-only

On the same 510 reviewed proteins, the reviewer treatment of the three electronic
populations defined above:

Population (same proteins) n ACCEPT KEEP_AS_NON_CORE down-graded
TreeGrafter, uncorroborated (GO_REF:0000118) 898 41% 21% 34%
TreeGrafter corroborated by ≥1 other pipeline (GO_REF:0000120, PANTHER:PTN…) 637 77% 9% 12%
InterPro2GO, uncorroborated (GO_REF:0000002) 698 26% 20% 53%

Three things follow:

  1. Corroboration is the strongest single predictor of a good TreeGrafter
    call.
    The same algorithm's output is accepted 77% of the time when another
    pipeline reproduces it and 41% when none does — the corroborated set is
    accepted at a higher rate than curated PAINT/IBA on the whole corpus (72%).
    UniProt's GO_REF:0000120 merge is, in effect, already a QC gate, and the
    GO_REF:0000118 residue is the part that failed it.

Caveat — the reviewers were not blind to the label. original_reference_id
is visible in the review YAML while the annotation is being adjudicated, and
"combined multiple IEA methods" (GO_REF:0000120) reads as visibly stronger
provenance than a lone GO_REF:0000118. So the 77%-vs-41% gap may be partly
caused by the label rather than only predicted by it. The direction of the
effect is very likely real — corroboration by an independent pipeline is
genuine evidence, and this is the same reasoning the repo's IBA guidance
applies to propagation provenance — but the magnitude should not be taken at
face value from this corpus. Separating the two requires the held-out,
provenance-blinded test in the next-steps list below; until that is run, treat
77% vs 41% as an upper bound on the true effect.
2. Uncorroborated InterPro2GO is worse than uncorroborated TreeGrafter, not
better: 53% down-graded, 58% for MF. Its rejected terms are dominated by
coarse ancestors — catalytic activity (47), oxidoreductase activity (21),
membrane (17), nucleotide binding (15). This qualifies the graft-check
hypothesis on the failure-modes page: InterPro does out-resolve PANTHER for
particular proteins (AprA, S-crystallin, Mcr1), but the InterPro2GO rows that
TreeGrafter fails to corroborate are mostly low-information, and the
informative InterPro2GO calls have already been merged into GO_REF:0000120.
3. The taxon skew cuts both ways: TreeGrafter is down-graded 31% on the
P. putida rows and 42% elsewhere, while InterPro2GO is down-graded 58% on
P. putida and 37% elsewhere (the P. putida reviews were strict about
generic MF terms).

Where TreeGrafter inferences fail

Deep dive: Failure Modes & Tree Placement
joins every down-graded annotation to the PANTHER family/subfamily it was
grafted onto (and the ancestral PTN graft node), and assigns each one a
failure mode. Short answer: in five cases out of six the placement is
fine and the inherited term is the problem
— too coarse or a sibling term
from the family node (47%), or a generic / out-of-context localization or
process (36%). About one in eight (41 annotations on 27 proteins, e.g. aprA,
fcs, mdh, mqo1–3, dapE) is a true within-superfamily mis-placement,
and genuine pseudo-enzymes are rare (4 annotations, 2 proteins).

Failure mode annotations share proteins
1 Granularity — right subfamily, family/node-level or sibling term 143 47% 109
3 Generic / out-of-context CC, binding or process term 109 36% 94
4 Within-superfamily mis-placement 41 13% 27
0 Unclassified — heuristic declines to guess (curation queue) 9 3% 8
2 Pseudo-enzyme / co-opted fold 4 1% 2

Protein counts are distinct review files, not gene symbols — the corpus has
510 files but only 493 symbols (mdh, ALB, dapF and seven other symbols
span more than one file), so a symbol-keyed count under-reports modes 3 and 4
as 92 and 26.

The TreeGrafter terms most often down-graded (REMOVE / MODIFY /
MARK_AS_OVER_ANNOTATED) cluster in two failure modes:

  1. Over-specific catalytic activity propagated to the wrong paralog/subfamily.
    The grafting node carries a precise enzymatic MF that the query has diverged
    away from: NADH dehydrogenase activity (GO:0003954),
    fatty acid synthase activity (GO:0004312),
    spermidine synthase activity (GO:0004766),
    phosphotransferase activity (GO:0016776),
    carotenoid dioxygenase activity (GO:0010436),
    (S)-2-hydroxyglutarate dehydrogenase activity (GO:0047545),
    triacylglycerol lipase activity (GO:0004806). TreeGrafter places a sequence
    on a tree node but cannot tell that the catalytic residues, or the whole
    substrate specificity, have changed — the classic paralog over-annotation.

  2. Generic / uninformative localization. cytoplasm (GO:0005737),
    cytosol (GO:0005829), plasma membrane (GO:0005886), membrane
    (GO:0016020), nucleus (GO:0005634) — low-information CC terms inherited
    from distant ancestors. cytosol and cytoplasm alone are the two most
    frequently down-graded TreeGrafter terms in the corpus. The MF analogue is
    identical protein binding (GO:0042802), an uninformative
    "it oligomerises" term propagated from the tree.

There are also process-level mis-propagations of two kinds: a mechanistically
wrong sibling process (lipopolysaccharide core region biosynthetic process
on the mcr phosphoethanolamine transferases, which modify lipid A, not the
core oligosaccharide), and pathway terms propagated onto hosts that lack the
pathway (sucrose biosynthetic process on P. putida fbp).

These are precisely the cases where automated tree grafting lacks the
gene-specific evidence (catalytic-residue conservation, substrate assays,
organism pathway context) that a curator brings — and why PAINT/IBA, where a
curator made the call, fares so much better.

Family hotspots (upstream targets)

treegrafter_family_hotspots.tsv
aggregates the reviewer outcome over all 898 TreeGrafter annotations per
PANTHER family, subfamily and graft node. Of the 63 families with at least
four reviewed annotations, 29 have half or more of their propagated terms
down-graded
and 19 have none — the errors are concentrated, not diffuse.
The worst families are candidate PAINT subfamily-annotation or node-term
fixes:

Family n proteins down-graded Failure Example genes
PTHR10543 beta-carotene dioxygenase 8 4 100% stilbene/lignostilbene dioxygenases grafted onto the carotenoid-cleavage subfamily (mode 4) Q53353, Saro_0802, Saro_2809, lsdB
PTHR30443 "inner membrane protein" (EptA) 8 4 100% EptA node carries LPS core and phosphotransferase instead of pEtN transferase (mode 1) mcr-1, mcr2, mcr-3, mcr-4
PTHR44169 acyl-DHAP reductase 6 1 100% lipid-enzyme family terms on a secondary-metabolite SDR (mode 1) fogD
PTHR48078 threonine dehydratase 6 2 100% serine-deaminase sibling terms on biosynthetic IlvA (mode 1) ilvA-I, ilvA-II
PTHR11556 FBPase-related 11 2 64% plant cytosolic-isoform processes on chloroplast/bacterial FBPases (mode 4/3) NCGR_LOCUS1270, fbp
PTHR11558 spermidine synthase 9 3 67% family-level spermidine terms on the PMT subfamily (mode 1) NaPMT3, PMT1, PMT2
PTHR21047 dTDP-sugar epimerase 10 3 60% generic polysaccharide / epimerase parents (mode 1) eryBVII, rfbC, rmlC
PTHR11632 SDH flavoprotein 7 2 71% SDH/FRD/APS-reductase heterogeneity (modes 1 and 4) aprA, sdhA
PTHR43775 fatty acid synthase 4 4 100% family-level FAS term on PKS subfamilies (mode 1) eryAI–III, Pks1
PTHR43128 L-2-hydroxycarboxylate DH 4 2 100% MDH grafted onto the L-LDH subfamily (mode 4) METEA/mdh, PSEPK/mdh
PTHR21272 catabolic 3-dehydroquinase 4 4 100% catabolic process on biosynthetic type-II DHQases (mode 4) aroQ, aroQ1, aroQ2, aroQ-III
PTHR43808 acetylornithine deacetylase (M20A) 5 3 80% DapE / PepV grafted onto the ArgE branch (mode 4) dapE, pepV

Caveats

Deeper analyses


NOTES

2026-09-27

2026-09-25

2026-09-19

2026-09-07

2026-09-06

Next steps

Slides