Over-Annotation Patterns Project

MATURE PIPELINEFLAGSHIPEVALUATION

Species: human, SCHPO, CANAL, PSEAE, STRCO, SACEN

Genes: PHYKPL UBA7 Epe1 LPL1 pqsC pqsB actI-ORF1 actI-ORF2 eryCII

Over-Annotation Patterns Project

Bottom line: many GO annotations are technically defensible but tell a reader almost
nothing, or assert an activity the protein does not have, because they come from
high-throughput interactome screens, domain signatures, broad keyword mappings or
over-generalised family propagation. We catalogued the recurring shapes of this
over-annotation that surfaced during gene review, eight categories in all, each tied to
worked examples in real reviews. We did this so that curators and pipeline authors can
recognise a pattern once rather than rediscovering it gene by gene; the catalogue was
first presented at the GO Consortium meeting in October 2025. Nine exemplar reviews are
complete (human PHYKPL and UBA7, fission yeast Epe1, Candida LPL1, and five
biosynthetic-cluster enzymes: pqsC, pqsB, actI-ORF1, actI-ORF2, eryCII), and their
recorded actions match the patterns: all ten generic protein binding IPI rows on
PHYKPL and UBA7 are REMOVE, and Epe1's electronically inferred JmjC-domain
demethylase, oxidoreductase and dioxygenase rows are REMOVE, while its two
experimental IDA/EXP H3K9 demethylase rows and its metal ion binding row are
UNDECIDED. The catalogue is qualitative; it does not yet measure how
often each pattern occurs across the repository.

Overview

This project documents systematic patterns of over-annotation discovered through AI-assisted gene review. Over-annotation occurs when GO terms are assigned that are technically correct but provide minimal functional insight, or when terms are too broad/generic to be useful for understanding gene function.

These patterns emerge from multiple sources:
- High-throughput screens that generate generic annotations
- Domain-based IEA annotations that don't reflect actual activity
- IBA annotations that over-generalize from distantly related proteins
- Keyword-based mappings that assign parent terms unnecessarily

Source: Presented at Gene Ontology Consortium Meeting, October 2025, Cambridge UK. See ai4curation/ai-gene-review.

Categories of Over-Annotation

1. Generic "Protein Binding" (GO:0005515)

The Problem: High-throughput interactome studies generate thousands of IPI annotations to "protein binding" that provide no functional information.

Examples from Reviews:
- PHYKPL: 4 protein binding annotations from HTP screens showing interactions with POT1, USO1, VAC14, LNX2 - none related to its metabolic function
- UBA7: 6 protein binding annotations from interactome studies - UBA7 obviously binds proteins (ISG15, UBE2L6) but the generic term adds nothing
- Epe1: Protein binding IPI row (partner Cdt2) when ubiquitin protein ligase binding (GO:0031625) is more informative

Recommended Action: REMOVE generic protein binding when more specific functional annotations exist or when interactions are from HTP screens without validation.

2. Overly Broad Enzymatic Terms

The Problem: Generic enzymatic terms (hydrolase activity, oxidoreductase activity, ligase activity) assigned when more specific terms exist.

Examples:
- LPL1: GO:0016787 (hydrolase activity) when GO:0102545 (phospholipase B activity) is more specific
- Epe1: GO:0016491 (oxidoreductase activity) assigned although no catalytic activity has been detected and the Fe(II) triad is non-canonical
- UBA7: GO:0016874 (ligase activity) when GO:0019782 (ISG15 activating enzyme activity) is specific

Recommended Action: REMOVE or MODIFY to more specific child terms.

3. Domain-Based Predictions Without Validation

The Problem: IEA annotations from domain presence (InterPro, Pfam) that don't reflect actual biochemical activity.

Examples:
- Epe1: JmjC domain → histone demethylase activity, dioxygenase activity (REMOVE - no activity detected, non-canonical Fe(II) triad, although the divergent Tyr370 is required for function); metal ion binding left UNDECIDED (two of three iron ligands retained, binding unmeasured)
- PHYKPL: Aminotransferase domain → transaminase activity (INCORRECT - functions as phospho-lyase)

Recommended Action: REMOVE when biochemical evidence contradicts domain prediction.

4. Indirect Downstream Process Annotations

The Problem: Genes annotated to broad biological processes based on indirect effects rather than direct function.

Pattern: Gene affects X → X affects Y → Gene annotated to Y

Examples:
- Kinase that phosphorylates one transcription factor annotated to "regulation of cell proliferation"
- Enzyme in metabolic pathway annotated to disease process it indirectly affects

Recommended Action: Use more proximal process terms; annotate to direct function, not downstream consequences.

5. Duplicate IEA Annotations

The Problem: Multiple automated pipelines annotate the same term, creating redundancy.

Examples:
- GND1: Same term (GO:0004616) from both IBA and IEA sources
- UBA7: Cytoplasm annotation from IBA, IEA, and IDA sources

Note: This is less problematic as multiple evidence codes can provide confidence, but creates clutter.

6. Predicted Localization Conflicts

The Problem: Automated transmembrane predictions leading to incorrect membrane annotations.

Examples:
- LPL1: GO:0016020 (membrane) from transmembrane prediction, but protein localizes to lipid droplets (monolayer, not bilayer membrane)

Recommended Action: REMOVE when experimental localization data contradicts prediction.

7. Fold-Based Pathway Mis-Propagation (wrong product class)

The Problem: A catalytic-domain signature shared between two pathways propagates the
wrong pathway's terms. The β-ketoacyl-synthase (FabH/KAS-III) fold is common to fatty-acid
and polyketide synthases, so InterPro/IEA assigns fatty-acid terms to polyketide/secondary-
metabolite enzymes.

Examples (BGC project):
- pqsC (PSEAE): GO:0006633 fatty acid biosynthetic process and GO:0004315 3-oxoacyl-ACP
synthase → the enzyme makes a quinolone QS signal (octanoate is a substrate, not the product).
- actI-ORF1 (STRCO): GO:0006633 / GO:0030497 fatty acid elongation → it is a polyketide
ketosynthase (GO:0016218); MODIFY to GO:1901112 actinorhodin biosynthetic process.

8. Catalytic Function Assigned to Non-Catalytic Complex Subunits

The Problem: Domain signatures assign a catalytic MF to every subunit of an obligate
heterodimer, including the partner that lacks the active site. Per GO guidelines the activity
belongs on the catalytic member (enables); the partner takes contributes_to at most.

Examples (BGC project): pqsB (PqsBC; catalytic Cys/His in PqsC), actI-ORF2/CLF
(no active site), eryCII (heme-less P450 activator). See PSEUDOENZYMES.md and
PROTEIN_COMPLEX_FUNCTIONS.md.

Genes Exemplifying Patterns

Gene Species Over-Annotation Pattern Status
PHYKPL human Protein binding, transaminase (wrong mechanism) COMPLETE
UBA7 human Protein binding, generic ligase COMPLETE
Epe1 pombe Domain-based demethylase (pseudo-enzyme) COMPLETE
LPL1 CANAL Generic hydrolase, membrane localization COMPLETE
pqsC PSEAE Fatty-acid synthase terms (KAS-III fold) on a quinolone synthase COMPLETE
actI-ORF1 STRCO Fatty-acid biosynthesis/elongation terms on a polyketide synthase COMPLETE
pqsB PSEAE Catalytic acyltransferase MF on the non-catalytic subunit COMPLETE
actI-ORF2 STRCO Catalytic acyltransferase MF on the chain-length factor (no active site) COMPLETE
eryCII SACEN Full P450 MF/cofactor set on a heme-less pseudoenzyme COMPLETE
  1. Specificity over breadth: Always prefer the most specific accurate term
  2. Remove uninformative annotations: Generic protein binding from HTP screens
  3. Validate domain predictions: Especially for enzymatic activity
  4. Distinguish direct vs. indirect: Annotate to proximal function
  5. Consider pseudo-enzymes: Domains don't guarantee activity

Impact on Annotation Quality

These over-annotation patterns:
- Dilute the signal from informative annotations
- Create false impressions of functional understanding
- Complicate enrichment analyses
- Propagate through IBA to other species


STATUS

Documented Patterns

Genes Analyzed

Slides

Last updated: 2026-01-22

NOTES

2026-01-22

Project Creation

Documented systematic over-annotation patterns discovered through AI review.

Key Insight: The most common over-annotation is generic "protein binding" (GO:0005515) from high-throughput interactome studies. These annotations:
- Provide no functional information
- Often represent false positives or indirect interactions
- Should be removed when specific functional annotations exist

Curation Recommendation: Consider flagging or filtering HTP-derived protein binding annotations during curation review.