NCBIFAM / CDD → GO Contribution & Gap Project

SCOPING PIPELINE

NCBIFAM / CDD → GO Contribution & Gap Project

The value question, answered first: yes — and here are the mappings

Are there ncbifam2go mappings (existing NCBI-assigned, or new/refined here) that
would add high-value annotations we do not already get from the InterPro rollup?
Yes.
"Not from the rollup" has a precise meaning here, because NCBIFAM reaches GOA
only through interpro2go: an added annotation is genuinely net-new when an
entry already carries the NCBIFAM signature but the term never propagated — the
family is unintegrated into InterPro, or its InterPro entry has no interpro2go row.
That net-new count is exactly the closure-aware gain measured live by
ncbifam_go_gain.py (an entry counts as gaining the
term only if it lacks the term and every descendant). The highest-value cases,
straight from the validated seed:

NCBIFAM family Proposed GO (→ ncbifam2go) Net-new (all) reviewed† Kind Why it's unique / high-value
NF033545 IS630 transposase GO:0004803 transposase activity 18,874 2 ~ adopt NCBI Mobile elements — InterPro integration lags most here; single biggest gap
TIGR03542 LL-DAP aminotransferase GO:0010285 LL-DAP transaminase 1,185 2→0 refine Large bacterial TrEMBL gap is real; but both reviewed "gaps" are plant ALD1 paralogs (see †)
NF041162 encapsulin shell protein GO:0140737 encapsulin nanocompartment 897 0 adopt NCBI CC term; new microbial-compartment biology, essentially no propagation
TIGR00417 spermidine synthase GO:0004766 spermidine synthase 575 1→0 refine NCBI gave near-root GO:0003824 (gain ~0); specific term reveals 575; the 1 reviewed "gap" is tobacco PMT paralog (see †)
NF006559 dihydroorotase GO:0004151 dihydroorotase activity 491 0 refine NCBI gave broad GO:0016810; specific child (EC 3.5.2.3) unmasks 491
NF002326 dGTPase GO:0016793 (family) + GO:0008832 / GO:0106375 (clades) (split) split ⚠️ Corrected: structural verification showed this family is substrate-heterogeneous, so it was split — family-level exactMatch to the parent GO:0016793 (gain ~0) + two narrowMatch sub-clades (strict dGTPase GO:0008832 marked do-not-blanket-propagate; broad SAMHD1-like GO:0106375). The apparent 456 net-new for the strict term applies only to the dGTPase core, not the whole family (see ‡‡ / the SSSOM split block)
TIGR02791 VirB5 (T4SS) GO:0043684 type IV secretion system complex 298 4 ✅ adopt (broad) 0/298 carry it; the one clean reviewed gap-fill — all 4 are Brucella VirB5
NF042963 anti-phage DUF1156 GO:0051607 defense response to virus 151 0 adopt NCBI Anti-phage family; BP term is coarse (see ‡) but mechanism now resolved to a DNA amino-MTase → gains a real MF (see ‡‡)

† Verification of the reviewed (Swiss-Prot) gains (2026-07-18). The reviewed
column is the curation-relevant one, so each nonzero value was checked by pulling the
actual Swiss-Prot entries that lack the term (ncbifam_go_gain.py counts reproduced
live; entries inspected via the UniProtKB REST API). The check materially changes
the reviewed story — only VirB5 survives as a clean gap-fill:

Row reviewed gain (all UniProtKB) …restricted to prokaryotes What the entries actually are
VirB5 → GO:0043684 4 4 ✅ Genuine — all 4 are Brucella VirB5, real T4SS subunits
dGTPase → GO:0008832 13 13 ⚠️ All 13 are "dGTPase-like protein", no EC assigned — curators deliberately withheld the specific activity; propagating GO:0008832 asserts a substrate they declined (the FtsX over-annotation pattern, within Bacteria)
IS630 → GO:0004803 2 2 ~ Both are "uncharacterized" IS630-element ORFs (Shigella, Sinorhizobium) — plausibly the transposase, but left uncharacterized; weak
LL-DAPGO:0010285 2 0 ✗ Both are plant ALD1 (Arabidopsis Q9ZQI7, rice Q6VMN7); Q9ZQI7 carries L-lysine α-aminotransferase (GO:0062045, pipecolate/defense), not DapL — cross-kingdom paralog trap
spermidine synthase → GO:0004766 1 0 ✗ Tobacco PMT1 (Q42963), neofunctionalized to putrescine N-methyltransferase (GO:0030750) — paralog trap

Two lessons fall out. (1) Scope the gain to the model's lineage. NCBI's PGAP
NF/TIGR models are only applied to prokaryotic genomes, yet UniProt runs the HMMs
cross-kingdom; querying all of UniProtKB inflates the reviewed gain with eukaryotic
paralogs PGAP would never touch. Restricting to Bacteria+Archaea collapses the LL-DAP
and spermidine reviewed gaps to 0 — the bacterial reviewed enzymes already carry
the term; only diverged plant paralogs were missing it. (2) "-like" is a curator
caution flag.
Even in-scope, the dGTPase "gap" is a clade the curators declined to
give the specific activity. So the honest reviewed tally among these eight is: 1
clean (VirB5), 1 over-annotation-risk (dGTPase), 1 weak (IS630), 2 spurious (LL-DAP,
spermidine).
The all-UniProtKB (bacterial TrEMBL) gains remain real — this only
corrects the small reviewed number, which was the fragile part of the claim.

‡ "defense response to virus" is coarse — and no better GO term exists yet. The
user's instinct is right: NF042963 is phage-specific, and GO:0051607 (defense
response to virus) does not say so. But GO has no defense response to bacteriophage term — the only defense response to X children are virus, bacterium,
symbiont, protozoan, insect, nematode, oomycetes, fungus, parasitic plant (QuickGO,
2026-07-18). So GO:0051607 is the least-wrong existing BP term (a phage is a virus),
and the right BP deliverable is a proposed_new_terms entry for a phage-specific
child (defense response to bacteriophage, is_a GO:0051607), not a confident
ncbifam2go exactMatch. This is a BP row whose value is entirely in bacterial TrEMBL
(0 reviewed), so it does not carry the "high-value reviewed" claim regardless.

‡‡ The mechanism is no longer unknown — and it hands the family a real MF. We
ran the DUF1156 family through OpenScientist as a blinded, structure-scoped
hypothesis job (representative A0A3B7MFS0, Thermosynechococcus, 1002 aa; prompt gave
only the Gao-2020 phage phenotype + sequence). Its structural analysis and our own
holdout agree: NF042963 is, at its core, an intact site-specific SAM-dependent
DNA amino-methyltransferase
(Class I motif I FxGxG + catalytic DPPY motif IV,
converging in one Rossmann SAM pocket in the AlphaFold model), with no nuclease
anywhere in the chain; the DUF1156 module and C-terminal extension are accessory
(target-recognition / scaffolding). The defense mechanism is therefore modification-
based self/non-self discrimination in the BREX/DISARM/restriction–modification mould

— host DNA is self-marked by methylation, unmethylated phage DNA is excluded, and any
effector step is supplied in trans. This upgrades the row from "DUF of unknown
function" to a family with a concrete molecular function: GO:0009008 DNA-methyltransferase activity (safe parent; amino-MTase confirmed, m5C excluded).
The specific target base is UNDECIDED — OpenScientist confidently called
N4-cytosine (GO:0015667, EC 2.1.1.113) from Foldseek, but the fold cannot
distinguish amino-MTase subtypes, and the family-wide annotation favors N6-adenine
(GO:0009007, EC 2.1.1.72; 19/151 members named adenine-specific, 0 cytosine),
matching our holdout — so we hold the base open pending LC-MS/MS. This MF goes
beyond NCBI's BP-only go_terms for the family, so it is a proposed enrichment
(a proposed_new_terms MF), not part of the adopt/refine seed. Full write-up and
blinded comparison:
NF042963-hypotheses/duf1156_antiphage_mechanism/.
It also illustrates an OpenScientist failure mode worth remembering: a confident
single-representative Foldseek call (m4C) over-specified past what the fold can
resolve
— caught only by the family-wide sequence/annotation check.

Structural verification of the seed rows (OpenScientist, 2026-07-18). Four
high-value rows were sent to OpenScientist as blinded, structure-scoped hypothesis
jobs
(Foldseek fold assignment + active-site check on a representative; the mapping
was withheld from the prompt), each with a pre-registered holdout prediction.
Full reports and blinded comparisons live under
NF0*-hypotheses/.
The round was productive in both altitude directions:

Row Representative OpenScientist verdict Effect on the seed
NF042963 DUF1156 (anti-phage) A0A3B7MFS0 Intact SAM-dependent DNA amino-methyltransferase (DPPY), no nuclease; BREX/DISARM-like defense; DUF1156 accessory Adds an MF the DUF name hid (GO:0009008; base UNDECIDED m6A/m4C) — seed under-specified
NF002326 dGTPase Q92Q32 ("-like") Active site intact, but SAMHD1-like broad dNTPase, strict dGTP not supported Over-refinement caught: keep GO:0008832 for strict members, map "-like" clade to sibling GO:0106375
NF033545 IS630 transposase P16943 ("uncharacterized") Bona fide DDE transposase (RNase-H fold, intact D181/D261/E297 triad) Confirms GO:0004803; grounds the 18,874 headline (add an intact-triad filter for defective IS copies)
NF041162 encapsulin A0A0D5NHT9 ("Membrane protein") HK97 encapsulin shell, not a membrane protein (⚠️ weakly blinded — the prompt named the accession, so Foldseek could read encap_f2a/IPR049822 off the curated signature; the independent part is the AlphaFold no-TM evidence) Confirms CC GO:0140737; "membrane protein" is a mis-annotation

The two enzyme rows are the payoff: DUF1156 shows the seed can under-specify (a
"DUF" concealing a real MF) and dGTPase shows it can over-specify (a strict child
propagated onto a broad-specificity clade) — the same altitude discipline as the FtsX
case, now demonstrated in both directions and grounded in structure rather than
annotation. Two operational notes for scaling: a confident single-representative
Foldseek call can over-specify
(DUF1156 m4C-vs-m6A; resolved only by a family-wide
check), and OpenScientist jobs occasionally need a narrowed, single-analysis prompt
to finish under the 7200 s API ceiling (IS630).

The honest caveat that frames all of the above: the bulk of the gain is in
unreviewed TrEMBL, not Swiss-Prot — which is exactly what we should expect and is
fine. Over a 60-model equivalog sample the totals are Σ 19 reviewed vs 26,578
all-UniProtKB
— the opposite emphasis from RHEA, because NCBIFAM is prokaryote-heavy
and Swiss-Prot's curated prokaryotic enzymes usually already carry the term. The value
is real and large, but it is mostly automated-annotation depth; the curation-ready
(reviewed) additions are a small, high-quality minority — and, as † shows, must be
inspected per-entry and taxonomically scoped before being trusted.

Two things this table also demonstrates about how to build ncbifam2go:

Scale beyond the hand-picked eight. The same EC-bridge standard, run over the
whole collection (ncbifam2go_candidates.py),
yields 2,455 models (2,503 rows) where NCBI's go_terms and GO's ec2go
independently agree — an automatable, ready-to-add ncbifam2go core (AMR-rich:
β-lactamases, trimethoprim-resistant DHFRs, aminoglycoside acetyltransferases), plus
843 "refine" models where ec2go supplies a specific term over NCBI's broad one.

A second, different sense of "unique." Separately, NCBIFAM is the sole
integrated signature behind 250 of this repo's existing interpro2go
annotations — including mainstream genes GAPDH, RPS3, ATP6V1A, HMGCS2. Those do
come through the rollup today, but they exist only because an NCBIFAM/TIGRFAM
model is the integrating signature; remove NCBIFAM and those GOA rows disappear. That
is an attribution/robustness argument (detailed below),
distinct from the net-new gap-fills in the table above.

Overview

This project audits the NCBI protein-family resources — NCBIFAM (the
PGAP/TIGRFAM HMM collection) and CDD (Conserved Domain Database) — as sources of
GO annotations
, in the same spirit as the RHEA, SPKW, and
UniPathway source-audit projects, and develops the missing
ncbifam2go / cdd2go mappings the way RHEA develops rhea2go gap-fills.

The structural fact that organises everything: NCBIFAM and CDD are InterPro
member databases, so they reach GOA through exactly one pipeline
— InterPro
integration → interpro2goGO_REF:0000002 (evidence IEA,
assigned_by=InterPro). There is no public ncbifam2go or cdd2go
external2go file (verified: both 403 on current.geneontology.org; only
interpro2go is served). So like RHEA, the analysis runs in two directions:

  1. Contribution (forward). Where is a GO_REF:0000002 annotation actually
    carried by an NCBIFAM/CDD signature (versus Pfam, PROSITE, etc.), and is that
    contribution useful or over-/mis-propagated? — the SPKW-style audit, here
    complicated by InterPro masking the member DB.
  2. Gaps (reverse). Where does NCBIFAM/CDD assert a function — an NCBI-curated
    GO/EC term, or a precise equivalog family name — that never reaches GO,
    either because the signature is unintegrated into InterPro or because the
    integrated InterPro entry has no interpro2go row? These are gap-filling
    opportunities, not over-annotations.

See NCBIFAM-METHODOLOGY.md for the queries, the
reproducible probe (ncbifam_cdd_probe.py), and
the masking/closure caveats.

Key Findings (scoping pass)

NCBIFAM/CDD is mostly masked by InterPro

UniProt enzymes and families almost always carry multiple member signatures
(an NCBIFAM equivalog and a Pfam domain and a CDD model …), all collapsed into
one or a few InterPro entries. So an NCBIFAM/CDD signature's "real"
GO_REF:0000002 contribution is only the slice where it is the integrating /
distinguishing signature for the InterPro entry, and where no other member DB
already supplies the same term. Two consequences:

Un-masking: member-DB attribution on this repo's annotations

The masking claim is now measured, not just asserted. Every GO_REF:0000002
annotation in this repo's genes/**/*-goa.tsv was re-joined to its InterPro entry's
member_databases via the InterPro API
(interpro_member_attribution.py,
resumable). Across 5,549 resolved InterPro-citation rows (1,827 distinct entries):

Member DB distinct entries annotation rows sole signature
pfam 708 (39%) 2,374 (43%)
panther 459 (25%) 1,082 (19%)
ncbifam 272 (15%) 705 (13%) 250
prints / profile / smart ~190 each ~820 each
cdd 183 (10%) 469 (8%) 116
hamap / pirsf / ssf / cathgene3d 117–159 403–627

NCBIFAM contributes a signature to 705 (13%) and CDD to 469 (8%) of the repo's
InterPro2GO annotations — a contribution entirely invisible in GOA, which shows
only the InterPro:IPR… id. Stronger still, NCBIFAM is the sole integrated
signature for 250 rows and CDD for 116
— annotations that exist purely because of
an NCBIFAM/CDD model, with no other member DB in the entry. And these are not exotic:
NCBIFAM-sole entries back annotations on GAPDH (IPR006424GO:0006006), RPS3
(IPR005703GO:0006412), ATP6V1A (IPR005725GO:0016887), and HMGCS2
(IPR010122GO:0008299) — mainstream genes whose InterPro2GO term traces solely to a
TIGRFAM/NCBIFAM signature. This is the quantitative form of the masking finding, and
the join the sibling InterPro Mapping Review project needs to attribute
each interpro2go term to the member DB that earned it.

Specificity and quality cut by family_type

NCBIFAM's family_type is a built-in altitude/quality signal absent from CDD:

family_type N Function-transfer safety
equivalog 13,253 High — all members one function; GO/EC transfer safe
domain 10,974 Low — domain, not whole-protein function
subfamily 4,564 Medium — clade-specific; check altitude
PfamEq / PfamAutoEq 1,807 / 1,204 Equivalent to a Pfam entry → likely already InterPro-covered
exception / hypoth_equivalog 1,341 / 434 Curated caveat / hypothetical — review individually

CDD models, by contrast, are domain/architecture-oriented and lack this typing,
so CDD's forward contribution skews toward broad domain terms (the
protein binding / generic-domain altitude problem) and its reverse gap is
harder to curate than NCBIFAM's.

Does CDD have its own GO mappings? No — for CDD-proper

It is worth pinning down, because the CDD search database superficially looks
like it carries GO. The answer separates two things CDD conflates:

Conclusion: the GO-bearing NCBI family resource is NCBIFAM, not CDD. CDD
"has GO" only as a search aggregator surfacing NCBIFAM/Pfam GO under cd/PRK/
TIGR accessions. This is why the curated mapping deliverable targets
ncbifam2go, and a cdd2go would be largely redundant with it (and with
Pfam→InterPro). Reproduce with ncbifam_cdd_probe.py and the FTP/Entrez checks in
NCBIFAM-METHODOLOGY.md.

Gaps Found (scoping)

# Gap Size What it is
G1 InterPro masking all GO_REF:0000002 rows GOA hides which member DB fired → member contribution unattributable without re-join
G2 NCBIFAM unintegrated 11,064 / 18,511 (60%) NCBIFAM signatures not in InterPro → contribute no GO
G3 CDD unintegrated 14,843 / 19,902 (75%) CDD signatures not in InterPro → contribute no GO
G4 NCBI-curated GO not ingested 11,228 models w/ GO NCBIFAM models carry NCBI GO that GO has no ncbifam2go to ingest
G5 NCBI-curated EC not ingested 6,417 models w/ EC Could seed EC2GO-bridged GO mappings (the RHEA EC-bridge pattern)
G6 Integrated-but-unmapped InterPro entries staged NCBIFAM/CDD integrated into an IPR entry that has no interpro2go row

G4/G5 are the high-value half: an equivalog with a clean NCBI go_terms /
ec_numbers value is a ready-to-curate mapping; an ec_numbers-only equivalog
can be bridged through ec2go exactly as RHEA bridges reactions through
rhea2ec/ec2go.

Curated new mappings (SSSOM)

The curation deliverable mirrors RHEA's rhea2go.sssom.yaml:
ncbifam2go.sssom.yaml records the NCBIFAM-family →
GO mapping we propose for ingestionnot a transcription of NCBI's
hmm_PGAP.tsv go_terms. Each is backed by the model's family_type,
product_name, EC, and PMIDs, plus the live UniProtKB propagation gain. A
250-row seed (248 families; the dGTPase family is 3 rows after the clade split
below, +220 EC-bridge rows promoted from the candidate set — batch 2 (10 hand-picked) plus
gain-ranked batches 3–5 covering the complete gain ≥ 100 set (60 at ≥500, 48 at 200–499,
102 at 100–199)) spanning all three GO aspects (MF, BP, CC — NCBIFAM is a whole-protein family
resource, not enzyme-only) is in place, with predicate classes parallel to RHEA:

We suggest our own term where NCBI's was too broad — and that unmasks the real
gain.
For five families NCBI's go_terms gave only a broad parent — twice the
ontology near-root GO:0003824 catalytic activity (enoyl-CoA hydratase NF005804,
spermidine synthase TIGR00417) — even though a precise, EC-bridged child already
exists. Rather than record the useless broad term, the seed proposes the specific
child as an exactMatch (dGTPase→GO:0008832, enoyl-CoA hydratase→GO:0004300,
dihydroorotase→GO:0004151, spermidine synthase→GO:0004766, LL-DAP
aminotransferase→GO:0010285). This is not cosmetic: the broad parent is
near-universal so its propagation gain looks ~0, but the specific term reveals
large gaps the parent masked
— spermidine synthase 575, LL-DAP aminotransferase
1,185, dihydroorotase 491, and dGTPase 456 (incl. 13 reviewed/Swiss-Prot)
entries missing the precise activity. Proposing our own term is what turns these from
invisible into actionable gap-fills.

…but more specific is not always right — the FtsX cell-division case. The
mirror-image judgement is TIGR00439 (permease-like cell division protein FtsX),
where chasing specificity would be over-annotation. NCBI assigned GO:0000910 cytokinesis; the ontology shows cytokinesis is part_of GO:0051301 cell division
(so NCBI's term is actually the narrower one — an earlier draft of this seed had
that backwards). Empirically, all 7 reviewed FtsX proteins carry GO:0051301 cell division but only 2/7 carry cytokinesis or the most specific GO:0043093 FtsZ-dependent cytokinesis. FtsX/FtsEX regulates septal peptidoglycan hydrolysis
and divisome assembly, so curators annotate the safe participation term, not the
constriction act. The gain numbers make the trap explicit: mapping to GO:0051301
gains only 22 (it is already near-universal — a confirmatory mapping), whereas
mapping to GO:0043093 would show a 3,304 apparent gap — but propagating
FtsZ-dependent cytokinesis to every FtsX would assert more than the family supports.
We therefore propose the curator-consensus GO:0051301 cell division as the
exactMatch and decline the higher-gain specific term. Specificity is the goal
only up to the altitude the evidence supports.

Verification matters and catches real errors. One scoping-sample NCBIFAM
go_terms value (GO:0009448 on a GABA transaminase) is obsolete; a
diacylglycerol-kinase model (NF009874, EC 2.7.1.107) is tagged with the wrong
activity GO:0003951 NAD+ kinase activity (the correct GO:0004143 exists); and
several assignments sit at near-root altitude. The obsolete and wrong ids were
excluded; the broad ones were replaced by our own specific term (above) — so
the seed drops obsolete/incorrect ids, prefers specific children, and
EC-bridge-confirms enzymes. Every GO id/label was checked non-obsolete
against QuickGO (2026-06-20); every family id/name/type/EC is from hmm_PGAP.tsv;
every EC→GO bridge against the live ec2go. Validate
with just validate-ncbifam-mappings — SSSOM structural validation plus GO
term/label validation (object bound to the full GO graph, MF+BP+CC; generated
nested view ncbifam2go.terms.yaml). The seed
passes validation.

A cdd2go set is not planned: per the CDD section,
CDD-proper has no native GO, and the GO surfaced through CDD belongs to NCBIFAM
(captured here) or Pfam (InterPro-routed) — a cdd2go would be largely redundant.

Scaling the seed to the whole collection (EC-bridge candidates)

The 250-row seed has two tiers of review, and they should not be conflated: ~40
fully hand-curated rows
(the original bespoke set + the hand-picked batch 2) with
individually-written rationale, plus 210 EC-bridge rows bulk-promoted from the
candidate set (batches 3–5). The promoted rows are not individually hand-reasoned —
their comments are uniform/generated and each was auto-verified against a fixed bar
(live ec2go bridge, QuickGO non-obsolete + canonical label, live UniProtKB gain ≥ 100,
non-catalytic-subunit screen). The EC bridge is what lets us apply that same
mechanical evidence standard
to the whole collection with no per-row human judgement,
because the agreement of two independent curated resources (NCBI's go_terms and GO's
ec2go) is the verification — but that standard is weaker than the altitude/over-
annotation judgement applied to the ~40 hand-curated rows (the "high gain is a flag,
not an instruction" discipline holds for the hand-set; the promoted rows rely on the
EC-bridge agreement plus the subunit screen). ncbifam2go_candidates.py
walks every NCBIFAM model and emits each (model, GO) where ec2go(model's EC)
confirms one of the model's own NCBI go_terms. The live funnel:

Stage Count
GO-bearing NCBIFAM models 11,228
…with both an EC and a GO term 3,782
…where ec2go(EC) confirms a model GO → exactMatch candidates 2,455 (2,503 rows)
…where ec2go(EC) would refine NCBI's broader/absent GO (the spermidine-synthase pattern, at scale) 843
…candidates already in the reviewed seed (cross-check) 17

The generated set is ncbifam2go.candidates.tsv
(2,503 rows, clearly marked generated; mapping_justification would be
semapv:CompositeMatching). The 17 rows that coincide with the reviewed seed are
exactly the seed's EC-bridge enzyme rows — an automatic confirmation that the
generator agrees with manual curation where they overlap. These 2,455 are
AMR-rich (trimethoprim-resistant dihydrofolate reductases → GO:0004146,
β-lactamases → GO:0008800, aminoglycoside 6′-N-acetyltransferases → GO:0047663,
…) and are the natural ready-to-add core of a real ncbifam2go. The 843
"refine"
models are the scaled version of the five hand-fixed altitude rows: NCBIFAM
gave a broad/near-root term but ec2go supplies the specific child — a second,
also-automatable candidate class (propose ec2go's term), pending the same
altitude/over-annotation check the FtsX case shows is still needed.

Methods

The interpro2go characterisation, InterPro member-integration counts, and the
NCBIFAM PGAP GO/EC coverage are computed live by
ncbifam_cdd_probe.py (stdlib only; no go-db, no
auth). The masking evidence (0 member signatures in WITH/FROM; 1,160 NCBIfam +
2,174 CDD DR lines) is computed from this repo's genes/**/*-goa.tsv and
*-uniprot.txt. The annotation-gain numbers are computed live against the
UniProtKB REST API by ncbifam_go_gain.py
(closure-aware go: query; see
NCBIFAM-ANNOTATION-GAIN.md). The CDD-own-GO
check uses the NCBI CDD FTP (cddannot*.dat, cddid_all.tbl) and Entrez cdd
records. The member-DB attribution (which member DB backs each GO_REF:0000002
row in the repo) is computed live by
interpro_member_attribution.py against the
InterPro API (resumable, cached). The forward closure-filtered cross-organism
contribution table reuses the UniPathway/RHEA uniqueness
query and is staged pending the go-db DuckDBs (absent in the web container). Full
queries and caveats: NCBIFAM-METHODOLOGY.md.

How this differs from RHEA, SPKW, and UniPathway

SPKW (GO_REF:0000043) RHEA (GO_REF:0000116) NCBIFAM/CDD (GO_REF:0000002)
GO aspect mostly BP MF (enzyme activity) MF + BP + CC (family/domain models)
Provenance in GOA direct (keyword visible) direct (assigned_by=RHEA) masked — only the integrated InterPro:IPR… shows
Dedicated *2go file yes yes no (ncbifam2go/cdd2go do not exist)
Dominant failure mode process conflation parent/child altitude; wrong substrate domain-altitude (CDD); unintegrated coverage gap
Built-in quality signal none reaction precision NCBIFAM family_type (equivalog)
Curation emphasis over-annotation removal gap-filling both, but attribution-first then equivalog gap-fill

NCBIFAM is unusual among these sources in carrying its own curated GO/EC that
GO does not ingest, so — like RHEA — the expected verdict skew is toward
NEW / gap-filling (for equivalog models) on the reverse side, with the
forward side dominated by the attribution/masking problem rather than outright
over-annotation.

Curation Recommendations (preliminary)

  1. Attribute before auditing. A GO_REF:0000002 annotation cannot be praised
    or blamed on NCBIFAM/CDD until GOA is re-joined to InterPro member integration;
    build that join first.
  2. Mine NCBIFAM equivalog GO/EC as a mapping source. The 13,253 equivalogs
    with NCBI go_terms/ec_numbers are the cleanest gap-fill substrate — start
    the ncbifam2go.sssom.yaml here.
  3. EC-bridge where only EC is given. An equivalog with ec_numbers but no
    go_terms can be mapped through ec2go, the RHEA EC-bridge pattern.
  4. Treat CDD as domain-altitude-risky. Without family_type, CDD's forward
    contribution skews broad; prefer the specific child the curated entry supports.
  5. Unintegrated signatures with a real function are proposed_new_terms /
    InterPro-integration requests
    , not silent gaps.

Follow-Up Targets

Target Rationale
✅ GOA × InterPro member-integration re-join Done (attribution section): NCBIFAM backs 13% / CDD 8% of the repo's InterPro2GO rows (sole signature for 250 / 116).
Forward closure-filtered cross-organism scan UniPathway-style uniqueness for member-attributed rows; needs go-db DuckDBs. Now seeded by the member-attribution join above.
Promote the 2,455 EC-bridge candidates Altitude/obsolete-check ncbifam2go.candidates.tsv and fold the clean rows into the reviewed SSSOM → a near-complete ingestible ncbifam2go.
Build the 843 "refine" class Auto-propose ec2go's specific term where NCBI's go_terms is broad/absent (the spermidine-synthase pattern), then altitude-review as FtsX shows is needed.
Non-EC families (defense/secretion/transport) The high-gain non-enzyme equivalogs (transposases, anti-phage, T4SS, encapsulins) have no EC bridge → need a different verification (literature/SPARCLE), curated like the seed's CC/BP rows.
Full-collection gain run Replace the 60-model gain sample with the complete equivalog set for a definitive reviewed-vs-TrEMBL gain figure.
Exemplar gene reviews Pick 2–3 genes whose only MF/BP support is an NCBIFAM equivalog (e.g. an anti-phage or secretion family) and run the full review workflow.

Project Status