ARBA00026572 nucleobase-containing compound metabolic process (GO:0006139)

View original ARBA rule on UniProt

Type: ARBA
Status: COMPLETE
Action: DEPRECATE

Description

Mega-rule with 781 condition sets applying the extremely broad GO term "nucleobase-containing compound metabolic process" to diverse protein families including DNA polymerases, tRNA synthetases, nucleotide kinases, and cofactor biosynthesis enzymes

Analysis Summary

Condition-set counts describe the sets recorded in this review, which may omit the full rule.

0
Domain Pairs Analyzed
0
Recorded condition sets
0
Subset Relationships
0
Redundant Annotations

Review Summary

ARBA00026572 is a catastrophically over-complex mega-rule with 781 condition sets that assigns the overly broad GO term "nucleobase-containing compound metabolic process" to diverse protein families including DNA polymerases, tRNA synthetases, nucleotide kinases, and cofactor biosynthesis enzymes. The rule violates fundamental GO curation principles by using a parent-level term for proteins with clearly defined specific functions, creating a maintenance nightmare that provides no biological insight beyond "involves nucleotides".

Action Rationale

Complete removal is required due to fundamental structural and conceptual flaws. The 781 condition sets make the rule unanalyzable and unmaintainable. The GO term GO:0006139 is inappropriately broad for automated annotation of functionally distinct enzyme classes. Each functional group (DNA repair, tRNA charging, nucleotide synthesis) deserves separate rules with specific GO terms. The mega-rule approach contradicts literature-supported best practices for GO annotation.

GO Annotations

GO:0006139 - nucleobase-containing compound metabolic process
Aspect: BP

Rule Definition

Assessments

OVERLY_COMPLEX

With 781 condition sets, this rule represents the antithesis of parsimony. The sheer number of conditions makes the rule unanalyzable and unmaintainable. Each functional class (DNA repair, tRNA charging, nucleotide synthesis) should be separate rules with 2-5 specific condition sets. The current structure conflates dozens of distinct biochemical functions under a meaningless umbrella term. Complete reconstruction required.

CONTRADICTED

While individual condition sets may represent valid protein families, the literature contradicts using a single broad GO term for functionally diverse enzymes. GO Consortium guidelines emphasize specificity and functional coherence. Literature supports distinct specific terms: DNA polymerases require GO:0006281 (DNA repair), tRNA synthetases require GO:0006418 (aminoacylation), nucleotide kinases require metabolic pathway-specific terms. The mega-rule approach lacks literature support and violates established curation principles.

Supporting Evidence:

  • file:rules/arba/ARBA00026572/ARBA00026572-deep-research-manual.md: Deep research analysis confirms GO:0006139 is too broad for automated annotation. Literature supports using specific terms: DNA polymerases require GO:0006281 (DNA repair), tRNA synthetases require GO:0006418 (aminoacylation), nucleotide kinases require metabolic pathway-specific terms. GO Consortium guidelines emphasize specificity and functional coherence, which this mega-rule violates.
SIGNIFICANT

Analysis impossible due to 781 condition sets exceeding computational limits. However, sampling reveals heterogeneous functional classes: UmuC domains (DNA repair), aminoacyl-tRNA synthetases (protein synthesis), NAD synthases (cofactor biosynthesis), deoxynucleoside kinases (salvage pathways). These represent completely different biochemical functions with no meaningful overlap. The condition sets should be grouped by function into separate rules rather than artificially unified under a broad GO term.

TOO_BROAD

GO:0006139 (nucleobase-containing compound metabolic process) is completely inappropriate for an automated annotation rule. This parent-level term should only be used when specific pathways cannot be determined. The rule captures proteins with clearly defined functions that deserve specific terms: DNA repair (GO:0006281), tRNA aminoacylation (GO:0006418), nucleotide biosynthesis (GO:0009165), etc. Using GO:0006139 provides no biological insight beyond "involves nucleotides" and violates the GO specificity principle.

UNNECESSARY

Taxonomic restrictions appear arbitrary and inconsistent across the 781 condition sets. Some target Eukaryota, others Metazoa, others Saccharomycotina with no clear biological rationale. Since GO:0006139 represents universal cellular processes, taxonomic restrictions are inappropriate at this broad level. However, when the rule is properly split into functional classes, specific restrictions may be justified (e.g., selenocysteine-containing enzymes in specific lineages).

References (1)

Raw YAML

View Source YAML
id: ARBA00026572
description: 'Mega-rule with 781 condition sets applying the extremely broad GO term "nucleobase-containing compound metabolic process" to diverse protein families including DNA polymerases, tRNA synthetases, nucleotide kinases, and cofactor biosynthesis enzymes'
status: COMPLETE
rule_type: ARBA
rule:
  rule_id: ARBA00026572
  condition_sets: []
  go_annotations:
  - go_id: GO:0006139
    go_label: nucleobase-containing compound metabolic process
    aspect: BP
  reviewed_protein_count: 0
  unreviewed_protein_count: 0
  created_date: ''
  modified_date: ''
  entries: []
review_summary: 'ARBA00026572 is a catastrophically over-complex mega-rule with 781 condition sets that assigns the overly broad GO term "nucleobase-containing compound metabolic process" to diverse protein families including DNA polymerases, tRNA synthetases, nucleotide kinases, and cofactor biosynthesis enzymes. The rule violates fundamental GO curation principles by using a parent-level term for proteins with clearly defined specific functions, creating a maintenance nightmare that provides no biological insight beyond "involves nucleotides".'
action: DEPRECATE
action_rationale: 'Complete removal is required due to fundamental structural and conceptual flaws. The 781 condition sets make the rule unanalyzable and unmaintainable. The GO term GO:0006139 is inappropriately broad for automated annotation of functionally distinct enzyme classes. Each functional group (DNA repair, tRNA charging, nucleotide synthesis) deserves separate rules with specific GO terms. The mega-rule approach contradicts literature-supported best practices for GO annotation.'
suggested_modifications:
- 'Replace with multiple smaller rules: DNA repair enzymes (GO:0006281), tRNA synthetases (GO:0006418), nucleotide biosynthesis enzymes (GO:0009165), salvage pathway enzymes (pathway-specific terms)'
- 'Limit each new rule to maximum 12 condition sets with functional coherence'
- 'Ensure taxonomic restrictions have clear biological rationale'
parsimony:
  assessment: OVERLY_COMPLEX
  notes: 'With 781 condition sets, this rule represents the antithesis of parsimony. The sheer number of conditions makes the rule unanalyzable and unmaintainable. Each functional class (DNA repair, tRNA charging, nucleotide synthesis) should be separate rules with 2-5 specific condition sets. The current structure conflates dozens of distinct biochemical functions under a meaningless umbrella term. Complete reconstruction required.'
literature_support:
  assessment: CONTRADICTED
  notes: 'While individual condition sets may represent valid protein families, the literature contradicts using a single broad GO term for functionally diverse enzymes. GO Consortium guidelines emphasize specificity and functional coherence. Literature supports distinct specific terms: DNA polymerases require GO:0006281 (DNA repair), tRNA synthetases require GO:0006418 (aminoacylation), nucleotide kinases require metabolic pathway-specific terms. The mega-rule approach lacks literature support and violates established curation principles.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00026572/ARBA00026572-deep-research-manual.md
    supporting_text: 'Deep research analysis confirms GO:0006139 is too broad for automated annotation. Literature supports using specific terms: DNA polymerases require GO:0006281 (DNA repair), tRNA synthetases require GO:0006418 (aminoacylation), nucleotide kinases require metabolic pathway-specific terms. GO Consortium guidelines emphasize specificity and functional coherence, which this mega-rule violates.'
condition_overlap:
  assessment: SIGNIFICANT
  notes: 'Analysis impossible due to 781 condition sets exceeding computational limits. However, sampling reveals heterogeneous functional classes: UmuC domains (DNA repair), aminoacyl-tRNA synthetases (protein synthesis), NAD synthases (cofactor biosynthesis), deoxynucleoside kinases (salvage pathways). These represent completely different biochemical functions with no meaningful overlap. The condition sets should be grouped by function into separate rules rather than artificially unified under a broad GO term.'
  supported_by: []
go_specificity:
  assessment: TOO_BROAD
  notes: 'GO:0006139 (nucleobase-containing compound metabolic process) is completely inappropriate for an automated annotation rule. This parent-level term should only be used when specific pathways cannot be determined. The rule captures proteins with clearly defined functions that deserve specific terms: DNA repair (GO:0006281), tRNA aminoacylation (GO:0006418), nucleotide biosynthesis (GO:0009165), etc. Using GO:0006139 provides no biological insight beyond "involves nucleotides" and violates the GO specificity principle.'
  supported_by: []
taxonomic_scope:
  assessment: UNNECESSARY
  notes: 'Taxonomic restrictions appear arbitrary and inconsistent across the 781 condition sets. Some target Eukaryota, others Metazoa, others Saccharomycotina with no clear biological rationale. Since GO:0006139 represents universal cellular processes, taxonomic restrictions are inappropriate at this broad level. However, when the rule is properly split into functional classes, specific restrictions may be justified (e.g., selenocysteine-containing enzymes in specific lineages).'
  supported_by: []
confidence: 0.0
references:
- id: file:rules/arba/ARBA00026572/ARBA00026572-deep-research-manual.md
  title: Deep research analysis of ARBA00026572 nucleobase metabolism mega-rule
  findings:
  - statement: 'GO:0006139 (nucleobase-containing compound metabolic process) is too broad for automated annotation and should only be used when specific pathways cannot be determined'
  - statement: 'The 781 condition sets represent diverse functional classes (DNA repair, tRNA charging, nucleotide biosynthesis, salvage pathways) that require distinct specific GO terms'
  - statement: 'GO Consortium guidelines emphasize specificity principle: use most specific applicable term, avoid over-annotation with broad terms when specific ones are available'
  - statement: 'Rule represents failed mega-rule approach that provides little biological value while introducing significant false positive risk'
supported_by: []