id: ARBA00028985
description: 'A highly complex mega-rule with 357 condition sets that annotates diverse RNA-binding and RNA-processing proteins with the overly broad GO term "RNA metabolic process" (GO:0016070). The rule indiscriminately captures functionally distinct protein families including aminoacyl-tRNA synthetases, ribonucleases, RNA helicases, and RNA-binding proteins, providing minimal functional specificity.'
status: COMPLETE
rule_type: ARBA
rule:
  rule_id: ARBA00028985
  condition_sets: []  # 357 condition sets (exceeds analysis threshold)
  go_annotations:
  - go_id: GO:0016070
  reviewed_protein_count: 0
  unreviewed_protein_count: 0  # Unable to determine due to rule complexity
  created_date: null  # Not specified in rule data
  modified_date: null  # Not specified in rule data
  entries: []  # No entities analyzed due to rule complexity
review_summary: 'ARBA00028985 represents a fundamentally flawed approach to RNA-related protein annotation. With 357 condition sets, it exceeds all reasonable complexity thresholds and becomes computationally and curatorially unmanageable. The rule captures diverse RNA-associated protein families (tRNA synthetases, ribonucleases, helicases, RNA-binding proteins) and annotates them all with GO:0016070 "RNA metabolic process" - a term so broad it provides no meaningful functional information. This approach violates core GO annotation principles: specificity, biological relevance, and computational tractability. The rule creates annotation noise rather than useful functional information and obscures the specific biological roles of these important protein families. Analysis tools refused to process the rule due to its excessive complexity (>12x the recommended maximum of 30 condition sets).'
action: DEPRECATE
action_rationale: 'This rule should be completely deprecated because it fundamentally violates annotation best practices and computational feasibility constraints. The 357 condition sets make human validation impossible and exceed analysis tool limits (12x the maximum). The single GO annotation (GO:0016070) is too broad to provide useful biological insight - it is equivalent to annotating all metabolic enzymes with "metabolic process". Individual protein families captured by this rule should have specific, focused rules with appropriate molecular function and biological process terms. For example, tRNA synthetases should be annotated with specific aminoacyl-tRNA ligase activity terms, not generic RNA metabolism. The rule provides negative value by obscuring functional specificity and creating computational overhead without biological benefit.'
suggested_modifications: []  # Rule should be removed entirely, not modified
parsimony:
  assessment: OVERLY_COMPLEX
  notes: 'With 357 condition sets, this rule exceeds the recommended maximum by more than 10-fold and represents one of the most complex rules in the ARBA system. The rule attempts to unify functionally diverse RNA-associated proteins that have distinct biochemical activities, substrate specificities, and biological roles. This violates the fundamental principle that annotation rules should target coherent, mechanistically related protein sets. The complexity makes the rule impossible to validate, maintain, or troubleshoot.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
    supporting_text: 'The rule has 357 condition sets, which exceeds the recommended maximum of 12. Analysis was skipped due to excessive complexity'
  - reference_id: file:rules/arba/ARBA00028985/ARBA00028985-deep-research-manual.md
    supporting_text: 'The rule''s 357 condition sets place it among the most complex rules in the ARBA system. This level of complexity: Exceeds analysis tool limits, Prevents human validation, Creates maintenance burden, Increases false positive risk'
literature_support:
  assessment: CONTRADICTED
  notes: 'While individual protein families captured by this rule are well-characterized in the literature, the scientific literature strongly supports functional specificity in RNA metabolism annotation. Publications on tRNA synthetases, ribonucleases, and RNA helicases emphasize their distinct catalytic mechanisms, substrate specificities, and biological roles. No credible literature supports annotating functionally diverse RNA-binding proteins with a single broad term. The approach contradicts decades of biochemical and molecular biological research that has elucidated specific functions for these protein families.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
    supporting_text: 'The rule captures many different RNA-related protein families that have distinct specific functions, but annotates them all with the same broad term. This violates GO annotation principles of specificity'
condition_overlap:
  assessment: COMPLETE
  notes: 'Due to the rule complexity (357 condition sets), formal overlap analysis could not be performed. However, the inclusion of diverse protein families (synthetases, nucleases, helicases, binding proteins) virtually guarantees extensive condition overlap and mechanistic incoherence. Many RNA-binding domains are promiscuous and appear in multiple functional contexts, creating high risk for false positive annotations and functional misclassification.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
    supporting_text: 'Analysis was skipped due to excessive complexity. Rule ARBA00028985 has 357 condition sets, which exceeds the maximum of 12'
go_specificity:
  assessment: TOO_BROAD
  notes: 'GO:0016070 "RNA metabolic process" is an extremely broad parent term that encompasses virtually all RNA-related biological processes. Using this term for specific protein families is comparable to annotating all kinases with "phosphorylation" or all transcription factors with "gene expression". The term provides no actionable biological information and fails to distinguish between fundamentally different RNA-related activities. Specific child terms should be used based on actual molecular functions and biological processes.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
    supporting_text: 'The GO term "RNA metabolic process" is too general and does not provide meaningful functional information. It encompasses virtually all RNA-related processes including RNA synthesis, processing, modification, transport, and degradation'
taxonomic_scope:
  assessment: TOO_BROAD
  notes: 'The rule shows inconsistent taxonomic scoping with some condition sets restricted to specific clades (e.g., Saccharomycotina) while others appear universal. This mixed approach suggests poor rule design and potential for inappropriate cross-kingdom annotations. RNA metabolic processes vary significantly between prokaryotes and eukaryotes, requiring careful consideration of lineage-specific mechanisms and regulatory contexts.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
    supporting_text: 'Some condition sets are restricted to specific taxa (e.g., Saccharomycotina). Others appear to have no taxonomic restrictions. This inconsistency suggests poor rule design'
confidence: 0.05
references:
- id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
  title: Analysis of ARBA00028985 rule complexity and annotation issues
  findings:
  - statement: 'Rule has 357 condition sets, exceeding analysis tool limits and human validation capabilities'
  - statement: 'Single broad GO term (GO:0016070) provides no meaningful functional specificity'
  - statement: 'Captures diverse RNA-associated protein families with distinct biochemical functions'
  - statement: 'Violates fundamental GO annotation principles of specificity and biological relevance'
- id: file:rules/arba/ARBA00028985/ARBA00028985-deep-research-manual.md
  title: Deep research analysis of RNA metabolic process annotation and protein family diversity
  findings:
  - statement: 'Literature supports family-specific annotation approaches for RNA-related proteins'
  - statement: 'GO Consortium guidelines emphasize using the most specific applicable terms'
  - statement: 'Successful ARBA rules contain <20 condition sets and target coherent protein families'
  - statement: 'Domain-based annotation requires careful consideration of protein architecture and functional context'
supported_by:
- reference_id: file:rules/arba/ARBA00028985/ARBA00028985-notes.md
  supporting_text: 'This rule should be deprecated because: The GO term is too broad to be informative, the rule structure is overly complex and unmaintainable, more specific rules should exist for individual protein families, the current approach reduces annotation quality rather than improving it'