id: ARBA00027169
description: 'Annotates proteins containing Sm or Sm-like domains with GO:0120114 "Sm-like protein family complex" using 28 condition sets covering classical Sm proteins, Lsm proteins, and various spliceosomal components across different taxonomic groups'
status: COMPLETE
rule_type: ARBA
rule:
  rule_id: ARBA00027169
  condition_sets: []
  go_annotations: []
  reviewed_protein_count: 0
  unreviewed_protein_count: 0
  created_date: ''
  modified_date: ''
  entries: []
review_summary: 'This rule is an overly complex "mega-rule" with 28 condition sets that attempts to capture too many functionally related but distinct protein families. While it correctly identifies many legitimate Sm/Lsm proteins and spliceosomal components, it suffers from excessive complexity, taxonomic inconsistency, and inclusion of incorrectly classified proteins (e.g., translation elongation factors). The rule violates parsimony principles and would benefit from being split into multiple focused rules with consistent taxonomic scope.'
action: DEPRECATE
action_rationale: 'Recommended for deprecation due to excessive complexity (28 condition sets exceeding the 12 condition limit for analysis), taxonomic inconsistency (mixing broad eukaryotic scope with very specific restrictions), functional heterogeneity (including non-Sm proteins like translation elongation factors), and violation of parsimony principles. The rule should be replaced with multiple focused rules covering: (1) core Sm proteins, (2) Lsm proteins, (3) U7-specific components, and (4) other spliceosomal factors with appropriate taxonomic constraints and more specific GO terms.'
suggested_modifications:
- 'Split into separate focused rules for core Sm proteins (SmB/B prime, SmD1-3, SmE-G)'
- 'Create dedicated Lsm protein rule for U6 snRNP and mRNA decay components (Lsm1-8)'
- 'Establish U7-specific rule for histone mRNA processing (Lsm10/11)'
- 'Remove incorrectly included proteins (translation elongation factor in condition set 13)'
- 'Establish consistent taxonomic scope across all condition sets'
- 'Use more specific GO terms where appropriate (e.g., specific snRNP component terms)'
- 'Add negative conditions to exclude false positives from promiscuous Sm domains'
parsimony:
  assessment: OVERLY_COMPLEX
  notes: 'Rule has 28 condition sets, far exceeding the recommended limit of 12 and violating parsimony principles. Many condition sets show functional overlap and could be consolidated. The mixture of broad InterPro families with highly specific CATH FunFams creates unnecessary complexity. Taxonomic restrictions are inconsistent, ranging from broad eukaryotic scope to genus-specific conditions (Rattus).'
literature_support:
  assessment: MODERATE
  notes: 'Sm and Lsm proteins are well-characterized components of RNA processing machinery with strong literature support for their roles in snRNP assembly and function. However, the rule includes proteins with diverse functions beyond classical Sm/Lsm roles, and some incorrectly classified proteins lack support for Sm-like complex participation.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00027169/ARBA00027169-deep-research-manual.md
    supporting_text: 'Sm proteins are core components of small nuclear ribonucleoproteins (snRNPs), which are essential for pre-mRNA splicing. The Sm protein family includes classical Sm proteins (SmB/B prime, SmD1, SmD2, SmD3, SmE, SmF, SmG) and Sm-like (Lsm) proteins (Lsm1-8) components of U6 snRNP and involved in mRNA degradation.'
condition_overlap:
  assessment: SIGNIFICANT
  notes: 'Unable to perform quantitative analysis due to excessive number of condition sets (28 > 12 limit). Manual inspection reveals significant functional overlap between InterPro-based conditions (sets 1-4) and many CATH FunFam conditions targeting the same protein families. Multiple condition sets target the same proteins with different domain signatures, creating redundancy.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00027169/ARBA00027169-analysis-manual.md
    supporting_text: 'With 28 condition sets, this rule is extremely complex and likely contains significant redundancy. The mixing of broad eukaryotic conditions with very narrow taxonomic restrictions suggests poor rule design.'
go_specificity:
  assessment: TOO_BROAD
  notes: 'GO:0120114 "Sm-like protein family complex" is too broad for the diverse proteins captured by this rule. Many proteins would be better annotated with specific snRNP component terms (e.g., GO:0071004 U2-type spliceosomal complex, GO:0071007 U2-type catalytic step 1 spliceosome). The cellular component term does not distinguish between different functional roles.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00027169/ARBA00027169-analysis-manual.md
    supporting_text: 'GO:0120114 "Sm-like protein family complex" is a cellular component term that may be too broad for the diverse proteins captured by this rule, which includes legitimate Sm/Lsm proteins, spliceosomal components without Sm domains, and potentially unrelated proteins.'
taxonomic_scope:
  assessment: TOO_BROAD
  notes: 'Taxonomic scope is highly inconsistent across condition sets. Some conditions apply broadly to all Eukaryota, while others are restricted to specific lineages (Taphrinomycotina, Ascomycota, Fungi) or even single genera (Rattus). This inconsistency suggests poor rule design and may lead to inappropriate annotations across taxonomic groups where the proteins have different functions.'
  supported_by:
  - reference_id: file:rules/arba/ARBA00027169/ARBA00027169-analysis-manual.md
    supporting_text: 'Some conditions are restricted to specific taxa (Taphrinomycotina, Fungi, Ascomycota), others apply broadly to Eukaryota, some are restricted to vertebrates (Chordata, Craniata, Mammalia), and one condition is specific to Rattus.'
confidence: 0.2
references:
- id: file:rules/arba/ARBA00027169/ARBA00027169-deep-research-manual.md
  title: Manual deep research analysis of Sm/Lsm proteins
  findings:
  - statement: 'Sm proteins are well-characterized core components of snRNPs essential for pre-mRNA splicing'
  - statement: 'Rule includes legitimate Sm/Lsm proteins but also captures non-Sm spliceosomal components and incorrectly classified proteins'
  - statement: 'Excessive complexity with 28 condition sets violates parsimony principles'
- id: file:rules/arba/ARBA00027169/ARBA00027169-analysis-manual.md
  title: Manual structural analysis of rule condition sets
  findings:
  - statement: 'Rule shows significant functional overlap and taxonomic inconsistency'
  - statement: 'Translation elongation factor incorrectly included in condition set 13'
  - statement: 'Recommended for splitting into multiple focused rules'
supported_by:
- reference_id: file:rules/arba/ARBA00027169/ARBA00027169.enriched.json
  supporting_text: 'Rule ARBA00027169 contains 28 condition sets covering InterPro families IPR001163 (Sm domain), IPR010920 (LSM superfamily), multiple CATH FunFams for spliceosomal components, with inconsistent taxonomic restrictions ranging from Eukaryota to genus-specific (Rattus)'
