View original ARBA rule on UniProt
Overly broad rule attempting to annotate all kinase proteins across all domains of life using 1,358 condition sets with only a keyword annotation
Condition-set counts describe the sets recorded in this review, which may omit the full rule.
ARBA00022777 represents a problematic approach to kinase annotation that prioritizes coverage over accuracy. The rule attempts to capture all kinase activity across bacteria, archaea, eukaryotes, and viruses using an enormous collection of disparate domains and families. This approach violates established principles of protein function annotation and poses significant risks for database quality. The absence of GO annotations is particularly concerning, as the keyword "Kinase" provides no mechanistic or functional information to users. Modern kinase classification recognizes distinct families with different substrate specificities (protein, lipid, carbohydrate, nucleotide kinases), regulatory mechanisms, and evolutionary origins that require separate treatment. The rule's scope encompasses domains that likely include ATP-binding folds from non-kinase proteins, regulatory domains found in both kinases and non-kinases, and structural domains without catalytic function. This creates substantial risk for false positive annotations of pseudokinases, ATP-dependent helicases, chaperones, and other ATP-binding proteins. Literature supports hierarchical classification of kinases into specific families with distinct structural and functional constraints. The Manning et al. (2002) human kinome analysis and subsequent evolutionary studies demonstrate the importance of family-specific annotation rather than broad superfamily capture. Recommendation: Remove this rule and replace with a suite of specific, well-curated kinase family rules that provide meaningful GO annotations with appropriate taxonomic scope and functional specificity.
This rule exhibits fundamental design flaws that make it unsuitable for high-quality protein annotation: 1. **Extreme over-breadth**: 1,358 condition sets capturing 462 unique InterPro domains and 157 PANTHER families 2. **Lack of functional specificity**: Only provides keyword "Kinase" without GO annotations 3. **High false positive risk**: Will annotate ATP-binding proteins, pseudokinases, and structural domains as kinases 4. **Poor biological coherence**: Attempts to capture evolutionarily distinct kinase classes in one rule 5. **Violation of curation principles**: Lacks the specificity required for meaningful functional annotation
id: ARBA00022777
description: "Overly broad rule attempting to annotate all kinase proteins across all domains of life using 1,358 condition sets with only a keyword annotation"
status: COMPLETE
action: REMOVE
action_rationale: |
This rule exhibits fundamental design flaws that make it unsuitable for high-quality protein annotation:
1. **Extreme over-breadth**: 1,358 condition sets capturing 462 unique InterPro domains and 157 PANTHER families
2. **Lack of functional specificity**: Only provides keyword "Kinase" without GO annotations
3. **High false positive risk**: Will annotate ATP-binding proteins, pseudokinases, and structural domains as kinases
4. **Poor biological coherence**: Attempts to capture evolutionarily distinct kinase classes in one rule
5. **Violation of curation principles**: Lacks the specificity required for meaningful functional annotation
review_summary: |
ARBA00022777 represents a problematic approach to kinase annotation that prioritizes coverage over accuracy.
The rule attempts to capture all kinase activity across bacteria, archaea, eukaryotes, and viruses using an
enormous collection of disparate domains and families. This approach violates established principles of
protein function annotation and poses significant risks for database quality.
The absence of GO annotations is particularly concerning, as the keyword "Kinase" provides no mechanistic
or functional information to users. Modern kinase classification recognizes distinct families with different
substrate specificities (protein, lipid, carbohydrate, nucleotide kinases), regulatory mechanisms, and
evolutionary origins that require separate treatment.
The rule's scope encompasses domains that likely include ATP-binding folds from non-kinase proteins,
regulatory domains found in both kinases and non-kinases, and structural domains without catalytic function.
This creates substantial risk for false positive annotations of pseudokinases, ATP-dependent helicases,
chaperones, and other ATP-binding proteins.
Literature supports hierarchical classification of kinases into specific families with distinct structural
and functional constraints. The Manning et al. (2002) human kinome analysis and subsequent evolutionary
studies demonstrate the importance of family-specific annotation rather than broad superfamily capture.
Recommendation: Remove this rule and replace with a suite of specific, well-curated kinase family rules
that provide meaningful GO annotations with appropriate taxonomic scope and functional specificity.
confidence: 0.95
rule:
id: ARBA00022777
# Full rule count (1,358) is documented above; individual sets are not recorded.
condition_sets: []
go_annotations: [] # Only keyword annotation present
unique_interpro_domains: 462
unique_panther_families: 157
taxonomic_restrictions: 169 # out of 1358 condition sets
parsimony:
assessment: OVERLY_COMPLEX
rationale: |
With 1,358 condition sets covering hundreds of unique domains, this rule represents the antithesis of
parsimonious design. A well-designed kinase annotation system would use hierarchical rules targeting
specific kinase classes rather than attempting to capture all kinases in a single rule.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "This rule attempts to capture all kinase activity across all of biology in a single annotation rule. This violates the principle of specificity in protein function annotation."
literature_support:
assessment: CONTRADICTED
rationale: |
Extensive literature supports specific kinase family classification rather than broad superfamily annotation.
The Manning kinome analysis and subsequent studies demonstrate the importance of recognizing distinct
kinase families with specific structural and functional constraints.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "Manning et al. (2002) Science: The human kinome contains ~518 protein kinases grouped into major families (AGC, CAMK, CK1, CMGC, STE, TK, TKL)"
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "Kannan & Neuwald (2005) Proteins: Evolutionary analysis shows distinct structural and functional constraints across kinase families"
condition_overlap:
assessment: NONE
rationale: |
Each domain appears only once across the 1,358 condition sets, indicating no direct overlap but
highlighting the rule's unfocused collection of disparate kinase-related domains.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "Each domain appears only once, suggesting it's a collection of disparate kinase families"
go_specificity:
assessment: MISMATCHED
rationale: |
The complete absence of GO annotations represents a critical failure in functional annotation.
The keyword "Kinase" provides no mechanistic information and fails to distinguish between
protein kinases, lipid kinases, carbohydrate kinases, or their specific biological processes.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "The absence of GO terms is a critical flaw. The keyword 'Kinase' provides no mechanistic or functional information."
taxonomic_scope:
assessment: TOO_BROAD
rationale: |
Attempting to annotate kinases across bacteria, archaea, eukaryotes, and viruses in a single rule
ignores the distinct evolutionary origins and regulatory mechanisms of kinases in different domains of life.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "Kinases represent one of the largest and most diverse enzyme superfamilies, with distinct substrate specificities, regulatory mechanisms, cellular functions, and evolutionary origins"
false_positive_risk:
assessment: HIGH
details: |
Significant risk of annotating non-kinases including pseudokinases, ATP-binding proteins
(helicases, chaperones), and proteins with kinase-like domains but no catalytic activity.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "High likelihood of annotating as kinases: Pseudokinases, ATP-binding proteins involved in other processes (helicases, chaperones), Inactive kinase variants due to mutations in catalytic residues"
false_negative_risk:
assessment: MODERATE
details: |
Will miss novel kinase families not represented in current InterPro/PANTHER collections
and divergent kinases with atypical domain architectures.
supported_by:
- reference_id: "file:rules/arba/ARBA00022777/ARBA00022777-deep-research-manual.md"
supporting_text: "Will miss: Novel kinase families not yet represented in InterPro/PANTHER, Kinases with atypical domain architectures, Divergent kinases below detection thresholds"
proposed_improvements:
- "Split into specific kinase family rules (protein kinases, lipid kinases, sugar kinases, nucleotide kinases)"
- "Add taxonomic specificity (eukaryotic vs bacterial vs archaeal kinases)"
- "Include catalytic site requirements and conserved motifs"
- "Implement quality filters to exclude pseudokinases"
- "Add appropriate GO molecular function and biological process annotations"
- "Use minimum domain coverage and sequence identity thresholds"
entries:
analysis_status: "Manual analysis completed - rule too large for automated pairwise analysis"
total_unique_domains: 619
domain_architecture_issues: "Disparate collection of unrelated kinase domains without coherent functional theme"