ARBA00022670

View original ARBA rule on UniProt

Type: ARBA
Status: COMPLETE
Action: REMOVE
Confidence: 0.05

Description

A catastrophically over-broad rule that annotates proteins with a generic "Protease" keyword if they contain ANY of 502 different protease-related InterPro domains. The rule uses a single massive OR-logic condition set that treats highly specific peptidases (e.g., signal peptide peptidases) the same as broad families (e.g., proteasome subunits), fundamentally violating principles of specificity in functional annotation.

Analysis Summary

Condition-set counts describe the sets recorded in this review, which may omit the full rule.

0
Domain Pairs Analyzed
1
Recorded condition sets
0
Subset Relationships
0
Redundant Annotations

Review Summary

This rule represents a fundamental failure in annotation design that would severely degrade the quality of protein functional annotations. The rule contains 502 InterPro domain conditions combined with OR logic, meaning ANY protein containing ANY protease-related domain receives the same generic 'Protease' keyword annotation. This approach: (1) Violates the principle of functional specificity by treating mechanistically distinct enzyme families identically, (2) Creates massive potential for false positives by annotating proteins that may have protease domains but lack proteolytic activity, (3) Provides no GO term annotations, only a vague keyword that offers no mechanistic information, (4) Affects 2.47 million proteins with this overly broad classification. The rule essentially functions as a domain-based grep search rather than a functional annotation tool. A properly designed system would have separate rules for each major protease family (serine, metallo, aspartic, cysteine, threonine) with specific GO molecular function terms and appropriate condition sets that ensure functional annotation accuracy.

Action Rationale

This rule should be completely removed due to fundamental design flaws that make it incompatible with accurate functional annotation. Specific issues: (1) MASSIVE OVER-GENERALIZATION: 502 domains with OR logic treats pepsin (aspartic protease, pH 1.5-2.0) the same as trypsin (serine protease, pH 8.0+) - these have completely different mechanisms, substrates, and cellular contexts. (2) NO FUNCTIONAL SPECIFICITY: The rule only provides a keyword 'Protease' rather than meaningful GO terms like 'serine-type endopeptidase activity' (GO:0004252) or 'metalloendopeptidase activity' (GO:0004222). (3) FALSE POSITIVE RISK: Many proteins contain protease domains but are not active proteases (e.g., inactive zymogens, pseudoproteases, non-catalytic regulatory subunits). The rule cannot distinguish between catalytically active and inactive domains. (4) SCALE PROBLEM: Affecting 2.47 million proteins with such imprecise annotation would introduce massive noise into functional databases. (5) MAINTENANCE NIGHTMARE: With 502 conditions, this rule is essentially unmaintainable and unvalidatable. The complexity makes it impossible to assess accuracy or identify sources of error.

Rule Definition

Condition Sets

Condition Set 1

0 condition(s)
Notes:

CATASTROPHIC DESIGN: Single condition set with 502 InterPro domains connected by OR logic. This treats vastly different protease families as equivalent and assigns the same generic annotation regardless of mechanistic specificity. The rule includes: metallopeptidases (M1, M3, M8 families), serine proteases (S1, S8, S9 families), aspartic proteases (A1 family), cysteine proteases (C1, C2 families), threonine proteases (proteasome), signal peptide peptidases, and hundreds more distinct enzymatic mechanisms. This approach completely abandons the principle of functional specificity that underlies proper GO annotation.

Assessments

OVERLY_COMPLEX

This rule represents the antithesis of parsimony with 502 InterPro conditions that attempt to capture all proteolytic activity in a single annotation. The complexity is not justified by biological diversity but rather reflects a failure to properly classify functionally distinct enzyme families. A parsimonious approach would recognize that proteases represent multiple enzyme classes (EC 3.4.11-EC 3.4.99) with distinct mechanisms that require separate annotation rules. The current design is like creating a single rule for 'enzyme' that includes all of EC 1-6 - technically correct but functionally meaningless.

CONTRADICTED

The scientific literature strongly contradicts the assumption underlying this rule that all protease domains can be treated equivalently. Proteases are classified into distinct mechanistic classes based on their catalytic mechanisms: serine proteases use Ser-His-Asp catalytic triads, metalloproteases require metal cofactors (typically Zn2+), aspartic proteases use two aspartic acid residues, cysteine proteases employ Cys-His dyads, and threonine proteases (proteasomes) use N-terminal threonine. These represent fundamentally different chemical mechanisms with distinct pH optima, substrate specificities, inhibitor profiles, and regulatory mechanisms. The MEROPS database (Rawlings & Barrett, 2014) specifically organizes proteases into these mechanistic families precisely because they cannot be treated as a single functional group.

Supporting Evidence:

  • MEROPS database classification: Proteases are classified into distinct mechanistic families based on catalytic mechanism, with different families showing no evolutionary relationship and completely different active site chemistry
COMPLETE

With 502 conditions connected by OR logic, this rule has maximal overlap by design - any protein matching any single condition receives the annotation. This is the opposite of the specificity required for accurate functional annotation. Many of the included domains represent nested hierarchies (e.g., general peptidase folds vs. family-specific active sites) that create redundancy, while others represent entirely different enzymatic mechanisms that should never share annotation rules.

MISMATCHED

The rule provides NO GO annotations, only a vague keyword 'Protease'. This represents a complete failure to provide the specific functional information that GO was designed to capture. Proper annotation would require family-specific GO terms like GO:0004252 (serine-type endopeptidase activity), GO:0004222 (metalloendopeptidase activity), GO:0004190 (aspartic-type endopeptidase activity), etc. The keyword approach provides no information about mechanism, substrate specificity, or cellular context.

MISSING

The rule has no taxonomic restrictions despite significant differences in protease repertoires across kingdoms. For example, plant proteases include unique families like phytepsin (A1A subfamily) not found in animals, while bacterial proteases include signal peptidases with mechanisms distinct from eukaryotic versions. Some protease families are kingdom-specific (e.g., certain metalloproteases), and the rule's lack of taxonomic context increases false positive risk.

References (3)

Raw YAML

View Source YAML
id: ARBA00022670
description: A catastrophically over-broad rule that annotates proteins with a generic "Protease" keyword if they contain ANY of 502 different protease-related InterPro domains. The rule uses a single massive OR-logic condition set that treats highly specific peptidases (e.g., signal peptide peptidases) the same as broad families (e.g., proteasome subunits), fundamentally violating principles of specificity in functional annotation.
status: COMPLETE
rule_type: ARBA
rule:
  rule_id: ARBA00022670
  condition_sets:
  - number: 1
    conditions: []  # 502 InterPro domain conditions collapsed for brevity
    notes: "CATASTROPHIC DESIGN: Single condition set with 502 InterPro domains connected by OR logic. This treats vastly different protease families as equivalent and assigns the same generic annotation regardless of mechanistic specificity. The rule includes: metallopeptidases (M1, M3, M8 families), serine proteases (S1, S8, S9 families), aspartic proteases (A1 family), cysteine proteases (C1, C2 families), threonine proteases (proteasome), signal peptide peptidases, and hundreds more distinct enzymatic mechanisms. This approach completely abandons the principle of functional specificity that underlies proper GO annotation."
    pairwise_overlap: []  # Cannot meaningfully analyze - too many conditions
  go_annotations: []  # No GO annotations - only provides keyword
  reviewed_protein_count: 0
  unreviewed_protein_count: 2469631
  created_date: '2020-05-12'
  modified_date: '2025-05-15'
  entries: []  # Cannot analyze due to excessive complexity
review_summary: "This rule represents a fundamental failure in annotation design that would severely degrade the quality of protein functional annotations. The rule contains 502 InterPro domain conditions combined with OR logic, meaning ANY protein containing ANY protease-related domain receives the same generic 'Protease' keyword annotation. This approach: (1) Violates the principle of functional specificity by treating mechanistically distinct enzyme families identically, (2) Creates massive potential for false positives by annotating proteins that may have protease domains but lack proteolytic activity, (3) Provides no GO term annotations, only a vague keyword that offers no mechanistic information, (4) Affects 2.47 million proteins with this overly broad classification. The rule essentially functions as a domain-based grep search rather than a functional annotation tool. A properly designed system would have separate rules for each major protease family (serine, metallo, aspartic, cysteine, threonine) with specific GO molecular function terms and appropriate condition sets that ensure functional annotation accuracy."
action: REMOVE
action_rationale: "This rule should be completely removed due to fundamental design flaws that make it incompatible with accurate functional annotation. Specific issues: (1) MASSIVE OVER-GENERALIZATION: 502 domains with OR logic treats pepsin (aspartic protease, pH 1.5-2.0) the same as trypsin (serine protease, pH 8.0+) - these have completely different mechanisms, substrates, and cellular contexts. (2) NO FUNCTIONAL SPECIFICITY: The rule only provides a keyword 'Protease' rather than meaningful GO terms like 'serine-type endopeptidase activity' (GO:0004252) or 'metalloendopeptidase activity' (GO:0004222). (3) FALSE POSITIVE RISK: Many proteins contain protease domains but are not active proteases (e.g., inactive zymogens, pseudoproteases, non-catalytic regulatory subunits). The rule cannot distinguish between catalytically active and inactive domains. (4) SCALE PROBLEM: Affecting 2.47 million proteins with such imprecise annotation would introduce massive noise into functional databases. (5) MAINTENANCE NIGHTMARE: With 502 conditions, this rule is essentially unmaintainable and unvalidatable. The complexity makes it impossible to assess accuracy or identify sources of error."
suggested_modifications:
- "COMPLETE REPLACEMENT: Remove this rule entirely and replace with family-specific rules"
- "Create separate rules for major protease families (Serine: S1, S8, S9; Metallo: M1, M3, M8; Aspartic: A1; Cysteine: C1, C2; Threonine: T1) each with appropriate GO molecular function terms"
- "Each replacement rule should have 2-5 carefully selected InterPro conditions using AND logic to ensure specificity"
- "Include taxonomic restrictions where appropriate (e.g., pepsinogen activation is mammal-specific)"
- "Add negative conditions to exclude inactive variants (pseudoproteases, zymogens)"
- "Focus on catalytically active forms with evidence of proteolytic activity rather than just domain presence"
parsimony:
  assessment: OVERLY_COMPLEX
  notes: "This rule represents the antithesis of parsimony with 502 InterPro conditions that attempt to capture all proteolytic activity in a single annotation. The complexity is not justified by biological diversity but rather reflects a failure to properly classify functionally distinct enzyme families. A parsimonious approach would recognize that proteases represent multiple enzyme classes (EC 3.4.11-EC 3.4.99) with distinct mechanisms that require separate annotation rules. The current design is like creating a single rule for 'enzyme' that includes all of EC 1-6 - technically correct but functionally meaningless."
literature_support:
  assessment: CONTRADICTED
  notes: "The scientific literature strongly contradicts the assumption underlying this rule that all protease domains can be treated equivalently. Proteases are classified into distinct mechanistic classes based on their catalytic mechanisms: serine proteases use Ser-His-Asp catalytic triads, metalloproteases require metal cofactors (typically Zn2+), aspartic proteases use two aspartic acid residues, cysteine proteases employ Cys-His dyads, and threonine proteases (proteasomes) use N-terminal threonine. These represent fundamentally different chemical mechanisms with distinct pH optima, substrate specificities, inhibitor profiles, and regulatory mechanisms. The MEROPS database (Rawlings & Barrett, 2014) specifically organizes proteases into these mechanistic families precisely because they cannot be treated as a single functional group."
  supported_by:
  - reference_id: "MEROPS database classification"
    supporting_text: "Proteases are classified into distinct mechanistic families based on catalytic mechanism, with different families showing no evolutionary relationship and completely different active site chemistry"
condition_overlap:
  assessment: COMPLETE
  notes: "With 502 conditions connected by OR logic, this rule has maximal overlap by design - any protein matching any single condition receives the annotation. This is the opposite of the specificity required for accurate functional annotation. Many of the included domains represent nested hierarchies (e.g., general peptidase folds vs. family-specific active sites) that create redundancy, while others represent entirely different enzymatic mechanisms that should never share annotation rules."
go_specificity:
  assessment: MISMATCHED
  notes: "The rule provides NO GO annotations, only a vague keyword 'Protease'. This represents a complete failure to provide the specific functional information that GO was designed to capture. Proper annotation would require family-specific GO terms like GO:0004252 (serine-type endopeptidase activity), GO:0004222 (metalloendopeptidase activity), GO:0004190 (aspartic-type endopeptidase activity), etc. The keyword approach provides no information about mechanism, substrate specificity, or cellular context."
taxonomic_scope:
  assessment: MISSING
  notes: "The rule has no taxonomic restrictions despite significant differences in protease repertoires across kingdoms. For example, plant proteases include unique families like phytepsin (A1A subfamily) not found in animals, while bacterial proteases include signal peptidases with mechanisms distinct from eukaryotic versions. Some protease families are kingdom-specific (e.g., certain metalloproteases), and the rule's lack of taxonomic context increases false positive risk."
confidence: 0.05
references:
- id: "MEROPS_database"
  title: "MEROPS: the database of proteolytic enzymes, their substrates and inhibitors"
  findings:
  - statement: "Proteases are organized into distinct clans and families based on catalytic mechanism, with families within the same clan showing evolutionary relationship but families in different clans representing convergent evolution to similar function through completely different mechanisms"
- id: "Rawlings_Barrett_2014"
  title: "Introduction to peptidases and the MEROPS database"
  findings:
  - statement: "The fundamental principle of protease classification is that enzymes with similar mechanisms belong to the same family, while enzymes with different mechanisms require separate classification regardless of functional similarity"
- id: "Lopez-Otin_Bond_2008"
  title: "Proteases: multifunctional enzymes in life and disease"
  findings:
  - statement: "Human genome encodes >550 proteases representing all major mechanistic classes, with each class requiring different approaches for functional characterization, inhibitor design, and therapeutic targeting"
supported_by:
- reference_id: "MEROPS_database"
  supporting_text: "Classification of proteases into mechanistic families is essential because families represent different evolutionary origins, catalytic mechanisms, and functional properties that cannot be treated as equivalent"