View original ARBA rule on UniProt
Extremely complex rule attempting to capture all chromatin remodeling activities across eukaryotes using 134 condition sets spanning ATP-dependent remodelers, histone-modifying enzymes, chromatin assembly factors, and pioneer transcription factors.
Condition-set counts describe the sets recorded in this review, which may omit the full rule.
ARBA00027994 represents one of the most problematic rules in the ARBA system due to its extreme complexity and functional over-breadth. With 134 condition sets, this rule attempts to capture the entire landscape of chromatin remodeling proteins across all eukaryotes in a single annotation unit. CRITICAL PROBLEMS IDENTIFIED: 1. UNMANAGEABLE COMPLEXITY: The rule contains 134 condition sets, far exceeding the practical limit for rule analysis and maintenance. This makes it impossible to properly validate or troubleshoot. 2. FUNCTIONAL CONFLATION: The rule incorrectly groups together proteins with fundamentally different mechanisms: - ATP-dependent chromatin remodeling complexes (SWI/SNF, ISWI, CHD) - Histone acetyltransferases and deacetylases - Histone methyltransferases and demethylases - Chromatin assembly factors - Pioneer transcription factors 3. GO TERM MISUSE: GO:0006338 "chromatin remodeling" is extremely broad and encompasses many distinct biological processes that should use more specific child terms. 4. HIGH FALSE POSITIVE RISK: Many condition sets rely on promiscuous domains: - Bromodomains (IPR001487) appear in non-chromatin proteins - SANT domains (IPR001005) exist in non-remodeling complexes - Single-domain conditions (e.g., IPR055197) lack specificity 5. TAXONOMIC INCONSISTENCIES: The rule mixes: - Universal eukaryotic conditions - Lineage-specific adaptations - Genus-specific conditions (Homo, Mus, Saccharomyces) 6. LACK OF MECHANISTIC PRECISION: The rule fails to distinguish between: - Writers vs. readers vs. erasers of histone marks - ATP-dependent vs. ATP-independent processes - Structural vs. enzymatic components BIOLOGICAL ASSESSMENT: While chromatin remodeling is a well-established and important process across eukaryotes, this rule's approach is fundamentally flawed. The attempt to create a comprehensive "chromatin remodeling" rule results in a system that is too complex to maintain and too imprecise to provide useful annotations. RECOMMENDATION: SPLIT INTO MECHANISM-SPECIFIC RULES This rule should be split into multiple, focused rules using valid current GO terms: - ATP-dependent chromatin remodeler activity (GO:0140658, MF) - for SWI/SNF, ISWI, CHD, INO80 families - Histone modifying activity (GO:0140993, MF) - or more specific child terms per modification type - Histone deacetylase activity (GO:0004407, MF) - for HDAC enzymes - DNA replication-dependent chromatin assembly (GO:0006335, BP) - for chromatin assembly factors NOTE: Previously suggested GO terms GO:0043044, GO:0016573, GO:0016575, and GO:0016570 are all OBSOLETE and must not be used. Each replacement rule should have 5-10 well-characterized condition sets with appropriate negative constraints and taxonomic scope validation.
This rule is fundamentally flawed due to: 1. EXCESSIVE COMPLEXITY: 134 condition sets exceed any reasonable maintenance threshold 2. FUNCTIONAL OVER-BROADNESS: Conflates distinct mechanisms (ATP-dependent remodeling, histone modifications, chromatin assembly) 3. HIGH FALSE POSITIVE RISK: Single-domain conditions and promiscuous domains (e.g. UBR7/Q8N806, a ubiquitin ligase, falsely annotated to chromatin remodeling due to PHD-like zinc finger domain IPR011011; see geneontology/go-annotation#5899) 4. TAXONOMIC INCONSISTENCIES: Mix of universal, lineage-specific, and genus-specific conditions 5. GO TERM MISUSE: GO:0006338 is too broad for the diverse protein functions covered The rule should be split into multiple mechanism-specific rules using appropriate GO terms from the current ontology.
With 134 condition sets, this rule far exceeds any reasonable complexity threshold. The analysis system cannot even process rules with more than 12 condition sets due to computational constraints. This represents a fundamental design failure where the rule attempts to solve too broad a problem in a single annotation unit.
While chromatin remodeling as a biological process is extremely well-supported in the literature, this rule's specific approach of conflating different mechanisms lacks proper literature support. The individual components (SWI/SNF complexes, histone modifying enzymes, etc.) are well-characterized, but grouping them under a single broad annotation is not scientifically justified.
With 134 condition sets covering overlapping domain architectures and taxonomic groups, there is almost certainly significant redundancy. Many proteins likely match multiple condition sets, and some condition sets may be entirely redundant with others. Without proper analysis tools for this scale, the extent of overlap cannot be precisely quantified, but it is inevitably substantial.
GO:0006338 "chromatin remodeling" is far too broad for this rule. It encompasses ATP-dependent chromatin remodeling, histone modifications, chromatin assembly, and pioneer transcription factor activities - each of which should use more specific child terms. The rule conflates writers, readers, and erasers of histone modifications with ATP-dependent remodelers and structural factors.
The rule shows inconsistent taxonomic scope, mixing universal eukaryotic conditions with highly specific genus-level conditions (Homo, Mus, Saccharomyces). Some universal conditions may annotate proteins in organisms lacking the specific chromatin structures or regulatory mechanisms, while genus-specific conditions are unnecessarily narrow and create annotation bias.
Rule contains 134 condition sets making it one of the most complex and problematic rules in the ARBA system
GO:0006338 is extremely broad encompassing ATP-dependent remodeling, histone modifications, chromatin assembly, and pioneer transcription factors
High risk of false positives due to promiscuous domains and single-domain conditions
Taxonomic scope inconsistencies mixing universal, lineage-specific, and genus-specific conditions
UBR7 (Q8N806) is a ubiquitin protein ligase falsely annotated to chromatin remodeling due to PHD-like zinc finger domain (IPR011011) shared with chromatin-associated proteins
UBR7 has GO annotations for ubiquitin protein ligase activity (GO:0061630) and protein ubiquitination (GO:0016567) but no legitimate chromatin remodeling function
id: ARBA00027994
description: Extremely complex rule attempting to capture all chromatin remodeling activities across eukaryotes using 134 condition sets spanning ATP-dependent remodelers, histone-modifying enzymes, chromatin assembly factors, and pioneer transcription factors.
status: COMPLETE
rule_type: ARBA
rule:
rule_id: ARBA00027994
condition_sets: []
go_annotations:
- go_id: GO:0006338
go_label: chromatin remodeling
aspect: BP
reviewed_protein_count: 0
unreviewed_protein_count: 0
created_date: '2021-10-20'
modified_date: '2025-09-20'
entries: []
action: SPLIT
action_rationale: |
This rule is fundamentally flawed due to:
1. EXCESSIVE COMPLEXITY: 134 condition sets exceed any reasonable maintenance threshold
2. FUNCTIONAL OVER-BROADNESS: Conflates distinct mechanisms (ATP-dependent remodeling, histone modifications, chromatin assembly)
3. HIGH FALSE POSITIVE RISK: Single-domain conditions and promiscuous domains (e.g. UBR7/Q8N806, a ubiquitin ligase, falsely annotated to chromatin remodeling due to PHD-like zinc finger domain IPR011011; see geneontology/go-annotation#5899)
4. TAXONOMIC INCONSISTENCIES: Mix of universal, lineage-specific, and genus-specific conditions
5. GO TERM MISUSE: GO:0006338 is too broad for the diverse protein functions covered
The rule should be split into multiple mechanism-specific rules using appropriate GO terms from the current ontology.
review_summary: |
ARBA00027994 represents one of the most problematic rules in the ARBA system due to its extreme complexity and functional over-breadth. With 134 condition sets, this rule attempts to capture the entire landscape of chromatin remodeling proteins across all eukaryotes in a single annotation unit.
CRITICAL PROBLEMS IDENTIFIED:
1. UNMANAGEABLE COMPLEXITY: The rule contains 134 condition sets, far exceeding the practical limit for rule analysis and maintenance. This makes it impossible to properly validate or troubleshoot.
2. FUNCTIONAL CONFLATION: The rule incorrectly groups together proteins with fundamentally different mechanisms:
- ATP-dependent chromatin remodeling complexes (SWI/SNF, ISWI, CHD)
- Histone acetyltransferases and deacetylases
- Histone methyltransferases and demethylases
- Chromatin assembly factors
- Pioneer transcription factors
3. GO TERM MISUSE: GO:0006338 "chromatin remodeling" is extremely broad and encompasses many distinct biological processes that should use more specific child terms.
4. HIGH FALSE POSITIVE RISK: Many condition sets rely on promiscuous domains:
- Bromodomains (IPR001487) appear in non-chromatin proteins
- SANT domains (IPR001005) exist in non-remodeling complexes
- Single-domain conditions (e.g., IPR055197) lack specificity
5. TAXONOMIC INCONSISTENCIES: The rule mixes:
- Universal eukaryotic conditions
- Lineage-specific adaptations
- Genus-specific conditions (Homo, Mus, Saccharomyces)
6. LACK OF MECHANISTIC PRECISION: The rule fails to distinguish between:
- Writers vs. readers vs. erasers of histone marks
- ATP-dependent vs. ATP-independent processes
- Structural vs. enzymatic components
BIOLOGICAL ASSESSMENT:
While chromatin remodeling is a well-established and important process across eukaryotes, this rule's approach is fundamentally flawed. The attempt to create a comprehensive "chromatin remodeling" rule results in a system that is too complex to maintain and too imprecise to provide useful annotations.
RECOMMENDATION: SPLIT INTO MECHANISM-SPECIFIC RULES
This rule should be split into multiple, focused rules using valid current GO terms:
- ATP-dependent chromatin remodeler activity (GO:0140658, MF) - for SWI/SNF, ISWI, CHD, INO80 families
- Histone modifying activity (GO:0140993, MF) - or more specific child terms per modification type
- Histone deacetylase activity (GO:0004407, MF) - for HDAC enzymes
- DNA replication-dependent chromatin assembly (GO:0006335, BP) - for chromatin assembly factors
NOTE: Previously suggested GO terms GO:0043044, GO:0016573, GO:0016575, and GO:0016570 are all OBSOLETE and must not be used.
Each replacement rule should have 5-10 well-characterized condition sets with appropriate negative constraints and taxonomic scope validation.
suggested_modifications:
- Split the rule into mechanism-specific rules due to excessive complexity and functional over-breadth
- 'Replace with focused rules using valid current GO terms: GO:0140658 (ATP-dependent chromatin remodeler activity, MF), GO:0140993 (histone modifying activity, MF), GO:0004407 (histone deacetylase activity, MF), GO:0006335 (DNA replication-dependent chromatin assembly, BP)'
- 'NOTE: Previously suggested terms GO:0043044, GO:0016573, GO:0016575, GO:0016570 are OBSOLETE and must not be used'
- Each replacement should have 5-10 condition sets maximum with proper validation
- Add negative constraints to exclude pseudoenzymes and non-functional domains (e.g. exclude UBR7-like proteins with PHD-like domains but primary ubiquitin ligase function)
- Validate taxonomic scope for each replacement rule against experimental evidence
parsimony:
assessment: OVERLY_COMPLEX
notes: |
With 134 condition sets, this rule far exceeds any reasonable complexity threshold. The analysis system cannot even process rules with more than 12 condition sets due to computational constraints. This represents a fundamental design failure where the rule attempts to solve too broad a problem in a single annotation unit.
supported_by:
- reference_id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
supporting_text: "134 condition sets is far beyond manageable complexity. High likelihood of overlapping and redundant conditions. Impossible to manually validate all conditions. Creates maintenance nightmare."
literature_support:
assessment: MODERATE
notes: |
While chromatin remodeling as a biological process is extremely well-supported in the literature, this rule's specific approach of conflating different mechanisms lacks proper literature support. The individual components (SWI/SNF complexes, histone modifying enzymes, etc.) are well-characterized, but grouping them under a single broad annotation is not scientifically justified.
supported_by:
- reference_id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
supporting_text: "SWI/SNF complexes are conserved from yeast to humans (Clapier & Cairns, 2009). ISWI and CHD families show evolutionary conservation (Flaus & Owen-Hughes, 2011). Histone modifications are universal eukaryotic mechanisms (Kouzarides, 2007)."
condition_overlap:
assessment: SIGNIFICANT
notes: |
With 134 condition sets covering overlapping domain architectures and taxonomic groups, there is almost certainly significant redundancy. Many proteins likely match multiple condition sets, and some condition sets may be entirely redundant with others. Without proper analysis tools for this scale, the extent of overlap cannot be precisely quantified, but it is inevitably substantial.
supported_by:
- reference_id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
supporting_text: "High likelihood of overlapping and redundant conditions. Groups together proteins with fundamentally different mechanisms."
go_specificity:
assessment: TOO_BROAD
notes: |
GO:0006338 "chromatin remodeling" is far too broad for this rule. It encompasses ATP-dependent chromatin remodeling, histone modifications, chromatin assembly, and pioneer transcription factor activities - each of which should use more specific child terms. The rule conflates writers, readers, and erasers of histone modifications with ATP-dependent remodelers and structural factors.
supported_by:
- reference_id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
supporting_text: "GO:0006338 is extremely broad, encompassing: ATP-dependent chromatin remodeling, Histone modifications, Chromatin assembly/disassembly, Pioneer transcription factor activity. Many proteins may perform only specific sub-functions."
taxonomic_scope:
assessment: TOO_BROAD
notes: |
The rule shows inconsistent taxonomic scope, mixing universal eukaryotic conditions with highly specific genus-level conditions (Homo, Mus, Saccharomyces). Some universal conditions may annotate proteins in organisms lacking the specific chromatin structures or regulatory mechanisms, while genus-specific conditions are unnecessarily narrow and create annotation bias.
supported_by:
- reference_id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
supporting_text: "Universal eukaryotic conditions may annotate proteins in organisms lacking specific chromatin structures. Some conditions may be too narrow (e.g., genus-specific like Homo, Mus)."
confidence: 0.6
references:
- id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
title: Deep research analysis for ARBA00027994 chromatin remodeling rule
findings:
- statement: Rule contains 134 condition sets making it one of the most complex and problematic rules in the ARBA system
- statement: GO:0006338 is extremely broad encompassing ATP-dependent remodeling, histone modifications, chromatin assembly, and pioneer transcription factors
- statement: High risk of false positives due to promiscuous domains and single-domain conditions
- statement: Taxonomic scope inconsistencies mixing universal, lineage-specific, and genus-specific conditions
- id: 'github:geneontology/go-annotation#5899'
title: False positive chromatin remodeling annotation for UBR7 (Q8N806) from ARBA00027994
findings:
- statement: UBR7 (Q8N806) is a ubiquitin protein ligase falsely annotated to chromatin remodeling due to PHD-like zinc finger domain (IPR011011) shared with chromatin-associated proteins
- statement: UBR7 has GO annotations for ubiquitin protein ligase activity (GO:0061630) and protein ubiquitination (GO:0016567) but no legitimate chromatin remodeling function
supported_by:
- reference_id: file:rules/arba/ARBA00027994/ARBA00027994-deep-research-manual.md
supporting_text: "This is an exceptionally complex ARBA rule with 134 condition sets, making it one of the largest and most comprehensive rules in the ARBA system. The rule attempts to capture the entire landscape of chromatin remodeling proteins across multiple taxonomic groups."