Summary
The pipeline compares the JmjC domain of Epe1 (UniProt O94603) with six active
human JmjC demethylases and with fungal epe1 gene products and fission-yeast JmjC
proteins. Domain boundaries and Fe(II) ligands are read from UniProt feature
records, never from a sequence motif. The results, all regenerated by
just all in this folder, are:
- Epe1 JmjC domain: 243-402 (UniProt FT DOMAIN).
- Epe1 Fe(II) ligands: H297 and E299 (UniProt FT BINDING). The UniProt CC
CAUTION states that the catalytic His at position 370 is replaced by Tyr, and
the sequence has Y at 370.
- Active comparators: each of the six has three annotated Fe(II) ligands in
the His-(Asp/Glu)-His pattern (table below). Epe1 has two.
- Alignment: at the column of Epe1 Y370, all six active comparators carry
their third (His) Fe ligand. Two other Schizosaccharomyces epe1 genes and
two related unnamed Schizosaccharomyces JmjC proteins also carry Tyr there.
- What the sequence does not show: whether Epe1 binds Fe(II), binds
2-oxoglutarate, or has any catalytic activity. No activity verdict is drawn
from sequence alone.
All positions in this file are 1-based residue numbers of the full-length
protein. A motif is identified by the position of its His (for example
HVD@280 is H280-V281-D282).
Pipeline
Run from this folder, in its own uv project:
just all # 01 fetch, 02 domain/ligands, 03 alignment, 04 regions, 05 structure
just test # doctests for uniprot_features.py + pytest test_pipeline.py
| Step |
Script |
Output |
| 1 |
01_fetch_sequences.py |
data/ (Epe1 and comparator FASTA + UniProt JSON) |
| 2 |
02_jmjc_domain_analysis.py |
results/jmjc_domain_analysis.txt |
| 3 |
03_conservation_analysis.py |
results/jmjc_domains.fasta, results/jmjc_domains_aligned.fasta, results/conservation_analysis.txt |
| 4 |
04_functional_regions_analysis.py |
results/functional_regions_analysis.txt, results/epe1_analysis_summary.png |
| 5 |
05_structural_features.py |
results/structural_analysis.txt, results/epe1_domain_architecture.png |
uniprot_features.py parses the Epe1 flat file (../Epe1-uniprot.txt) and
the comparators' UniProt JSON. Step 3 needs mafft on PATH.
Results
1. Fe(II) ligands from UniProt (results/jmjc_domain_analysis.txt)
| Protein |
Accession |
JmjC |
Fe(II) ligands (UniProt) |
| Epe1 |
O94603 |
243-402 |
H297, E299 |
| KDM2A_HUMAN |
Q9Y2K7 |
148-316 |
H212, D214, H284 |
| KDM3A_HUMAN |
Q9Y4C1 |
1058-1281 |
H1120, D1122, H1249 |
| KDM4A_HUMAN |
O75164 |
142-308 |
H188, E190, H276 |
| KDM5B_HUMAN |
Q9UGL1 |
453-619 |
H499, E501, H587 |
| KDM5C_HUMAN |
P41229 |
468-634 |
H514, E516, H602 |
| KDM7A_HUMAN |
Q6ZMT4 |
230-386 |
H282, D284, H354 |
Other annotated Epe1 positions: T294 and K314 (FT BINDING, ligand
"substrate"), and Y307 (FT MUTAGEN, "Y->A: Loss of function").
Motif scan (cross-check only). An H.[DE] scan inside each JmjC domain
does not locate the Fe(II) site on its own. In Epe1 it hits HVD@280, which is
not an annotated ligand, and HIE@297, which is. In active KDM2A the annotated
ligand H212 is itself an HVD motif. In KDM4A an HVD@144 hit is not a ligand.
The earlier claim that "HVD at 280" is Epe1's defective iron site came from
reading the first scan hit as the site, and it is retracted.
2. Alignment at Epe1's annotated sites (results/conservation_analysis.txt)
The JmjC domains of 18 proteins were aligned with MAFFT. These are Epe1, the six
comparators, the S. japonicus and S. osmophilus epe1 gene products (B6K4V0,
A0AAF0AZ59), unnamed S. octosporus and S. cryophilus JmjC proteins (S9R0Y2,
S9XEB0), and S. pombe and S. japonicus JmjC proteins returned by
the UniProt search. LUC7_YEAST was returned by the search but has no JmjC
domain feature and was skipped. At each annotated Epe1 site the output lists
every protein's residue:
- H297, E299 (Epe1 Fe ligands): these align with the first two annotated
Fe ligands of every active comparator except KDM3A. KDM3A's long JmjC
domain aligns poorly in this region.
- Y370 (CC CAUTION position): this aligns with the third annotated Fe
ligand, a His, in all six active comparators. B6K4V0, A0AAF0AZ59, S9R0Y2
and S9XEB0 carry Tyr at this column.
- Those four are unreviewed (TrEMBL) entries. Their UniProt Fe-ligand
annotation at that Tyr comes from the PROSITE ProRule PRU00538, which
labels a position as a ligand by alignment. It is not evidence of iron
binding.
- K314 (BINDING, "substrate"): the most common residue at this column is
Lys (fraction 0.72).
- Y307 (MUTAGEN): this column is variable (most common Y, fraction 0.39).
3. Composition and structure (results/functional_regions_analysis.txt, results/structural_analysis.txt)
- Histidines: the Epe1 JmjC domain has 4, against 5-9 in the comparators'
domains.
- C-terminal composition: in the last 100 residues it is 28% hydrophobic,
25% basic and 15% acidic, with no PxVxL motif. Composition is not evidence of
a binding site. Raiymbek et al. 2020 map the minimal Swi6 interaction site
to residues 434-600.
- Secondary structure: the propensity estimate for 243-402 is 15.0% helix
and 42.5% strand. It is crude, and no conclusion about fold or activity is
drawn from it.
Corrections to earlier versions of this analysis
- Wrong JmjC boundaries. Earlier scripts placed the JmjC domain at 200-350
or 400-600, from the first H.[DE] scan hit or a hardcoded guess.
UniProt annotates 243-402, which is what every step now uses.
- Retracted conclusions. Earlier outputs read "HVD motif at position 280"
as the iron site. They concluded "Non-catalytic JmjC domain", "Structure
retained but catalytic activity lost" and "lacks canonical HXD/HXE motifs,
indicating loss of demethylase catalytic activity". The last came from a
code path that always fired. These statements were hardcoded or drawn from
the motif scan, and they are withdrawn. No script now writes an activity
verdict.
- Mislabelled comparators. Four of the five comparator accessions were
mislabelled: P84027 (a spider toxin, not KDM4A), Q6ZMT4 (KDM7A, not KDM5C),
Q92833 (JARID2, a catalytically inactive JmjC protein, not KDM3A) and
P41229 (KDM5C, not KDM5B). The table above uses verified accessions, and
01_fetch_sequences.py now refuses to save a record whose UniProt entry name
disagrees with its label.
- Meaningless conservation score. The earlier conservation score was
computed on unaligned sequence windows. It is replaced by the MAFFT
alignment above.
- Legacy scripts removed.
analyze_epe1.py and analyze_jmjc_protein.py
hardcoded the 400-600 boundaries, reported 0-based positions (for example
"279" for H280) and drew an activity verdict from the motif scan. Their
outputs (epe1_refactored.json, and kdm5a_test.json, which was run on
Q9Y2K7, i.e. KDM2A) have been removed. They remain in git history.
Limitations
- UniProt Fe-ligand annotations are curated for reviewed entries but
rule-based (PRU00538) for unreviewed ones.
- The alignment is a single MAFFT run on JmjC domains only. Regions with
large insertions (KDM3A) align poorly.
- The UniProt homolog search is run live, so the homolog set can change
between runs.
- Sequence analysis cannot establish metal binding, cofactor binding or
catalysis.
References
- UniProt O94603 (
../Epe1-uniprot.txt) and the comparator entries listed above
- Raiymbek et al. 2020 (PMID:32195666): Swi6 interaction site 434-600