ADNP2: did the ADNP NAP-peptide annotation defect propagate to the paralogue?

Generated by analyze_adnp2_propagation.py. Every number below is computed at run time from UniProt, QuickGO, PubMed and the two committed GOA tables; nothing is hardcoded except the merged ADNP review's published PxVxL figures, which are asserted as a precondition.

A. The NAP octapeptide is absent from ADNP2

accession entry length NAPVSIPQ any NAP tripeptide
Q6IQ32 ADNP2_HUMAN (human ADNP2 (subject)) 1131 none none
Q8CHC8 ADNP2_MOUSE (mouse Adnp2 (the Compara/ISS donor for ADNP2)) 1165 none none
Q9H2P0 ADNP_HUMAN (human ADNP (paralogue)) 1102 [354] [354, 701]
Q9JKL8 ADNP_RAT (rat Adnp (the Compara donor behind ADNP's NAP rows)) 1103 [354] [354, 586, 701]
Q9Z103 ADNP_MOUSE (mouse Adnp) 1108 [354] [354, 585, 700]

Positive controls ['Q9H2P0', 'Q9JKL8', 'Q9Z103'] all carry the octapeptide, so a negative for ADNP2 is a result rather than a broken scan. Human and mouse ADNP2 contain no NAPVSIPQ and not one NAP tripeptide in 1131 / 1165 residues. No experiment on the NAP peptide can be an experiment on ADNP2.

B. PxVxL, with its null

Precondition satisfied: this scan reproduces the merged ADNP review's published figures exactly (1 hit at 820-824, null 0.758).

protein length hits expected under composition null
ADNP (Q9H2P0) 1102 PGVLL@820-824 0.758
ADNP2 (Q6IQ32) 1131 PPVLV@662-666, PSVLL@1107-1111 2.705

One occurrence in ~1100 aa is not enrichment for either paralogue: the compositional null is 0.758 for ADNP and 2.705 for ADNP2. What distinguishes ADNP2 is not the count but that PMID:38960717 mutated the motif and lost HP1beta binding.

C. Which of ADNP2's PxVxL candidates is the homologue?

ADNP2 has 2 P-x-V-x-[LMIV] candidates (PPVLV@662-666, PSVLL@1107-1111) against a compositional expectation of 2.705. Presence of the motif is therefore uninformative on its own, and the count cannot be used as evidence.

The discriminating test is projective. Global BLOSUM62 alignment of ADNP vs ADNP2 (27.1% identity) maps ADNP's motif start (820) onto ADNP2 position 1107. Does that land on one of ADNP2's candidates? True ({'match': 'PSVLL', 'start': 1107, 'end': 1111}).

ADNP   IASHFSNKRKKCVRDCEKYKPGVLLGFNMKELNKVKHEMDFDAEW
ADNP2  VASFFGKRRYICMKAIKNHKPSVLLGFDMSELKNVKHRLNFEYEP

The absolute positions differ only because ADNP carries 278 residues after the motif against ADNP2's 20 -- the poorly conserved C-terminal extension described in PMID:38960717. The second ADNP2 candidate has no ADNP counterpart and no experimental support.

D. Which route carried each ADNP2 row, and what did not cross

ADNP2 GOA: 18 rows (18 distinct), 8 terms. ADNP GOA: 53 rows, 33 terms.

Sequence-similarity donors for ADNP2 (GO_REF:0000107 Ensembl Compara and GO_REF:0000024 UniProt ISS): UniProtKB:Q8CHC8
Sequence-similarity donors for ADNP: UniProtKB:Q9JKL8, UniProtKB:Q9Z103

ADNP's rat-Compara block covers 19 terms. Terms from that block that also appear on ADNP2: none.

term name evidence reference with/from
GO:0005634 nucleus IBA GO_REF:0000033 PANTHER:PTN000405125, UniProtKB:Q9H2P0, ZFIN:ZDB-GENE-061215-112
GO:0010468 regulation of gene expression IBA GO_REF:0000033 MGI:MGI:1338758, PANTHER:PTN000405125, RGD:71030
GO:0003677 DNA binding IEA GO_REF:0000002 InterPro:IPR001356
GO:0005634 nucleus IEA GO_REF:0000044 UniProtKB-SubCell:SL-0191
GO:0006357 regulation of transcription by RNA polymerase II IEA GO_REF:0000108 GO:0000981
GO:0005515 protein binding IPI PMID:21888893 UniProtKB:P83916
GO:0005515 protein binding IPI PMID:21888893 UniProtKB:Q13185
GO:0005515 protein binding IPI PMID:24981860 UniProtKB:P83916
GO:0005515 protein binding IPI PMID:27705803 UniProtKB:P83916
GO:0005515 protein binding IPI PMID:27705803 UniProtKB:Q13185
GO:0005515 protein binding IPI PMID:32296183 UniProtKB:Q13185
GO:0005515 protein binding IPI PMID:33961781 UniProtKB:P83916
GO:0005515 protein binding IPI PMID:33961781 UniProtKB:Q13185
GO:0005515 protein binding IPI PMID:35271311 UniProtKB:P83916
GO:0007399 nervous system development IEA GO_REF:0000107 UniProtKB:Q8CHC8, ensembl:ENSMUSP00000068560
GO:0007399 nervous system development ISS GO_REF:0000024 UniProtKB:Q8CHC8
GO:0000785 chromatin ISA GO_REF:0000113 tfclass:3.1.8
GO:0000981 DNA-binding transcription factor activity, RNA polymerase II-specific ISA GO_REF:0000113 tfclass:3.1.8

E. What molecule was assayed behind each IBD seed

4 of 4 seed/term pairs carry their own experimental annotation in the propagated term's subtree.

propagated term seed own experimental evidence reference title names the peptide?
GO:0005634 human ADNP (Q9H2P0) GO:0090575 IDA PMID:29795351 Activity-dependent neuroprotective protein recruits HP1 and CHD4 to control lineage-specifying genes. no
GO:0005634 zebrafish adnpa (F1QLG5) GO:0005634 IDA PMID:32533114 ADNP promotes neural differentiation by modulating Wnt/β-catenin signaling. no
GO:0010468 mouse Adnp (Q9Z103) GO:0000981 IDA PMID:17222401 Activity-dependent neuroprotective protein (ADNP) differentially interacts with chromatin to regulate genes essential for embryogenesis. no
GO:0010468 mouse Adnp (Q9Z103) GO:0006357 IDA PMID:17222401 Activity-dependent neuroprotective protein (ADNP) differentially interacts with chromatin to regulate genes essential for embryogenesis. no
GO:0010468 mouse Adnp (Q9Z103) GO:0010468 IMP PMID:32533114 ADNP promotes neural differentiation by modulating Wnt/β-catenin signaling. no
GO:0010468 rat Adnp (Q9JKL8) GO:0010629 IDA PMID:15314252 NAP mechanisms of neuroprotection. YES

Partial confirmation. 1 seed annotation(s) rest on a reference whose own title names the NAP peptide. This is inside the donor set of an ADNP2 row, but it is not the whole donor set -- see the table above for the co-seed's evidence, which is a gene-product experiment.

F. Logical-opposite citation cross-product

NEGATIVE -- ADNP2 carries no logically opposed term pair, so the cross-product defect found on ADIPOQ cannot occur here. (opposed pairs found: 0)

G. TFClass node reach, and the import's own exclusion set

tfclass:3.1.8 reaches 14 human gene products (28 annotations). Positive control: the subject (UniProtKB:Q6IQ32) and its paralogue (UniProtKB:Q9H2P0) are both present, so this is a census of the right set. Per-entity term signatures: {'GO:0000785,GO:0000981': 14}.

gene product symbol terms received
UniProtKB:Q9H2P0 ADNP GO:0000785, GO:0000981
UniProtKB:Q6IQ32 ADNP2 GO:0000785, GO:0000981
UniProtKB:Q8IX15 HOMEZ GO:0000785, GO:0000981
UniProtKB:Q6ZSZ6 TSHZ1 GO:0000785, GO:0000981
UniProtKB:Q9NRE2 TSHZ2 GO:0000785, GO:0000981
UniProtKB:Q63HK5 TSHZ3 GO:0000785, GO:0000981
UniProtKB:P37275 ZEB1 GO:0000785, GO:0000981
UniProtKB:O60315 ZEB2 GO:0000785, GO:0000981
UniProtKB:Q9C0A1 ZFHX2 GO:0000785, GO:0000981
UniProtKB:Q15911 ZFHX3 GO:0000785, GO:0000981
UniProtKB:Q86UP3 ZFHX4 GO:0000785, GO:0000981
UniProtKB:Q9UKY1 ZHX1 GO:0000785, GO:0000981
UniProtKB:Q9Y6X8 ZHX2 GO:0000785, GO:0000981
UniProtKB:Q9H4I2 ZHX3 GO:0000785, GO:0000981

So every gene the node reaches receives the identical pair, which means GO:0000981 on ADNP2 is a property of class membership rather than a judgement about ADNP2.

Widening to the whole import: GO_REF:0000113 = 1436 annotations over 727 distinct entities, evidence codes {'ISA': 1436}. Of those entities, 709 receive GO:0000785+GO:0000981, 18 receive GO:0000785 alone, and 0 carry some other signature. Entity counts are derived as a distinct set of gene-product ids, not from the annotation total.

The 18 chromatin-only entities are the import's own negative control — the pipeline already withholds the molecular-function term where it does not apply. Listed as a set rather than a count, because this is the payload of the ask to NTNU_SB and a curator needs something diffable:

gene product symbol
UniProtKB:O15105 SMAD7
UniProtKB:O43541 SMAD6
UniProtKB:P51843 NR0B1
UniProtKB:P61129 ZC3H6
UniProtKB:Q12986 NFX1
UniProtKB:Q15466 NR0B2
UniProtKB:Q15596 NCOA2
UniProtKB:Q15788 NCOA1
UniProtKB:Q5H9I0 TFDP3
UniProtKB:Q5HYR2 DMRTC1
UniProtKB:Q6NT76 HMBOX1
UniProtKB:Q6ZN18 AEBP2
UniProtKB:Q6ZNB6 NFXL1
UniProtKB:Q8IX07 ZFPM1
UniProtKB:Q8N5P1 ZC3H8
UniProtKB:Q8WW38 ZFPM2
UniProtKB:Q9BPY8 HOPX
UniProtKB:Q9Y6Q9 NCOA3

Which excluded entities carry a DNA-binding domain? Rendering the DOMAIN note alongside the span, because the notes are not uniform and the difference matters. Of the 18 excluded entities, 3 carry an annotated DNA_BIND feature at all:

gene product symbol DNA_BIND non-degenerate? note identical to ADNP2's?
UniProtKB:Q5H9I0 TFDP3 108-190 yes no
UniProtKB:Q6NT76 HMBOX1 267-341 (Homeobox) yes yes
UniProtKB:Q9BPY8 HOPX 3-62 (Homeobox; degenerate) no no

ADNP2's own feature is 1043-1102 (Homeobox). So the fold-symmetry precedent is HMBOX1 — annotated with the identical note — and not HOPX, whose domain UniProt calls degenerate (3-62 (Homeobox; degenerate)). That asymmetry is stated here rather than left in the JSON for a reader to find, because it is the first thing a curator would raise: if HOPX were the only precedent, the reply would be that HOPX is excluded because its homeodomain is broken, which would not transfer to ADNP2's intact one. The script asserts that at least one excluded entity has a non-degenerate domain and refuses to report if none does.

HOPX still carries the biological half of the precedent — UniProt describes it as an atypical homeodomain protein that does not bind DNA, and it is in the exclusion set (asserted, not assumed: the script fails if it is not). What it does not carry is fold symmetry with ADNP2.

Is the exclusion a per-entity judgement or per-node coverage? This decides whether the precedent supports the ask at all: if GO_REF:0000113 could only withhold GO:0000981 for a whole node, then excluding ADNP2 would also exclude ADNP — whose sequence-specific binding is measured — and the correct request would be a different one. Answer, from the data: per-entity, demonstrated = True (an existential: it holds if any node retains the term for some members while withholding it from a strict subset).

The import uses both granularities, and that distinction has to be stated rather than filtered away. The excluded entities are spread across 11 nodes, of which 7 withhold the term from a strict subset (10 entities) while 4 withhold it from every member (8 entities). All 11 are printed:

node members keep the term excluded kind
tfclass:3.1.3 47 46 HOPX strict subset
tfclass:1.2.5 22 19 NCOA1, NCOA2, NCOA3 strict subset
tfclass:2.3.2 20 19 AEBP2 strict subset
tfclass:3.1.10 19 18 HMBOX1 strict subset
tfclass:3.3.2 11 10 TFDP3 strict subset
tfclass:7.1.1 8 6 SMAD6, SMAD7 strict subset
tfclass:2.5.1 8 7 DMRTC1 strict subset
tfclass:0.4.1 2 0 NFX1, NFXL1 whole node
tfclass:2.1.7 2 0 NR0B1, NR0B2 whole node
tfclass:2.7.2 2 0 ZFPM1, ZFPM2 whole node
tfclass:2.8.1 2 0 ZC3H6, ZC3H8 whole node

2 of the single-exclusion strict-subset nodes sit in the same TFClass class as ADNP2: HOPX excluded alone out of 47 members of tfclass:3.1.3, and HMBOX1 excluded alone out of 19 members of tfclass:3.1.10 — while ADNP2's own node tfclass:3.1.8 currently has 14/14 members holding the term. So single-entity exclusion inside a populated homeodomain node is something this import already performs, twice within class 3.1, and the request needs no new mechanism and would not touch ADNP: it takes one node from 14/14 to 13/14. The wholly-excluded nodes are not the precedent ADNP2 needs — but they do show the import has the coarser granularity too, so naming which one the ask relies on matters.

That also settles what these 18 entities do not have in common. They are biologically heterogeneous: NFX1 "Binds to the X-box motif of MHC class II genes", which is sequence-specific binding at a cis-regulatory region, and DMRTC1 is named a transcription factor. Three successive drafts generalised the set as "the non-DNA-binding members", then as "none is a sequence-specific polymerase II transcription factor", then as "each is an exclusion inside a node whose other members keep the term"; all three are false, each refuted by a table in this same document. So the statement below is computed from the partition rather than written:

They share no property at all. 10 of the 18 sit in nodes whose other members keep the term, but the remaining 8 sit in 4 nodes where NO member keeps it, so not even the structural description holds across the set. What can be said is only per-node, which is why the table above is the claim and this sentence is not.

And the positive argument for ADNP2 is neither of those. It is the measured failure to find a motif: no sequence motif explains ADNP2's ChIP-seq distribution, its peaks avoid transcription start sites, and PxVxL mutation nearly abolishes its chromatin binding. The exclusion set shows only that this import has a mechanism for withholding GO:0000981 and already applies it to intact-domain proteins; it does not itself argue that ADNP2 belongs there.

Note what does not depend on any of this: the GO:0000981 verdict rests on the quoted three-clause failure against GO:0003700's definition. This section supports the upstream ask, not the annotation action.