Does human ACRBP have a counterpart of the 'rodent-specific' ACRBP-V5 isoform?

Generated by analyze_isoform_transfer.py. Every figure is read from the UniProt and Ensembl REST APIs (raw responses cached under data/); nothing is hard-coded.

Question

Human ACRBP carries GO:0001675 acrosome assembly by ISS (GO_REF:0000024) and by Ensembl Compara IEA (GO_REF:0000107), both from mouse Acrbp Q3V140. The mouse source annotation is an IMP on PMID:27303034, where the acrosomal-granule defect of the knockout is rescued by transgenic ACRBP-V5 alone, and ACRBP-V5 is an intron-5-retaining splice variant reported in mouse but not in other mammalian ACRBPs (PMID:23426433). The obvious objection to the transfer is therefore that the granule-forming activity has no vehicle in human. This analysis tests the sequence-level half of that objection and finds it does not hold: the human locus carries the same intron-5 read-through transcript. It does not test the expression-level half, which PMID:30606959 reports directly and against the human transcript - see the limits at the end of Part 2b.

Part 1 - annotated isoforms across the reviewed ACRBP family

Reviewed members of PANTHER PTHR21362. Isoforms counts UniProt ALTERNATIVE PRODUCTS entries.

Accession Species Length (aa) Isoforms Names Alternative-sequence features
Q8NEB7 human 543 0 - 0
Q3V140 mouse 540 2 1 (ACRBP-W); 2 (ACRBP-V5) 2
Q6AY33 rat 540 2 1; 2 2
Q60485 guinea pig 543 0 - 0
Q29016 pig 539 0 - 0

Read-out. Multi-isoform: mouse, rat. Single-isoform: human, guinea pig, pig. UniProt annotates the alternative product on the two murid entries and on neither the human, the pig, nor - notably - the guinea pig, which is itself a rodent, so at UniProt level the variant reads as murid rather than pan-rodent. But an absent ALTERNATIVE PRODUCTS section records what UniProt has curated, not what the genome encodes, so Part 2 goes to the genome annotation instead.

Part 2 - does any human transcript encode an ACRBP-V5-like N-terminal half?

For every protein_coding transcript at the locus, the translated product is compared with that species' canonical product. A transcript is scored V5-like when it is non-canonical, is 0.40-0.80 of the canonical length, and matches the canonical N-terminus at >=95% identity over its own length. Mouse ACRBP-V5 (316/540 = 0.585 of canonical) is the positive control this test must recover.

Species Gene Assembly Transcripts Biotypes Canonical protein (aa)
human ENSG00000111644 GRCh38 13 nonsense_mediated_decay=1, protein_coding=8, protein_coding_CDS_not_defined=2, retained_intron=2 543
mouse ENSMUSG00000072770 GRCm39 7 protein_coding=4, protein_coding_CDS_not_defined=1, retained_intron=2 540
rat ENSRNOG00000017399 GRCr8 4 protein_coding=4 541
Species Transcript Protein (aa) Length ratio N-terminal identity Canonical V5-like
human ACRBP-201 543 1.0 100.0% yes no
human ACRBP-202 510 0.939 34.1% no no
human ACRBP-204 319 0.587 98.4% no yes
human ACRBP-209 173 0.319 91.9% no no
human ACRBP-210 521 0.959 96.7% no no
human ACRBP-211 151 0.278 79.5% no no
human ACRBP-212 392 0.722 92.6% no no
human ACRBP-213 543 1.0 100.0% no no
mouse Acrbp-201 167 0.309 72.5% no no
mouse Acrbp-202 540 1.0 100.0% yes no
mouse Acrbp-204 170 0.315 92.9% no no
mouse Acrbp-205 316 0.585 98.7% no yes
rat Acrbp-201 541 1.0 100.0% yes no
rat Acrbp-202 167 0.309 71.9% no no
rat Acrbp-203 500 0.924 8.6% no no
rat Acrbp-204 489 0.904 9.8% no no

Read-out. V5-like transcripts found in: human (ACRBP-204); mouse (Acrbp-205). None found in: rat. The test recovers the mouse positive control - and it also returns a hit in human. That is the opposite of what the rodent-specific reading of PMID:23426433 predicts, so Part 2b checks whether the human hit is the same splicing event or an unrelated truncation.

Part 2b - is the human hit the same intron-5 read-through event?

Mouse ACRBP-V5 arises by retention of intron 5: the transcript keeps the first five exons but runs past the exon-5 donor site into intron 5, which supplies an in-frame stop. UniProt records the consequence on the protein as SLQQL -> RYRKL immediately before the truncation. Each V5-like transcript is therefore checked for both signatures - the exon chain and the substituted C-terminal pentapeptide. In the leading exons shared column, exon 1 is matched on its donor boundary only, because annotated transcription start sites differ by a few bases between transcripts of the same gene; later exons must match on both boundaries.

Species Transcript Exons (variant/canonical) Leading exons shared Terminal-exon verdict Protein (aa) Variant tail Canonical at same positions
human ACRBP-204 (ENST00000536350) 5/10 4 terminal exon = canonical exon 5 extended 82 bp into intron 5 319 RYRKF SLLQL
mouse Acrbp-205 (ENSMUST00000112414) 5/10 4 terminal exon = canonical exon 5 extended 242 bp into intron 5 316 RYRKL SLQQL

Read-out. Intron read-through confirmed in: human (ACRBP-204), mouse (Acrbp-205). Read-through occurs after canonical exon 5 in every case. Variant C-terminal pentapeptides: RYRKF, RYRKL (shared prefix RYRK), replacing canonical SLLQL, SLQQL. The human and mouse variants read through the same intron, terminate at equivalent positions, and both replace the canonical S-L-x-Q-L pentapeptide with RYRK-initiated basic tails that differ only in the final residue. That is one conserved splicing event, not two coincidences: the human ACRBP locus does carry a structural counterpart of the 'rodent-specific' ACRBP-V5. Whether that counterpart is transcribed in human spermatogenic cells is a separate question, and the limits below record what the primary literature says about it.

Three limits on this result, and the second is the important one. It is an annotation-level finding: GENCODE annotating the transcript is not evidence that the protein is made in human spermatids. The primary literature reports the opposite at the mRNA level, and states it directly rather than by implication - PMID:30606959: "Porcine, guinea pig, and human spermatogenic cells produce only a single form of Acrbp (termed Acrbp-W) mRNA, whereas two mRNA forms, wild-type Acrbp-W and intron 5-retaining variant Acrbp-V5 mRNAs, are synthesized by pre-mRNA alternative splicing of the Acrbp gene in mouse." RT-PCR outranks a genome annotation on the question of what is transcribed, so what this section establishes is a discrepancy between GENCODE and the experimental record, not that human spermatids make an ACRBP-V5 equivalent. And UniProt Q8NEB7 carries no ALTERNATIVE PRODUCTS section (Part 1), so the human product of this transcript has no isoform identifier that a GO annotation could be qualified with.

Part 3 - where the ACRBP-V5 truncation lands on the human protein

Mouse ACRBP-V5 (Q3V140-2) is 316 aa: UniProt records residues 317-540 of the 540-aa canonical ACRBP-W as absent from it, i.e. V5 stops in the middle of the protein and keeps only the N-terminal half.

Quantity Value
Human Q8NEB7 length 543 aa
Mouse Q3V140-1 (ACRBP-W) length 540 aa
Mouse Q3V140-2 (ACRBP-V5) length 316 aa
Global identity, human vs mouse ACRBP-W 76.1%
Human residue aligned to mouse residue 316 (last V5 residue) 319
Human propeptide (removed on maturation) 26-273
Human mature chain 274-543

Identity across the two regions UniProt annotates on human ACRBP as pro-ACR binding:

Human region Note Identity vs mouse ACRBP-W Inside ACRBP-V5 span? Inside human propeptide?
26-106 Pro-ACR binding 93.8% yes yes
319-427 Pro-ACR binding 73.4% no no

Read-out. Human and mouse ACRBP are 76.1% identical overall, so the orthology behind the ISS transfer is not in doubt. The ACRBP-V5 truncation point maps to human residue 319, i.e. 46 residues past the end of the 26-273 propeptide that human ACRBP removes during maturation. ACRBP-V5 therefore corresponds to the whole of that propeptide plus the first 46 residues of the mature chain - essentially the half of the precursor that maturation discards. Of the two annotated pro-ACR-binding regions, 1 falls inside the V5 span (26-106) and lies within the propeptide, while 1 (319-427) sits in the mature chain that human ACRBP keeps. So the two mouse activities are carried on physically separate halves of the protein, and human ACRBP keeps both halves: the mature chain permanently, and the V5-equivalent N-terminal half either as the transient unprocessed 60-kDa precursor or as the product of the intron-5 read-through transcript found in Part 2b.

Part 4 - is any ACRBP ortholog membrane-anchored?

Bears on GO:0002080 acrosomal membrane (human ISS from pig Q29016, whose source is an immunofluorescence IDA that cannot resolve acrosomal matrix from acrosomal membrane).

Accession Species Signal peptide Transmembrane/lipidation features Max Kyte-Doolittle 19-mer after signal Window start
Q8NEB7 human 1-25 none 1.44 299
Q3V140 mouse 1-25 none 0.81 297
Q6AY33 rat 1-24 none 0.81 298
Q60485 guinea pig 1-25 none 1.56 299
Q29016 pig 1-25 none 1.58 293

Read-out. Orthologs carrying a transmembrane or lipidation feature: none. Orthologs lacking a signal peptide: none. The most hydrophobic mature window in the family is pig at 1.58, below the ~1.6 heuristic threshold, though not by a wide margin. The load-bearing evidence is the curated feature table, not the hydropathy scan: every ACRBP ortholog is a signal-peptide-bearing, anchor-free secretory protein. That is consistent with sp32 having been purified from acid extracts of ejaculated sperm as a soluble protein, and it makes a soluble lumenal/matrix assignment (GO:0043159 acrosomal matrix) the appropriate compartment. located_in GO:0002080 acrosomal membrane asserts residence in the bilayer itself, which no ortholog has a feature to support; GO:0005634 nucleus is likewise unreachable for a protein that is translocated into the secretory pathway at synthesis.