Generated by analyze_isoform_transfer.py. Every figure is read from the UniProt and Ensembl REST APIs (raw responses cached under data/); nothing is hard-coded.
Human ACRBP carries GO:0001675 acrosome assembly by ISS (GO_REF:0000024) and by Ensembl Compara IEA (GO_REF:0000107), both from mouse Acrbp Q3V140. The mouse source annotation is an IMP on PMID:27303034, where the acrosomal-granule defect of the knockout is rescued by transgenic ACRBP-V5 alone, and ACRBP-V5 is an intron-5-retaining splice variant reported in mouse but not in other mammalian ACRBPs (PMID:23426433). The obvious objection to the transfer is therefore that the granule-forming activity has no vehicle in human. This analysis tests the sequence-level half of that objection and finds it does not hold: the human locus carries the same intron-5 read-through transcript. It does not test the expression-level half, which PMID:30606959 reports directly and against the human transcript - see the limits at the end of Part 2b.
Reviewed members of PANTHER PTHR21362. Isoforms counts UniProt ALTERNATIVE PRODUCTS entries.
| Accession | Species | Length (aa) | Isoforms | Names | Alternative-sequence features |
|---|---|---|---|---|---|
| Q8NEB7 | human | 543 | 0 | - | 0 |
| Q3V140 | mouse | 540 | 2 | 1 (ACRBP-W); 2 (ACRBP-V5) | 2 |
| Q6AY33 | rat | 540 | 2 | 1; 2 | 2 |
| Q60485 | guinea pig | 543 | 0 | - | 0 |
| Q29016 | pig | 539 | 0 | - | 0 |
Read-out. Multi-isoform: mouse, rat. Single-isoform: human, guinea pig, pig. UniProt annotates the alternative product on the two murid entries and on neither the human, the pig, nor - notably - the guinea pig, which is itself a rodent, so at UniProt level the variant reads as murid rather than pan-rodent. But an absent ALTERNATIVE PRODUCTS section records what UniProt has curated, not what the genome encodes, so Part 2 goes to the genome annotation instead.
For every protein_coding transcript at the locus, the translated product is compared with that species' canonical product. A transcript is scored V5-like when it is non-canonical, is 0.40-0.80 of the canonical length, and matches the canonical N-terminus at >=95% identity over its own length. Mouse ACRBP-V5 (316/540 = 0.585 of canonical) is the positive control this test must recover.
| Species | Gene | Assembly | Transcripts | Biotypes | Canonical protein (aa) |
|---|---|---|---|---|---|
| human | ENSG00000111644 | GRCh38 | 13 | nonsense_mediated_decay=1, protein_coding=8, protein_coding_CDS_not_defined=2, retained_intron=2 | 543 |
| mouse | ENSMUSG00000072770 | GRCm39 | 7 | protein_coding=4, protein_coding_CDS_not_defined=1, retained_intron=2 | 540 |
| rat | ENSRNOG00000017399 | GRCr8 | 4 | protein_coding=4 | 541 |
| Species | Transcript | Protein (aa) | Length ratio | N-terminal identity | Canonical | V5-like |
|---|---|---|---|---|---|---|
| human | ACRBP-201 | 543 | 1.0 | 100.0% | yes | no |
| human | ACRBP-202 | 510 | 0.939 | 34.1% | no | no |
| human | ACRBP-204 | 319 | 0.587 | 98.4% | no | yes |
| human | ACRBP-209 | 173 | 0.319 | 91.9% | no | no |
| human | ACRBP-210 | 521 | 0.959 | 96.7% | no | no |
| human | ACRBP-211 | 151 | 0.278 | 79.5% | no | no |
| human | ACRBP-212 | 392 | 0.722 | 92.6% | no | no |
| human | ACRBP-213 | 543 | 1.0 | 100.0% | no | no |
| mouse | Acrbp-201 | 167 | 0.309 | 72.5% | no | no |
| mouse | Acrbp-202 | 540 | 1.0 | 100.0% | yes | no |
| mouse | Acrbp-204 | 170 | 0.315 | 92.9% | no | no |
| mouse | Acrbp-205 | 316 | 0.585 | 98.7% | no | yes |
| rat | Acrbp-201 | 541 | 1.0 | 100.0% | yes | no |
| rat | Acrbp-202 | 167 | 0.309 | 71.9% | no | no |
| rat | Acrbp-203 | 500 | 0.924 | 8.6% | no | no |
| rat | Acrbp-204 | 489 | 0.904 | 9.8% | no | no |
Read-out. V5-like transcripts found in: human (ACRBP-204); mouse (Acrbp-205). None found in: rat. The test recovers the mouse positive control - and it also returns a hit in human. That is the opposite of what the rodent-specific reading of PMID:23426433 predicts, so Part 2b checks whether the human hit is the same splicing event or an unrelated truncation.
Mouse ACRBP-V5 arises by retention of intron 5: the transcript keeps the first five exons but runs past the exon-5 donor site into intron 5, which supplies an in-frame stop. UniProt records the consequence on the protein as SLQQL -> RYRKL immediately before the truncation. Each V5-like transcript is therefore checked for both signatures - the exon chain and the substituted C-terminal pentapeptide. In the leading exons shared column, exon 1 is matched on its donor boundary only, because annotated transcription start sites differ by a few bases between transcripts of the same gene; later exons must match on both boundaries.
| Species | Transcript | Exons (variant/canonical) | Leading exons shared | Terminal-exon verdict | Protein (aa) | Variant tail | Canonical at same positions |
|---|---|---|---|---|---|---|---|
| human | ACRBP-204 (ENST00000536350) | 5/10 | 4 | terminal exon = canonical exon 5 extended 82 bp into intron 5 | 319 | RYRKF |
SLLQL |
| mouse | Acrbp-205 (ENSMUST00000112414) | 5/10 | 4 | terminal exon = canonical exon 5 extended 242 bp into intron 5 | 316 | RYRKL |
SLQQL |
Read-out. Intron read-through confirmed in: human (ACRBP-204), mouse (Acrbp-205). Read-through occurs after canonical exon 5 in every case. Variant C-terminal pentapeptides: RYRKF, RYRKL (shared prefix RYRK), replacing canonical SLLQL, SLQQL. The human and mouse variants read through the same intron, terminate at equivalent positions, and both replace the canonical S-L-x-Q-L pentapeptide with RYRK-initiated basic tails that differ only in the final residue. That is one conserved splicing event, not two coincidences: the human ACRBP locus does carry a structural counterpart of the 'rodent-specific' ACRBP-V5. Whether that counterpart is transcribed in human spermatogenic cells is a separate question, and the limits below record what the primary literature says about it.
Three limits on this result, and the second is the important one. It is an annotation-level finding: GENCODE annotating the transcript is not evidence that the protein is made in human spermatids. The primary literature reports the opposite at the mRNA level, and states it directly rather than by implication - PMID:30606959: "Porcine, guinea pig, and human spermatogenic cells produce only a single form of Acrbp (termed Acrbp-W) mRNA, whereas two mRNA forms, wild-type Acrbp-W and intron 5-retaining variant Acrbp-V5 mRNAs, are synthesized by pre-mRNA alternative splicing of the Acrbp gene in mouse." RT-PCR outranks a genome annotation on the question of what is transcribed, so what this section establishes is a discrepancy between GENCODE and the experimental record, not that human spermatids make an ACRBP-V5 equivalent. And UniProt Q8NEB7 carries no ALTERNATIVE PRODUCTS section (Part 1), so the human product of this transcript has no isoform identifier that a GO annotation could be qualified with.
Mouse ACRBP-V5 (Q3V140-2) is 316 aa: UniProt records residues 317-540 of the 540-aa canonical ACRBP-W as absent from it, i.e. V5 stops in the middle of the protein and keeps only the N-terminal half.
| Quantity | Value |
|---|---|
| Human Q8NEB7 length | 543 aa |
| Mouse Q3V140-1 (ACRBP-W) length | 540 aa |
| Mouse Q3V140-2 (ACRBP-V5) length | 316 aa |
| Global identity, human vs mouse ACRBP-W | 76.1% |
| Human residue aligned to mouse residue 316 (last V5 residue) | 319 |
| Human propeptide (removed on maturation) | 26-273 |
| Human mature chain | 274-543 |
Identity across the two regions UniProt annotates on human ACRBP as pro-ACR binding:
| Human region | Note | Identity vs mouse ACRBP-W | Inside ACRBP-V5 span? | Inside human propeptide? |
|---|---|---|---|---|
| 26-106 | Pro-ACR binding | 93.8% | yes | yes |
| 319-427 | Pro-ACR binding | 73.4% | no | no |
Read-out. Human and mouse ACRBP are 76.1% identical overall, so the orthology behind the ISS transfer is not in doubt. The ACRBP-V5 truncation point maps to human residue 319, i.e. 46 residues past the end of the 26-273 propeptide that human ACRBP removes during maturation. ACRBP-V5 therefore corresponds to the whole of that propeptide plus the first 46 residues of the mature chain - essentially the half of the precursor that maturation discards. Of the two annotated pro-ACR-binding regions, 1 falls inside the V5 span (26-106) and lies within the propeptide, while 1 (319-427) sits in the mature chain that human ACRBP keeps. So the two mouse activities are carried on physically separate halves of the protein, and human ACRBP keeps both halves: the mature chain permanently, and the V5-equivalent N-terminal half either as the transient unprocessed 60-kDa precursor or as the product of the intron-5 read-through transcript found in Part 2b.
Bears on GO:0002080 acrosomal membrane (human ISS from pig Q29016, whose source is an immunofluorescence IDA that cannot resolve acrosomal matrix from acrosomal membrane).
| Accession | Species | Signal peptide | Transmembrane/lipidation features | Max Kyte-Doolittle 19-mer after signal | Window start |
|---|---|---|---|---|---|
| Q8NEB7 | human | 1-25 | none | 1.44 | 299 |
| Q3V140 | mouse | 1-25 | none | 0.81 | 297 |
| Q6AY33 | rat | 1-24 | none | 0.81 | 298 |
| Q60485 | guinea pig | 1-25 | none | 1.56 | 299 |
| Q29016 | pig | 1-25 | none | 1.58 | 293 |
Read-out. Orthologs carrying a transmembrane or lipidation feature: none. Orthologs lacking a signal peptide: none. The most hydrophobic mature window in the family is pig at 1.58, below the ~1.6 heuristic threshold, though not by a wide margin. The load-bearing evidence is the curated feature table, not the hydropathy scan: every ACRBP ortholog is a signal-peptide-bearing, anchor-free secretory protein. That is consistent with sp32 having been purified from acid extracts of ejaculated sperm as a soluble protein, and it makes a soluble lumenal/matrix assignment (GO:0043159 acrosomal matrix) the appropriate compartment. located_in GO:0002080 acrosomal membrane asserts residence in the bilayer itself, which no ortholog has a feature to support; GO:0005634 nucleus is likewise unreachable for a protein that is translocated into the secretory pathway at synthesis.