ADAMTSL1 (Q8N6G6, punctin-1) — review notes

Human ADAMTSL1, PAINT + affinage campaign. GOA record is unusually small: 4 rows
(1 IEA + 3 TAS), no IBA, no IPI, no MF annotation of any kind.

1. Domain content: the "-like" really does mean non-catalytic

Established from the UniProt feature table before reading any narrative source, because the
campaign has twice shipped an inverted premise about catalysis.

The FT block of ADAMTSL1-uniprot.txt lists, over the 1762-aa precursor: a signal peptide
(1–28), nine annotated TSP type-1 repeats, four Ig-like C2-type domains, and a C-terminal
PLAC domain. There is no metalloprotease domain, no disintegrin-like domain, no
prodomain, and no zinc-binding site feature
. There is no CATALYTIC ACTIVITY comment and
no EC number. UniProt states it outright
[file:human/ADAMTSL1/ADAMTSL1-uniprot.txt "lacks the metalloprotease and disintegrin-like domains which are"],
and the primary literature says the same in two places
PMID:11805097
and
PMID:28722276.

So the lead in the task brief holds for this gene, and it is established from the entry's own
feature table rather than from the family name.

The ADAMTS spacer region named in the review description is not an FT DOMAIN feature;
it comes from DR Pfam; PF05986; ADAMTS_spacer1; 1. and DR InterPro; IPR010294; ADAMTS_spacer1.. Sourced in the review's UniProt reference rather than left as an unattributed
claim.

Pfam/PROSITE counts differ from the FT list (PROSITE TSP1 9, SMART TSP1 13, Pfam
TSP1_ADAMTS 11) — the field's usual count for the full-length protein is thirteen TSRs
PMID:28722276.
Nothing in the review depends on the exact repeat count.

2. The peptidase error IS present — but in UniProt's keywords, not in GOA

Q8N6G6 carries KW-0378 Hydrolase (Molecular function category), which generates
[file:human/ADAMTSL1/ADAMTSL1-uniprot.txt "GO:0016787; F:hydrolase activity; IEA:UniProtKB-KW."]
in the entry's own GO cross-reference list
[file:human/ADAMTSL1/ADAMTSL1-uniprot.txt "Extracellular matrix; Glycoprotein; Hydrolase; Immunoglobulin domain;"].

This is contradicted inside the same entry by the CAUTION comment quoted above, and there
is no reaction anywhere in the record to support it.

It is also anomalous within the family. Of the six human ADAMTS-like proteins, only ADAMTSL1
and THSD4 carry the Hydrolase keyword; ADAMTSL2, ADAMTSL3, ADAMTSL4 and ADAMTSL5 carry the
identical CAUTION about the missing metalloprotease domain and no MF keyword. None of the
six has a CATALYTIC ACTIVITY comment. Computed in
[file:human/ADAMTSL1/ADAMTSL1-bioinformatics/RESULTS.md "Entries with a CATALYTIC ACTIVITY comment: none."].

The important curation fact: this never reached GOA. GO:0016787 is absent from
ADAMTSL1-goa.tsv and from QuickGO, because keyword-derived annotations (GO_REF:0000043)
were withdrawn for cellular organisms. So the defect is live in UniProt and invisible in GO.
It is reported in suggested_questions as a UniProt correction to file, not as an
existing_annotations row, because reviewing rows that GOA does not carry would break the
row-count reconciliation.

Two other places the same conflation shows up, worth separating because they have different
standing:

InterPro2GO got this right. IPR013273 ADAMTS/ADAMTS-like is the only one of ADAMTSL1's
fourteen InterPro entries with a GO mapping, and it maps to GO:0030198 alone — no peptidase
term — even though the signature spans the catalytic ADAMTS proteases. That is the correct
treatment of a signature covering a mechanistically heterogeneous family.

3. PAINT knows catalysis was lost on this branch, and says so on one gene out of six

The cached PAINT table for the family contains two negated rows:

PTHR13723 PTN002673039 GO:0004222 F IKR true PANTHER:PTN000347317
PTHR13723 PTN002673039 GO:0006508 P IRD true PANTHER:PTN000347317

IKR (inferred from key residues) and IRD (inferred from rapid divergence) at node
PTN002673039 block metalloendopeptidase activity and proteolysis from propagating below it.
In GOA that loss surfaces as exactly one annotation — ADAMTSL2's NOT|enables GO:0004222
IBA — and on no other ADAMTS-like member. ADAMTSL1 gets neither the positive term (good) nor
the explicit negation (a missed opportunity: the NOT is the machine-readable form of the
CAUTION comment UniProt already writes).

4. The headline finding: ADAMTSL1 is the only human family member with no GO:0031012

PAINT holds GO:0031012 extracellular matrix as an IBD at node PTN000347317. Census over
all 26 human members of PTHR13723:

And the gap is species-specific, not subfamily-specific: mouse Adamtsl1 (Q8BLI0), same
PANTHER subfamily SF157, does receive the IBA from PTN000347317
, plus three HDA rows from
matrisome proteomics. Full table in
[file:human/ADAMTSL1/ADAMTSL1-bioinformatics/RESULTS.md "Members with no GO:0031012 annotation of any kind: ADAMTSL1."].

UniProt records the human localisation as experimentally supported by three papers
(ECO:0000269|PubMed:11805097, PubMed:17395588, PubMed:19671700), and independent human
evidence exists:

So the one gene in the family whose ECM residence is directly shown in human cells is the
one gene in the family with no ECM annotation. Proposed as a NEW row (located_in GO:0031012, IDA, PMID:11805097) rather than left for PAINT, because the human primary
evidence stands on its own.

5. GO:0030198 extracellular matrix organization (the one IEA)

KEEP_AS_NON_CORE, not ACCEPT. The first draft accepted it and rested the acceptance partly
on the annotation being "well-constructed and unrefuted", which the PR reviewer correctly
identified as absence of contradiction rather than positive support. There is positive
support, but all of it is family-level:

Kept on those three grounds, with the gap recorded in knowledge_gaps. Being present in the
ECM and being an MMP10 substrate places ADAMTSL1 in ECM remodelling; neither demonstrates
that it organises the matrix. So the term stays, but it is not asserted as this gene's core
function, and core_functions deliberately carries only locations: GO:0031012 — the one
claim the gene's own data establishes.

Divergence from the concurrent ADAMTSL5 review, and a premise that turned out false

PR #2305 (ADAMTSL5) marks the identical InterPro GO:0030198 IEA
MARK_AS_OVER_ANNOTATED. I was told the two positions could both stand because ADAMTSL1 has
the GO:0030198 IBA and ADAMTSL5 does not. That premise is false, and #2305's own table
says so: ADAMTSL1 is – in its GO:0030198 column. My census agrees — the IBA from
PTN000347317 reaches only ADAMTSL2, ADAMTSL4 and THSD4 within the ADAMTS-like branch.
ADAMTSL1 and ADAMTSL5 hold the InterPro IEA and nothing else, i.e. identical evidentiary
positions
, and if anything ADAMTSL5 has more gene-specific matrix evidence (its own IDA
to GO:0031012, plus microfibril and heparin binding). The divergence therefore cannot be
justified per gene; it is about where the family draws the KEEP_AS_NON_CORE /
MARK_AS_OVER_ANNOTATED line, and it is filed once in suggested_questions. The script now
asserts the parity so the comparison cannot go stale.

One asymmetry worth recording, raised by the PR reviewer: for GO:0030198 the IBA misses
4 of 26 human members (ADAMTSL1, ADAMTSL3, ADAMTSL5, PAPLN), against 2 of 26 for
GO:0031012. So "coverage gap" is a weaker reading for GO:0030198 than it is for the
headline GO:0031012 finding, and this review does not lean on it — the claim made about
GO:0030198 is only that the IBD node sits above ADAMTSL1, with the propagation pattern
itself filed as a question rather than argued as a defect.

Where the two reviews agree: ADAMTSL1 has no IBA at all (#2305 flagged this as worth
checking against my branch — no discrepancy, my review reports the absence, never IBA
support), and absence of an IBA at an incoherently-propagating node is a coverage gap rather
than a curatorial judgement. My headline uses the absence exactly that way.

6. The GO:0005788 ER lumen rows (three of them)

All three are true and all three say the same thing. ADAMTSL1's TSRs are O-fucosylated by
POFUT2 and extended by B3GLCT in the ER lumen, and this is required for export
PMID:17395588,
with C-mannosylation acting in the same quality-control step
PMID:19671700.

Reasons for KEEP_AS_NON_CORE rather than ACCEPT:

Sibling divergence, declared. The merged ADAMTSL4 review resolved the identical three
rows as ACCEPT, but its own reason reads "This localization reflects its secretory pathway
transit, not a functional localization" — substantively the same judgement, encoded with a
different action. Flagged in suggested_questions for harmonisation rather than silently
diverging.

7. Isoforms — the short form is punctin-1, not the long one

The task brief had this the wrong way round, so stating it plainly: punctin-1 is the SHORT
splice variant
(525 aa, UniProt Q8N6G6-1), and the canonical displayed sequence is the
long isoform 3 (1762 aa, Q8N6G6-3)
PMID:28722276.
UniProt applies the AltName "Punctin-1" at entry level, which is where the confusion comes
from; the literature uses it for the short form.

This matters because essentially every ADAMTSL1 experiment was done on isoform 1:

isoform: Q8N6G6-1 is therefore set on the proposed GO:0031012 row to record what was
tested
, per CLAUDE.md's isoform-tracking convention — not to assert the localisation is
isoform-restricted. It almost certainly is not: isoforms 1–4 all retain the signal peptide.
Isoforms 5 and 6 delete residues 1–1299 including the signal peptide (VSP_039322) and so
would not enter the secretory pathway at all, and isoform 5 is flagged as NMD-prone; that
distinction is recorded in functional_isoforms.

Minor tooling bug found: fetch-gene seeded alternative_products with isoform 5's
sequence_note as VSP_039322, VSP_039329, VSP_039330, — the trailing VSP_039331 was
dropped because the UniProt CC block wraps mid-list. Corrected by hand in the review YAML.

8. Checks run, including the ones that came back negative

8b. The one molecular-function experiment ever done on ADAMTSL1, and it is negative

Found late, and not by me: the concurrent ADAMTSL3 review surfaced it. PMID:22242013 is
titled "Microenvironmental regulation by fibrillin-1", so no ADAMTSL1-keyed search reaches
it; it is absent from ADAMTSL1's GOA and from the affinage record. It is the campaign's
"a paper titled for something else holds your gene's answer" lesson in its partner-named form.

Two direct SPR negatives for human ADAMTSL1, in a panel where its relatives were positive:

This is a measured negative, not an absence, which makes it much stronger than the silence
this review was otherwise working against.

Only the fibrillin-1 half is discriminating, and the PR reviewer was right that my first
write-up read as two independent exclusions. ADAMTSL-2 is equally negative against ADAMTS-10
and still reaches GO:0030198; ADAMTS-10 binding was shown only for ADAMTSL-3; THSD4 was never
tested against it. So the ADAMTS-10 result does not separate ADAMTSL1 from its relatives and
must not be double-counted. The fibrillin-1 negative alone carries the argument, and it is
enough.

The caveat is the isoform question again. The methods say
"Recombinant full length ADAMTSL-1, -2, LTBP-1, -4 ... were covalently coupled to CM5 sensor
chips"
but never give a length, and the Apte lab stated five years later that no construct
for the 1762-residue form existed. Both RefSeqs were available in 2012 (NP_443098 = isoform 1,
525 aa; NP_001035362 = isoform 3, 1762 aa), so the paper alone does not settle it, and the
balance of evidence favours punctin-1. Recorded as a bounded negative: firm for whatever
was assayed, and probably untested for the nine extra TSRs, four Ig-like domains and PLAC
domain of the long form.

What it changed in this review. Ground two of the GO:0030198 reason — "a non-catalytic
route to the term is available to this architecture" — is now explicitly qualified, because
the specific route the relatives use is excluded for this protein. It does not flip the verdict
to MARK_AS_OVER_ANNOTATED: GO:0030198 is far broader than fibrillin-1 binding, and
excluding one mechanism is not refuting the process. But it removes the strongest mechanistic
analogy, and it is a further reason the term cannot be called core. The MF knowledge_gap is
now bounded on both sides — what is established, and what has been measured and excluded —
and the proposed binding experiment became an unbiased partner search rather than the
candidate SPR I had originally proposed, which would have repeated a published negative.

One more correction from the same round, and it is the lesson of this section applied to
itself: I wrote that "the LTBPs remain untested". Unsafe. LTBP-1 and LTBP-4 were coupled to
CM5 chips in these very experiments, and the LTBP results sit in a supplementary table that is
not in the cached record. "Not reported in the main text" is what the evidence supports. This
whole section exists because a candidate that looked untested turned out to have been tested,
so asserting a second one untested on the same paper would have been the same error twice.

And the parity claim needed narrowing. I had written that ADAMTSL1 and ADAMTSL5 are "in
exactly the same evidentiary position" for GO:0030198. That was true when I wrote it and this
paper made it false: the two remain in the same IBA position, but ADAMTSL5 has its own IDA to
GO:0031012 plus microfibril and heparin binding, while ADAMTSL1 now carries a measured
negative against the route by which the family reaches GO:0030198. The difference runs
against this review
, not for it — the better-supported gene is the one taking the harsher
action. That does not change the verdict here, but it converts the family-wide line-drawing
question from a tie into an argument, and it is carried into suggested_questions as such.

8c. A gap in checkquotes.py worth knowing about

Reconciling a count that did not add up (I added three quotes and the checker's total rose by
one) turned up a real scope gap rather than an arithmetic slip. checkquotes.py walks
supported_by and findings only; provenance lists inside knowledge_gaps are invisible
to it
, as is any other reference_id + supporting_text pair under a differently-named key.
For this file: 17 supported_by + 32 findings = 49, exactly what the checker reports, and
6 provenance quotes are unchecked. They were verified here with a walker that matches on
the shape of the entry rather than on its parent key, and all 6 pass. Anyone relying on
checkquotes.py alone for a review that uses knowledge_gaps provenance is not checking those
quotes.

But it is a local-script gap, not a project-gate gap — I initially implied the wider risk and
that was wrong. The gate that actually runs in CI is linkml-reference-validator via
just validate, and it does walk provenance. Established by breaking it rather than by reading
code: injecting THIS SENTENCE APPEARS IN NO PUBLICATION ANYWHERE ZZQQXX. as a provenance
supporting_text turns just validate from ✓ Valid into
✗ Invalid ... Text part not found as substring. The probe asserted its anchor was present before
mutating, and the file was restored to a clean git diff afterwards. So provenance quotes in this
repo are gated; it is only the scratch checker that misses them.

8d. The recurring defect in this review, and what finally caught it

Five of the review rounds on this gene found the same defect shape: a claim corrected in the
field I was thinking about and left standing in a parallel field saying the same thing elsewhere.
The instances were the GO:0030198 reason vs core_functions; the top-level MF gap boundary vs
the nested BP gap boundary; the suggested_questions tail vs its own narrowed opening; the
knowledge_gaps resolution vs suggested_experiments; and the reference findings statement vs
the row reason. The reviewer caught four. The fifth I found only by giving up on per-field
fixes and sweeping the whole parsed document for every phrasing I had ever retracted
— it was
in the top-level MF boundary, which had already been rewritten twice without anyone noticing the
old double-count framing still sitting in it.

The sweep is seven regexes run over every string value in the parsed YAML:

two partners that place | both been tested for ADAMTSL1 and both were negative |
exactly the same evidentiary position | well-constructed and unrefuted |
LTBPs remain untested | no positive binding result of any kind |
divergence is not attributable to a difference between the genes

It is the ACBD3 lesson (a claim asserted at N sites with no generation relationship between
them) in a review that had read that lesson and still reproduced it five times. The transferable
part is not "be careful": it is that the unit of correction has to be the document, not the
field
, and a retraction list scanned over all prose is cheap enough to run on every round.

And the sweep is only half the control. The commit that fixed instances 1-5 produced a
sixth of the same shape: a (data not shown) caveat added to the reference findings and to
the row reason but not to the nested BP boundary. A retraction list structurally cannot catch
that — it looks for old phrasing left standing, whereas this is new phrasing propagated
to some siblings and not others. The complementary check has to run at write time and enumerate
the siblings: for the C-terminal widening the sibling set is references.findings[1],
existing_annotations[0].review.reason and
existing_annotations[0].review.knowledge_gaps[0].boundary, and the edit now asserts that both
the claim and its caveat appear in all three before it is accepted. Two directions, two
checks: sweep for what should be gone, assert co-occurrence for what should be everywhere.

And the co-occurrence check failed too, for the reason the reviewer named: it is only as good
as the hand-authored sibling list, and enumerating siblings is the same judgement that produced
the six instances.
My list had three members; the widened claim was load-bearing in a fourth —
the top-level MF boundary, which concluded "what is excluded is fibrillin-1 binding" while
stating the evidence as the N-terminal half only, so the field used the widened scope without
carrying what licenses it. Three of four beats the sweep alone, but the residual failure mode
survived the new control.

The fix is to stop hand-listing. Derive the sibling set from the document: collect every
string field, select those that draw the widened conclusion (excluded is fibrillin-1 binding
or covers fibrillin-1 entire), and require each to state the basis (C-terminal half). That
found the missing field immediately and now passes on three. Two further guards make the check
non-vacuous: it asserts the selector matched something at all, and the selector keys on the
conclusion rather than on a list of field paths, so a claim copied into a new field is caught
automatically. Generalisable form: select the fields by what they assert, not by where they
live.

And that control leaked too, in a smaller way — which turns out to be the real finding. The
selector matched two fixed English phrasings, so it missed the reference findings statement,
which draws the same conclusion as "covers fibrillin-1 in its entirety, not only the N-terminal
half"
. Nothing was wrong in the file — that field carried the basis and the caveat and passed on
merit — but had the caveat later been dropped from it, the check would still have reported all
green, because the field asserting the conclusion in different words was never selected. Matching
on fixed phrasings is still selecting by how a field asserts something; it is one abstraction
level up from a path list, not a different kind of thing. Broadened to four patterns, and the
selected set went 3 → 4.

The pattern across the whole review is the point. Each control caught one more instance than
its predecessor and left a smaller version of the same gap:

control caught residual gap
fix the field in front of me 1 at a time 5 parallel fields left stale
retraction sweep over all prose the 5th instance cannot see new phrasing propagated unevenly
hand-listed sibling co-occurrence 3 of 4 siblings list is the same judgement that caused the defect
conclusion-derived selector the 4th sibling selector is phrase-literal, misses paraphrase
broadened phrase set the paraphrase still phrase-literal in principle

No text-matching control closes this completely, because the thing being checked is whether two
sentences mean the same. The honest summary is that each round bought a real reduction and none
of them bought closure, and that a reviewer reading the parsed fields end to end caught what
every mechanical control missed.
Note the notes file legitimately still contains three of those strings, because explaining why a
phrasing was retracted requires quoting it — so the sweep belongs on the review YAML, not on
prose that discusses it.

9. What ADAMTSL1 is, in one paragraph

A large secreted ADAMTS-like glycoprotein of the extracellular matrix, built from
thrombospondin type-1 repeats, Ig-like C2-type domains and a PLAC domain, with none of the
catalytic machinery of the ADAMTS proteases. Its biosynthesis is unusually
glycosylation-dependent: O-fucosylation of the TSRs by POFUT2 with B3GLCT extension, and
C-mannosylation of TSR1 tryptophans, together gate export from the ER, and a natural
Trp42Arg substitution at a C-mannosylation site blocks secretion and acts dominant-negatively
in a family with congenital glaucoma, craniofacial, dental, auditory, renal and limb
anomalies. Once secreted it is deposited punctately into the matrix and is a substrate of
MMP10 in fibroblast secretomes. Its molecular function is undetermined.