Miscitation Audit

IN_PROGRESS PIPELINEEVALUATION

Miscitation Audit

Bottom line: a citation whose identifier resolves to a different paper cannot
support the annotation attached to it, and reviewers here had been flagging such
cases one gene at a time without anyone aggregating them. We wrote
harvest_citations.py, which collects every reference_review.correctness defect
flag across the 4,467 review files and keys it on the citation rather than the gene,
and detect_citation_anomalies.py, which looks for defects nobody has flagged yet. We
keyed on the citation because one bad PMID is usually copied across paralogs or complex
partners, so one correction clears several genes and one discovery says where else to
look. The current register (REPORT.md) holds 28
WRONG_IDENTIFIER rows on 21 distinct citations, 6 of which carry that flag on more than one gene,
plus 257 MISCITED rows (215 citations) that have not been sampled for precision; the
prose below was written at 27 rows / 20 citations, one flag earlier. Most of the defects
came from GOA (233 of 285 flagged citations), so the main deliverable is a bug report to
the assigning groups, and those reports have not been filed yet.

We did this because the existing validators only check internal consistency (the title
matches the PMID, the quote is in the paper), and a wrong PMID imported together with its
own title passes both. The unresolvable-identifier check found exactly one dead PMID
across 24,436 cited; the paralog-mismatch check was measured and does not work.

The two classes are not interchangeable and this page keeps them apart throughout:
WRONG_IDENTIFIER (27 rows / 20 citations) is mechanically checkable and is the
tier to act on; MISCITED (257 rows / 215 citations) is a judgement call that has
not been sampled for precision.

Why the citation, not the gene, is the right key

A single bad citation rarely damages one gene. It is copied across paralogs of a
family or partners in a complex, so one upstream correction clears several genes at
once — and, more usefully, one discovery predicts where else to look.

Seven of the twenty distinct WRONG_IDENTIFIER citations are flagged on more than one
gene. For six the flag is WRONG_IDENTIFIER on every listed gene; PMID:23209302 is
WRONG_IDENTIFIER on NDUFA8 and MISCITED on ACOX1 and SLC25A3 (the committed
REPORT.md lists all seven because it predates a fix that restricts that list to
WRONG_IDENTIFIER genes):

Citation Genes affected What the paper is actually about
PMID:23209302 ACOX1, NDUFA8, SLC25A3 KIF14/Radil/Rap1a signalling in breast cancer
PMID:10970790 ELOVL1, ELOVL2, ELOVL3 cloning of "HELO1" — i.e. ELOVL5
PMID:25732826 NAA10, NAA40 the Naa60 acetyltransferase
PMID:39329031 NPLOC4, UFD1 a clinical study of intellectual disability in Morocco
PMID:23264731 SERP1, SRPRB MTR120/KIAA1383
PMID:17469741 UPF1, UPF2 a melanoma serum-marker study
PMID:19037698 TIM9, TIM10 a colorectal-surgery article

Who owns the fix

This is the split that decides what the project is for. A citation that appears as
an annotation's original_reference_id came from GOA; one that appears only in a
review's references list was added here.

So the primary deliverable is not an internal clean-up — it is a bug report to the
assigning groups
(MGI, SGD, UniProt, GOA). The COX17 case found while reviewing the
mitochondrial copper delivery pathway is representative: a
protein farnesylation IDA on the copper chaperone COX17, citing a paper entirely
about COX10 (heme A:farnesyltransferase, one digit away), assigned by MGI
against a S. cerevisiae accession — and with a term that is wrong even for COX10,
since that enzyme farnesylates heme rather than protein.

Failure modes seen so far

Mode Example
Off-by-one gene symbol COX17 ← a COX10 paper
Paralog substitution ELOVL1/2/3 ← an ELOVL5 paper; NAA10/NAA40 ← a NAA60 paper
Gene-symbol collision ADPRH ← a paper whose "ARH1" is the hypercholesterolaemia gene; BRIP1 ← a paper on the bZIP factor BACH1, which shares BRIP1's alias
Wholly unrelated paper gbpC ← a Legionella SidC effector study; TIM9/TIM10 ← colorectal surgery
Identifier that resolves to nothing PMID:34521819 on STAT2
Wrong organism or subject insc ← a review of zebrafish cardiac development

Tooling

harvest_citations.py — the register

Aggregates every reference_review.correctness defect flag, keyed on citation.

uv run python projects/MISCITATION_AUDIT/harvest_citations.py

Outputs to MISCITATION_AUDIT/reports/:
citation_flags.tsv (one row per gene × citation, with the GOA/ours split),
bad_citations.tsv (one row per distinct defective citation) and REPORT.md.

Its most useful column is contamination spread: citations flagged
WRONG_IDENTIFIER in one review that are still cited without a flag elsewhere.
Seven such citations currently reach twelve unflagged uses.

Spread is a triage queue, not a verdict. The clearest illustration is
PMID:10970790, flagged wrong on ELOVL1/2/3 and cited unflagged on ELOVL5 — where
it is the correct citation, because ELOVL5 is what the paper actually characterises.
By contrast the two unflagged uses of PMID:34521819 (on JAK1 and STAT1) cannot be
correct, because the identifier resolves to nothing at all. Each row needs a human.

detect_citation_anomalies.py — finding what nobody has flagged

uv run python projects/MISCITATION_AUDIT/detect_citation_anomalies.py --check-pubmed

Check A — unresolvable identifiers. Works, and is nearly free. The insight is that
fetch-gene already caches every citation it can resolve, so absence from
publications/ is itself the signal
; the network call only confirms it. Across the
whole repository exactly one cited PMID has no cached record — PMID:34521819 —
and NCBI confirms it returns no document summary. It is cited by three genes and was
flagged on only one. This check should run in CI.

Check B — paralog mismatch. Does not work; recorded so nobody rebuilds it. The idea
was to flag a citation whose cached text names some members of a numbered family but
not the gene citing it. Measured behaviour:

Bare co-citation across a family is normal and is not a defect signal. A working
version of this check would need alias-aware matching (UniProt/HGNC synonym lists)
rather than literal symbols. Output is retained in paralog_mismatches.tsv and
family_clusters.tsv but should not be treated as findings.

Scope and honest limits

Next steps

  1. Adjudicate the 12 unflagged uses of the 7 spreading citations.
  2. Assemble the GOA-sourced WRONG_IDENTIFIER set into per-database reports (MGI, SGD,
    UniProt) and file them upstream.
  3. Fix the one review-only case (ADPRH).
  4. Put Check A in CI — it is cheap, has no false positives by construction, and a
    non-resolving identifier is unambiguous.
  5. Decide whether alias-aware matching is worth building for Check B, or whether the
    class is better caught by reviewers reading the cached text.

Relationship to other projects

Slides