Retrieval recall in production use (PAINT campaign, n=91)

Retrieval recall in production use (PAINT campaign, n=91)

The other cohorts in this project evaluate what Affinage says — whether its GO
grounding matches the curated function, whether its narrative recovers what the
grounding drops. This one evaluates what it finds. Across 91 human genes
reviewed with Affinage as the deep-research provider, how much of the literature
the finished review depends on did Affinage actually supply?

Numbers here are generated by retrieval_recall.py;
see paint-campaign/summary.md for the per-gene
table and paint-campaign/per-gene.json for raw
counts. Re-run with uv run python retrieval_recall.py --all.

Bottom line

Affinage supplies about half the references a review has to go find, and its
trust gates cannot tell you which half is missing.

PMIDs Affinage returned 1344
PMIDs cited by the finished reviews 1626
... already supplied by GOA (no search needed) 908
References the reviews had to find 718
... supplied by Affinage 371
Pooled novel-reference recall 52%
Fraction of Affinage's returned refs the reviews used 39%

The denominator matters. Counting every PMID in a finished review makes recall
look like 32%, but roughly 56% of those references arrive prepackaged in the GOA
file — the reviewer is handed them and no retrieval is involved. Scoring a
retrieval provider against references it was never asked to retrieve measures
nothing. Restricted to the 718 references the reviewer genuinely had to locate,
recall is 52%.

The complementary 61% figure — Affinage references the review never cited — is
the ordinary cost of a literature sweep and not by itself a defect. A provider
that returns twenty papers of which eight are load-bearing has done useful work.

Precision is gated; recall is not

gates_passed: True certifies that the citations in a report are real,
resolvable, and quoted correctly. That is a precision guarantee, and it holds up.
What it cannot certify is that the report found the papers that decide the
review, because nothing in the pipeline measures recall.

The gap is not hypothetical. Two verified cases from this campaign, both on genes
whose reports were otherwise clean:

Both are the same shape: the decisive paper for a sparsely-annotated gene is
usually titled for something else — a partner, a complex, a paralogue, or the
well-characterised family member. A symbol-keyed search does not reach it, and a
measured negative, which is the most decision-relevant evidence available on a
dark gene, is almost never in the title.

Recall does not depend on how well-studied the gene is

The sorted per-gene table invites a tempting misreading. Genes at the 100%-recall
end (RAD51C, RFWD3, SLX4, UBE2T, XRCC2) are all well-characterised, and five of
the seven at the 0% end (AADACL2/3/4, ACP7, ACTL10) have no GOA references at
all, which looks like a clean story about the provider failing on dark genes.

The 100% end is an artifact of small denominators. Those genes had one to five
novel references each, and a gene with a single novel reference scores 0% or 100%
and nothing in between.

The 0% end does not hold up either. The other two zero-recall genes are not
obscure: by the script's own banding, ACTR8 has 12 GOA references and sits in the
well-studied band, and ACTR1B has 8 and sits in the medium band.

Pooling references within bands of curation depth removes both effects:

curation depth genes novel refs supplied recall
dark (0-2 GOA refs) 25 167 87 52%
medium (3-9) 41 436 220 50%
well-studied (10+) 24 115 64 56%

Recall is flat. Affinage misses about half the findable literature regardless of
how much prior curation exists.

What differs across the bands is the consequence of a miss, not its rate. On
a well-studied gene, half the literature still leaves several independent papers
establishing the same function, and the review survives. On a dark gene, half the
literature can mean the difference between one functional experiment and none —
which is exactly what happened with ADAMTSL1 and ACTG2.

Silent empty returns

Six reports returned zero PMIDs: AADACL2, AADACL3, AADACL4, ACP7, ACTL10, ACTR8.
Their reviews went on to cite 6, 3, 4, 1, 2 and 27 references respectively, all
found by other means.

Two features make this worse than a low score. The AADACL cluster is an entire
paralogue family returning nothing, suggesting the failure attaches to a
neighbourhood of sequence space rather than to individual queries. And an empty
report is not flagged: ACTL10's carries gates_passed: True, which is
technically correct — a report with no citations has no false citations — but
reads as success. A gate that passes vacuously on empty input is a gate worth
changing.

Only one further gene, ACTR1B, received a non-empty report that supplied none of
its novel references — so outside the six empty returns, a report that contains
citations at all almost always contributes at least one.

Recommendations

Treat Affinage as a strong first pass, not as the literature search. It front-loads
the obvious references cheaply, and 56 of 91 reviews with a committed report cite
it as a source, so the reports are used rather than filed and ignored.

For any gene where the review turns on a small number of experiments, search
independently on partners, complexes, family members and paralogues, and run
QuickGO by reference on load-bearing PMIDs to see what else they annotate. This
is already standing practice for the per-gene review agents in this campaign and
is what surfaced both documented misses above.

Two changes would make the provider's own output more honest without touching
retrieval: report an explicit zero-citation warning rather than passing gates
vacuously, and record in the gene notes when the reviewer found a load-bearing
reference the provider missed. Recall cannot be measured at all unless the misses
are written down somewhere.

Limits of this analysis

52% is an upper bound, for two reasons, and the second is the larger one.

The reviews were written with the Affinage report in hand — 56 of the 91 cite it
as a source. The reference set being scored is therefore partly caused by the
thing being scored: a paper Affinage surfaced is more likely to end up in the
review than an equally relevant paper it did not, so the numerator is enriched by
construction. Measuring recall against a reviewer who read the report is not the
same as measuring it against an independent literature search, and the honest
reading of 52% is "of the references that ended up mattering, Affinage had already
supplied about half" — not "Affinage finds half of what is findable." A clean
estimate would need reviews written blind to the report.

Recall is also measured against a single finished review per gene, which is an
imperfect standard in the other direction: a reference neither Affinage nor the
reviewer found is invisible here.

Reference sets are compared by PMID string matching, so a paper cited in the
report by DOI or title alone would score as a miss. In this cohort that caveat
never bites — no Affinage report cites by DOI anywhere.

The 91 scored genes are those with a committed Affinage report, not a random
sample of the campaign. Genes where the provider errored and the reviewer moved
on without committing a report are absent, which biases the sample toward
successful runs.

Gate statistics are thin: only 33 of 91 reports carry a gates_passed field at
all (29 True, 4 False), the remaining 58 predating the field.