Mammalian ProtNLM2 availability census

Mammalian ProtNLM2 availability census

Frozen on 8 September 2026 to support the
benchmark design. This is an inventory, not a
biological assessment. No new gene-review verdicts are supplied.

Selected horse cohort

The 40-gene review list is a targeted selection from this census:
selection CSV, 89 GO predictions and 17 function descriptions,
current sequences, ordinary UniProt records,
and selection manifest with checksums.
The census remains the full 813-record inventory; summarize.py does not change
this manually selected cohort.

Files

Reproduce offline

From this directory, using Python 3.10 or later (standard library only):

python summarize.py

This regenerates the three census CSVs from the frozen inputs and fails if the
accession sets differ or the payload contains duplicate accessions. Protein names
are metadata, not verified orthology assignments. The source JSON contains
placeholder entry-audit dates and no reliable prediction-time sequence record;
do not use those placeholders for temporal holdouts or sequence-version matching.
The summarization script is included alongside the data.

Fetch a separate updated comparison

The existing project fetcher can retrieve these accessions again:

python ../fetch_protnlm_api.py --accessions accessions.tsv --out-dir /tmp/protnlm-mammal-refetch

Its output is a new, mutable-service snapshot; preserve the frozen inputs when
comparing releases. Retrieve a new official list as well if assessing release
coverage changes. The census intentionally does not search the whole API for
unlisted records.

Before selecting or scoring a benchmark, add the actual sequence and checksum,
entry version, isoform/gene model, orthology evidence, and prediction version where
recoverable. Historical error annotations need their own dated source snapshot.