Frozen pombe ProtNLM cohort
The pombe project page reports the biological assessments.
This snapshot identifies 28 original-export accessions matching current
Schizosaccharomyces pombe UniProt primary accessions. All 20 entries with GO or
function text form the initial review cohort: 32 GO claims and 11 function
paragraphs. Eight other records remain in the census.
Sources and scope
The published 26,856-record ProtNLM2 accession list
contains no pombe records. However, all 28 original-export pombe accessions return
predictions from the live ProtNLM API. List omission is not API unavailability.
All 28 ordinary UniProt entries are currently reviewed/Swiss-Prot; the prediction
API's placeholder TrEMBL classification is not their curated record status.
The source export post-processed-2026_02_28k.xml contains 28,553 entries with
placeholder organism, sequence and date fields. Membership is therefore resolved
by joining original primary accessions to the current UniProt taxonomy_id:4896
index, followed by the PomBase accession-to-gene table. This identifies the cohort
without treating placeholder taxonomy as biology. It does not establish an
exhaustive census of all API-served pombe accessions or search historical secondary
accessions. Gene symbols follow the frozen PomBase table, including mre11
(UniProt rad32), crt10 (pi073), cem1 (SPBC887.13c) and asr1
(SPCC126.07c).
Frozen inputs
- Original subset XML: all 28 complete entry elements,
reserialized with original element values, attributes, paragraphs and evidence
retained. The full export's filename, record count and SHA-256 are in the
manifest; the local original export is not redistributed here. - Published accession list: exact downloaded
bytes, retaining the evidence for absence from that list. - Current pombe UniProt index: 5,228 records.
- PomBase identifiers: original downloaded
identifier table, including systematic identifiers and current gene symbols. - Current UniProt records: 28 complete ordinary API
JSON objects, serialized as sorted-key JSON lines and deterministic gzip. - Prediction API responses: all 28 complete
responses with accessions, URLs and HTTP status, in the same gzip format.
The manifest hashes the stored bytes, including compressed bytes for gzip
files. Current sequences do not prove the exact input sequences used for model
prediction. Placeholder dates and annotation overlap do not establish training
membership.
Derived tables
- Inventory: all 28 accessions with resolved gene symbols,
systematic IDs, current lengths/status and GO/function counts. - Functional cohort: all 20 GO/function-bearing entries.
- Original GO/function statements: 32 GO labels and
11 intact paragraphs, with original evidence keys. - Claim provenance: the same claims with their full
referenced XML evidence fragments. Scores and phmmer/TMalign matches describe
prediction provenance, not independent biological validation. - All statements: also retains the 10 location
statements and one keyword. Predicted protein names remain in the raw snapshots. - API/source comparison: all GO IDs/labels agree.
All 11 API function paragraphs omit the XML's final period; otherwise their
text agrees exactly. Evidence metadata may differ between representations and
remains available in both frozen sources. - Current sequences and census summary.
The coverage checker writes a review inventory checking exact-accession coverage,
source GO IDs/labels, retained original paragraphs, annotation actions, available
research files, source/reference paths and sequence consistency. It does not
assign biological assessments or replace schema/evidence validation.
Validation summary records checked input hashes;
prediction evidence results preserve
per-file title and excerpt checks.
Reproduce
From the repository root, regenerate the tables without network access:
uv run python projects/PROTNLM_EVALUATION/pombe-benchmark/summarize.py
uv run python projects/PROTNLM_EVALUATION/pombe-benchmark/review_inventory.py
The summarizer verifies the snapshot checksums before reading the frozen inputs.
To retrieve a separate comparison snapshot, supply the original XML export
and a new empty output directory:
uv run python projects/PROTNLM_EVALUATION/pombe-benchmark/fetch.py \
--source-xml /path/to/post-processed-2026_02_28k.xml \
--out-dir /tmp/protnlm-pombe-refetch
The fetcher computes cohort membership and counts from the supplied export and
current sources. It refuses to overwrite an existing nonempty directory. A later
snapshot can differ in annotation, identifier mapping or prediction availability;
it is not automatically the same release.