Next 20 fly genes: source data and reproduction
This package records the metabolically enriched 20-gene selection, its original prediction outputs and its source provenance. Selection is distinct from completed biological review.
Inputs
- selection.csv: the explicit 20 accession choices, current FlyBase symbols/IDs, biological leads and review questions. These choices are manual, not scores inferred from annotation overlap.
- manifest.json: source dates, URLs and SHA-256 checksums.
- Original-export fly subset: 123 current-primary-accession matches from
post-processed-2026_02_28k.xml. XML is reserialized with original element values, attributes and evidence retained; its placeholder sequences, taxa and dates are not biological inputs. - Current fly taxonomy index: 42,881 UniProt records, downloaded on 2026-09-09 UTC, used to identify fly accessions in the original export. The search uses current primary IDs; it does not claim exhaustive coverage of historical aliases or all possible live-API records.
- Additional API responses: complete prediction and ordinary UniProt responses for all 29 accessions absent from the published subset. All returned HTTP 200 with exact accession identity.
Published-subset records reuse the existing prediction snapshot, ordinary UniProt snapshot, FlyBase FB2026_02 identifier table, and source manifest. Their retrieval date is 2026-09-08; they are not represented as freshly fetched records. The new snapshot retains all 29 additional entries, including alternatives that were not selected.
The combined candidate set contains 123 accessions but 119 FlyBase genes. The first cohort is excluded by FlyBase gene ID, not merely by accession, so an alternative record for Lcp3 or ftz-f1 is not counted as a new gene. Current FlyBase nomenclature also resolves the UniProt dnc label to Pde4.
Derived outputs
- candidate-inventory.csv: all 123 records, distinct gene mappings and output counts.
- cohort.csv: the selected 20 records with source tier, sequence hashes and review questions.
- prediction-statements.csv: exact GO IDs/labels, function paragraph, SL identifiers/labels, protein names and their original source objects. Recommended and submission protein names are both preserved.
- sequences.fasta: selected current sequences, not proven prediction-time inputs.
- summary.json: scope counts and selection status.
- selection-validation.json: source/identity checks, deterministic regeneration, history validation and rendered-link checks.
Reproduction
The derivation script validates frozen-source checksums, exact accessions, taxonomy, FlyBase mapping and non-overlap with the first cohort before regenerating the tables:
UV_NO_SYNC=1 uv run python projects/PROTNLM_EVALUATION/fly-benchmark/next20/summarize.py
To repeat discovery against another database release, retrieve the UniProt taxonomy-index URL in the manifest and join its Entry column to accession elements in the original export. Compare against the published-list snapshot, fetch both API responses for each additional accession, and save a new snapshot. Reconcile gene identities and manual selection explicitly; do not overwrite these frozen inputs with current data.
No schema-valid gene reviews or prediction judgments are generated by this selection script. Gene review, exact-isoform verification and claim-by-claim assessment are the subsequent work.