Next 20 fly genes — source data and reproduction

Next 20 fly genes: source data and reproduction

This package records the metabolically enriched 20-gene selection, its original prediction outputs and its source provenance. Selection is distinct from completed biological review.

Inputs

Published-subset records reuse the existing prediction snapshot, ordinary UniProt snapshot, FlyBase FB2026_02 identifier table, and source manifest. Their retrieval date is 2026-09-08; they are not represented as freshly fetched records. The new snapshot retains all 29 additional entries, including alternatives that were not selected.

The combined candidate set contains 123 accessions but 119 FlyBase genes. The first cohort is excluded by FlyBase gene ID, not merely by accession, so an alternative record for Lcp3 or ftz-f1 is not counted as a new gene. Current FlyBase nomenclature also resolves the UniProt dnc label to Pde4.

Derived outputs

Reproduction

The derivation script validates frozen-source checksums, exact accessions, taxonomy, FlyBase mapping and non-overlap with the first cohort before regenerating the tables:

UV_NO_SYNC=1 uv run python projects/PROTNLM_EVALUATION/fly-benchmark/next20/summarize.py

To repeat discovery against another database release, retrieve the UniProt taxonomy-index URL in the manifest and join its Entry column to accession elements in the original export. Compare against the published-list snapshot, fetch both API responses for each additional accession, and save a new snapshot. Reconcile gene identities and manual selection explicitly; do not overwrite these frozen inputs with current data.

No schema-valid gene reviews or prediction judgments are generated by this selection script. Gene review, exact-isoform verification and claim-by-claim assessment are the subsequent work.