Cross-cohort results
Project overview · Source counts · Summary generator · Narrative category index
Counts are a dated snapshot: as of 2026-09-27 (commit c7551cb3db). The scope contains 282 distinct protein records, including 242 prediction targets and 40 paired human reference records. Overlapping selections are counted once in the combined totals. These purposive, retrospective cohorts test informative biological distinctions; their proportions do not estimate proteome-wide accuracy.
GO-term assessments
Each row counted here is one emitted GO term. Narrative functions, protein names and SL localization outputs are excluded. Zero GO PLI or REP judgments does not imply the absence of narrative errors.
| Category | GO claims |
|---|---|
| COR | 52 |
| CNN | 32 |
| LSP | 84 |
| UNC | 98 |
| NPI | 20 |
| PLI | 2 |
| REP | 0 |
| Total | 288 |
Reviewed records with no GO predictions
The 162 prediction-review YAML files include 17 records with zero emitted GO predictions and 145 records with GO assessments. Completed predictions: [] records with a summary evaluation document reviewed output absence; a missing review file is not counted as zero output. Their descriptions record the available evidence for potential missed functions. No VDCL category or confidence score is assigned to an absent prediction, and these records do not enter the emitted GO-claim denominator. Zero output alone does not establish a biological false negative or a recall estimate.
Narrative function reviews
57 gene/accession review records assess emitted FUNCTION text. A record may contain multiple paragraphs and multiple claim categories. Each category below counts records with at least one such judgment, once per record; categories overlap and must not be summed or pooled with GO counts. These are neither atomic-claim counts nor one verdict per whole paragraph. Name and localization assessments remain in the gene notes and are outside both denominators.
| Category present in function review | Review records |
|---|---|
| COR | 1 |
| CNN | 19 |
| LSP | 1 |
| UNC | 25 |
| NPI | 14 |
| PLI | 13 |
| REP | 0 |
| SUPPORTED | 8 |
SUPPORTED records contain an explicitly supported claim whose review does not assign a novelty-specific COR/CNN/LSP category. They are preserved as such. References to separate GO judgments, hypothetical corrections, and claims explicitly not emitted are excluded.
Individual narrative records
Cohort scope
The cohort rows retain overlapping selections, including six fly targets selected twice. Paired reference records provide evidence and contribute no extra prediction assessments. Use the deduplicated totals above for the combined corpus.
| Cohort | Records | Records with GO assessments | Reviewed zero GO output | GO claims | Narrative review records |
|---|---|---|---|---|---|
| ARGO50 | 50 | 41 | 9 | 77 | 0 |
| HORSE40 | 40 | 27 | 0 | 89 | 17 |
| HORSE40_HUMAN_PAIR | 40 | 0 | 0 | 0 | 0 |
| FLY41 | 41 | 33 | 8 | 50 | 13 |
| FLY_LOCATION_KEYWORD | 29 | 0 | 0 | 0 | 0 |
| FLY_NEXT20 | 20 | 0 | 0 | 0 | 0 |
| POMBE20 | 20 | 18 | 0 | 32 | 10 |
| POMBE_REMAINING8 | 8 | 0 | 0 | 0 | 0 |
| NEUROSPORA20 | 20 | 16 | 0 | 21 | 3 |
| MOD_EVOLUTION20 | 20 | 10 | 0 | 19 | 14 |
Source metadata
In the JSON summary, prediction_review_files counts all review YAML files, including explicit empty records. go_review_files is retained as a legacy alias for that same total; records_with_go_assessments counts only records with emitted GO claims.
source_method: ProtNLM2 names the model. source_version identifies the release, XML artifact or dated API snapshot. API retrieval timestamps are observation times, not model training dates. The pilot, exploratory XML and later API snapshots remain distinct; frozen responses and source references retain their original provenance.