Inferred from Expression Pattern (IEP) Evidence Code Review

MATURE PIPELINEEVALUATION

Species: rat, ARATH, human, worm, DICDI, ECOLI, ORYSJ, MEDTR, mouse, yeast, DROME

Genes: Hmgcs2 Gsta4 Gstt1 Qdpr Hsd11b2 Ckmt2 Pgam2 Ephx1 Gss Gamt Casp3 Hspa8 Mapk1 Tp53 App Notch1 PIF3 CRY1 CRY2 CCA1 TOC1 SOC1 UVR8 SOS1 CBF1 HSP17.6A CDK1 PGK1 CPT1A ACADVL FN3K LGALS3 TOLLIP ACTL8 RB1 BAG6 NFE2L2 ADAM10 APOE HTT Mir26a-1 Mir384 Mir30e Mir100 Mir127 hsp-16.2 hsp-4 hsp-6 hsp-60 irg-1 irg-2 lys-1 skn-1 xbp-1 DnaK DnaJ arnF cotB ecmB mhcA acaA NFP EME1 THI22

Inferred from Expression Pattern (IEP) Evidence Code Review

Bottom line: IEP is the experimental evidence code that infers a gene's role in a process from a change in its own expression, so every IEP row carries a leap from "its abundance moved" to "it takes part". We surveyed all 25,401 IEP annotations in UniProt-GOA and reviewed the 550 IEP rows on 221 genes in this repo as of the 2026-09-13 survey (by 2026-09-26 the repo held 565 IEP rows across 235 gene reviews; the rates below have not been re-derived for those), including a tiered cohort of five miRNAs from one 130-gene batch. We did this to measure whether that leap holds, and to name the ways it fails. The characteristic IEP row is true but peripheral: reviewers accepted only 22.5% of IEP rows (the lowest of any code except IPI) and kept 55.6% as non-core, while 16.7% were flagged as over-annotated, modified or removed. Seven failure patterns recur, led by inducible bystanders (a detox enzyme collecting one response to X term per induction paper) and developmental time-courses read as tissue-building roles. Globally, 1,110 of the 1,147 cellular-component IEP rows that break GORULE:0000006 trace to a single ECO class (ECO:0000279) and could be cleared by one mapping change.

The dispositions come from this repository's own AI reviews, so the rates measure one primed reviewer population. The worked examples below carry the argument, and the open action items (a developmental cohort, the E. coli DNA-damage batch, independent disposition data) are listed at the foot of the page.

Overview

IEP (Inferred from Expression Pattern, ECO:0000270) is the one experimental
evidence code whose underlying observation is not about the gene product's
behaviour at all. IDA watches the protein do something; IMP removes it and
watches what breaks; IPI catches it holding a partner. IEP watches the gene's
own transcript or protein abundance go up or down, and infers participation in
whatever process the experimenter was manipulating.

That extra inferential step is the whole subject of this project. The review
question for an IEP row is therefore not the question asked of a propagated
annotation:

Is this gene product an agent in the annotated process, or is its
abundance merely modulated by it?

The same question is what GO itself asks. The GO best-practices
paper
(PMID:23842463) puts
the bar explicitly:

"The ‘response to’ GO terms are intended to annotate gene products that are
required for the response to occur and are a direct result of the organism’s
reaction to the stimuli... It is acceptable to not annotate from such
expression studies since changes in expression of a gene product does not in
itself indicate its contribution to the function or process. Also, expression
studies can seldom support annotations to a Cellular Component or Molecular
Function term. Thus IEP should be used to annotate to terms in Biological
Process only."

The GO wiki entry for IEP
is blunter still — "Use this code with caution!" — and adds two operational
constraints that this project treats as testable: IEP is "usually used in
conjunction with high level GO terms in the Biological Process ontology",
and only normal expression counts (an overexpression or ectopic-expression
experiment is IDA or IMP territory, not IEP). The BP-only restriction in the
quote above is also a hard validation rule:
GORULE:0000006
enforces it for IEP and its high-throughput twin HEP.

Relationship to the sibling evidence-code projects

This page is the third in a set, and the contrast between them is the point.

Project Where the defect lives Review question
SPKW The mapping layer — a UniProt keyword is converted to a GO term by a rule that ignores the individual gene. Is this keyword→term mapping valid for this gene?
IBA / ISO The transfer — a sound source annotation is propagated to a target that has diverged. Is the source sound, and is this term safe to move across this edge?
IEP (this page) The inference itself — there is no mapping and no transfer, only the leap from a correlation in abundance to a claim of participation. Is the gene an agent in this process, or a bystander whose expression happens to track it?

Because there is no source annotation and no propagation edge to audit, the
review.propagation_review machinery that IBA and ISO reviews rely on only
partly fits IEP; see Action items.

Corpus Snapshot

Three views. IEP/iep_corpus_survey.py (output:
iep-corpus-survey.md) produces two of them: the
GOA view counts rows in the repo's cached *-goa.tsv downloads, and the
review view counts annotations in *-ai-review.yaml, i.e. what reviewers
concluded. Neither is a probability sample of IEP, because a gene directory
exists only because somebody chose that gene for review.

IEP/iep_global_atlas.py (output:
iep-global-atlas.md) supplies the denominator: the
complete set of IEP annotations in UniProt-GOA, downloaded from QuickGO. IEP
is rare enough that no sampling is needed — there are only 25,401 of them.

Global (UniProt-GOA) Repo GOA files Coverage
IEP annotations 25,401 555 2.2%
Gene products with IEP 10,618 218 2.1%
Distinct GO terms 2,383 288 12.1%
Distinct references 11,777 368 3.1%

Within the repo, IEP is rare in the same way it is globally: 555 of 149,995
cached GOA rows (0.4%), roughly one IEP row per 40 IDA rows, spread over 220
gene directories (218 distinct accessions, two of which have a directory under
two different symbols). The review view finds 550 IEP rows in
221 review files. The two sides do not line up exactly, and neither is a subset
of the other: cached GOA files and reviews are refreshed independently, so a
review can retain a row that a later GOA download dropped, and a GOA row can
arrive after the review that would have carried it. The gap is small — under 1%
of rows, one review file — but it means the two views should be read as two
measurements of the same thing rather than as one filtered from the other.

Review-view denominators include all entries in existing_annotations, including
reviewer-proposed NEW annotations. The disposition table reports NEW
separately; those entries are not necessarily annotations supplied by GOA.

The reviewed sample is 2% of global IEP, so before drawing conclusions from it,
see how representative it is. The short
answer: close to global shares for coarse term types, badly skewed by organism
and annotation group. Matching those marginal shares does not establish that
reviewer flag rates are representative.

What "flagged" means, and does not mean

Every disposition on this page — every ACCEPT rate, every flag rate, the
cross-code comparison in the next section — is read out of this repository's own
*-ai-review.yaml files. Those are AI-generated reviews, written under the
project's reviewing guidance, which tells reviewers that "many GO terms are
over-annotations" and that they should "not take existing annotations as gospel,
whether experimental or bioinformatic". Three consequences follow, and they
apply to the recommendations at the foot of this page as much as to the tables:

Where a claim survives only in the statistics and not in the worked examples, it
is marked as suggestive in the text.

IEP is used to say one kind of thing

GO branch (is_a + part_of closure) Global rows Global share Repo share Flagged in repo
response to stimulus (GO:0050896) 17,099 67.3% 69.3% 16.8%
developmental process (GO:0032502) 4,110 16.2% 15.3% 22.6%
biological regulation (GO:0065007) 1,041 4.1% 6.0% 15.2%
unclassified (obsolete/unresolvable) 1,395 5.5% 4.4% 8.3%
cellular component 1,147 4.5% 2.5% 0%
metabolic process / localization / MF 609 2.4% 2.5% 14.3%

Branch assignment is first match wins in the order above, so a term parented
under both response to stimulus and developmental process — a defence
response that is also a developmental one, say — is counted as
stimulus-response. Those dual-parented rows are excluded from the developmental
bucket. A different bucket priority could change the flag-rate comparison; its
effect has not been measured. The repo shares here are the review view (550
rows); the representativeness tables further down use the GOA view (555 rows),
which is why the same stratum can differ by a point.

Seven out of ten IEP annotations are a "response to X" term, both globally and in
the repo, which follows directly from the experiment type: expose an organism to
a stimulus, see which transcripts move. Globally the most frequent terms are
response to xenobiotic stimulus (512 rows), response to cold (387),
response to bacterium (384) and response to abscisic acid (379).

The flag rates in that table hint at a second-order finding — the developmental
branch may be the riskier one
— but the corpus is too small to establish it.
19 of 84 developmental rows were flagged against 64 of 381 stimulus rows, a
difference not separable from noise (two-sided Fisher exact p = 0.21). Splitting
by species does not establish a general developmental excess: the direction
repeats in the three largest species (rat 33.3% vs 25.0%, human 20.6% vs 10.4%,
Arabidopsis 40.0% vs 8.6%), but none of those contrasts reaches significance.
It reverses in Dictyostelium (10.0% vs 18.2%) and in mouse, where 0/4
developmental rows are flagged against 8/8 stimulus rows (Fisher p = 0.002).
The mouse result is an opposite-direction contrast in only 12 rows, not evidence
for a general developmental risk. The developmental rows span 54 gene
directories across seven species, but that spread does not rule out review-batch
or gene-selection confounding.

The mechanistic reason to expect the gap is independent of the numbers, and it
is argued below from worked examples rather than from the flag rate: a "response
to X" row is usually at least true — the transcript really did move when the
stimulus was applied — whereas an "X development" row inferred from a
developmental time-course is a genuine over-reach whenever rising abundance as a
tissue matures reflects demand for the enzyme's product rather than an
instructive role in building the tissue.

The disposition data: IEP is not wrong so much as peripheral

This corpus's reviewers flagged 16.7% of IEP rows (REMOVE +
MARK_AS_OVER_ANNOTATED + MODIFY). That is worse than the other experimental
codes (IDA 6.8%, IMP 6.8%, IGI 6.4%) but comparable to ISO (16.8%) and better
than IEA (20.4%) — not, on its own, a damning number.

The sharper contrast is at the other end of the distribution:

Code Reviewed rows % ACCEPT % whose term reaches core_functions
TAS 15,432 73.9% 56.2%
IBA 9,728 72.3% 50.8%
IDA 21,893 69.3% 43.9%
IEA 31,612 50.6% 29.4%
IMP 9,863 48.3% 32.8%
IGI 1,557 45.3% 26.6%
ISO 4,245 33.1% 16.2%
IEP 550 22.5% 10.0%
IPI 17,834 10.7% 4.3%

The % core column credits a code whenever a term it carries also appears in
core_functions, even when the term got there on the strength of a different
code annotating it too. It is therefore generous to every code, and most
generous to codes that co-annotate often — including IEP, whose 10.0% falls to
9.1% if only rows the reviewer also ACCEPTed are counted.

Read with the limitations above in
mind, IEP has the lowest ACCEPT rate and the lowest core-function grounding rate
of any code surveyed except IPI — and IPI's position is a known artifact of
protein binding rather than a property of physical-interaction evidence. The
missing IEP mass went to KEEP_AS_NON_CORE, which absorbs 55.6% of IEP
rows, the highest share of any code.

So the characteristic IEP annotation is not false. It is true and peripheral:
a real observation about how the gene is regulated, phrased as a claim about
what the gene does. That is a harder problem than outright error, because
nothing in the annotation is checkably wrong.

Two more measurements sharpen it:

Is the reviewed sample representative?

The reviewed corpus is 2.1% of global IEP and was assembled by gene-review
interest, not by sampling IEP. The
global atlas
compares each stratum's repo share against its global share; a ratio of 1.00x
means proportional, above 1 over-sampled, below 1 under-sampled.

Close to global marginal distributions. Several coarse annotation shares
come out close to proportional:

Stratum Global Repo Ratio
response to stimulus branch 67.3% 68.3% 1.01x
developmental process branch 16.2% 15.1% 0.94x
involved_in qualifier 75.3% 72.1% 0.96x
acts_upstream_of_or_within qualifier 19.1% 23.1% 1.21x
biological_process aspect 95.2% 95.9% 1.01x
Genes with ≥5 IEP rows, as share of IEP rows 42.3% 41.3% 0.98x

The global counts independently support two descriptive findings: IEP is
overwhelmingly a stimulus-response code, and its annotations concentrate in a
few genes. They do not measure reviewer dispositions. Similar branch shares
therefore cannot establish that developmental annotations are riskier or rule
out selection and review-batch effects on the flag-rate comparison.

Badly skewed by organism and annotation group. Here the sample is a poor
picture of IEP:

Stratum Global Repo Ratio
Homo sapiens 3.7% 18.2% 4.96x
Dictyostelium discoideum 0.7% 4.1% 6.05x
Arabidopsis / TAIR 19.7% 27.7% 1.41x
Rattus norvegicus / RGD 47.5% 34.2% 0.72x
Mus musculus 8.0% 4.9% 0.60x
MGI as annotation group 3.4% 2.5% 0.73x
Drosophila / FlyBase 4.4% 1.3% 0.29x
EcoCyc 1.6% 0.4% 0.22x

Entirely absent from the repo: AgBase (785 rows), ZFIN (326),
CollecTF (211); Gossypium hirsutum (395), Danio rerio (335),
Gallus gallus (279), M. tuberculosis H37Rv (179).

The human over-sampling is expected — the repo is human-centric. The coarse
term-type shares above remain close to the global shares, but organism
imbalances still limit comparisons of reviewer dispositions.

The MGI row is the one that has moved, and it is worth saying why, because it
shows what targeted sampling buys. MGI matters disproportionately: the largest
single-screen IEP batches in all of GOA are MGI's (see
pattern 3),
so the stratum was exactly where the most extreme instance of a pattern this
page documents actually lives. Before the miRNA cohort,
the repo sampled MGI at 0.05x — 1 row against a 3.4% global share. Five
targeted reviews took it to 0.73x, close to proportional at the row level. At
the product level the gap is barely dented: 5 reviewed of the 437 RNAcentral
gene products, and 3 of 294, 5 of 139 and 2 of 103 rows for the three batch
terms those reviews touched. Row-share parity is not cohort coverage.

A slice the repo under-samples but can represent. 840 global IEP rows
(3.3%), on 437 gene products, are RNAcentral entries rather than proteins —
almost all microRNA precursors from miRNA-profiling studies. This is IEP in its
purest form: the only observation is that a non-coding RNA's abundance changed.
The genes/ tree handles ncRNA entries natively (id: URS…, product_type: MIRNA, fetched with ai-gene-review fetch-ncrna), so the gap was coverage, not
capability. Five of them are now reviewed — see
the miRNA cohort below.

Corrections this forces. Two claims made from the repo sample alone need
restating:

  1. Aspect violations are not a curiosity. The repo's 23 non-BP IEP rows looked
    like a rounding error. Globally there are 1,223 (4.5% CC, 0.3% MF) — and
    the repo sample was proportionally accurate (1.01x for BP, 0.80x for CC) all
    along. The mechanism proposed from 20 rows holds at scale: see
    below.
  2. No GO term is majority-IEP. Within a gene, 72.4% of IEP rows are the sole
    carrier of their exact term (53.5% with closure). But measured per term across
    the whole ontology, IEP is never the dominant support: the most IEP-dependent
    frequent term is seed trichome elongation at 19.6% of its annotations,
    then response to ethanol (15.5%) and
    cellular response to leukemia inhibitory factor (13.7%); most sit below 2%.
    (These shares alone among the atlas figures are not frozen by the committed
    snapshot — the per-term totals come from a live QuickGO count — so they drift
    by a point or two between runs.) Both statements are true and they answer
    different questions — IEP is load-bearing for the gene it sits on, never
    for the term it points at.

Failure Patterns

Pattern Description Examples Typical action
Inducible bystander A constitutively-functioning enzyme is transcriptionally induced by many unrelated stimuli; each induction paper yields one response to X row. rat/Gsta4, rat/Gstt1, rat/Qdpr, rat/Hsd11b2, rat/Gss MARK_AS_OVER_ANNOTATED
Developmental time-course → tissue term Abundance rises as a tissue matures; curated as involvement in building that tissue. rat/Ckmt2, rat/Ephx1, rat/Qdpr, rat/Hmgcs2, rat/Gamt, rat/Pgam2 MARK_AS_OVER_ANNOTATED
Differential-expression screen batch One screen generates one term across many unrelated genes. PMID:21492153 → 8 genes, all epithelial cell differentiation; globally up to 291 genes from one paper MARK_AS_OVER_ANNOTATED / REMOVE
Promiscuous hub inversion A signalling hub whose own transcript answers every stimulus collects the whole stimulus catalogue — while its actual role is to drive those responses. ARATH/PIF3 (7 rows, one paper) MARK_AS_OVER_ANNOTATED / MODIFY
Regulon membership ≠ function Being a transcriptional target of a stimulus-responsive regulator is a property of the promoter, not of the protein. ECOLI/arnF, yeast/THI22 MARK_AS_OVER_ANNOTATED
Marker-gene circularity A cell-type marker's expression is definitionally correlated with the stage it marks. DICDI/cotB, DICDI/mhcA MARK_AS_OVER_ANNOTATED
Wrong-granularity term The term sits at the wrong level for what the experiment showed — too specific (against GO's "use high-level terms" advice) or too broad for a gene whose stimulus is known precisely. too specific: ARATH/PIF3 response to water-immersion restraint stress, DICDI/acaA response to imidacloprid; too broad: ARATH/CRY1, ARATH/CRY2 response to light stimulus, ARATH/SOC1 MODIFY / MARK_AS_OVER_ANNOTATED
Aspect violation CC or MF terms carrying IEP, contrary to GORULE:0000006. 1,223 rows globally, 97% of the CC ones from one ECO class see below

1. Inducible bystander — the "response to X" cloud

The largest and most systematic class. A detoxification or intermediary
metabolic enzyme has one stable job; because that job is useful under stress,
its transcript is induced by a long list of chemically unrelated insults; and
because each induction was published separately, each becomes an annotation.

rat/Gstt1 carries response to salicylic acid, response to selenium ion,
response to vitamin E and response to xenobiotic stimulus — four separate
PMIDs, one enzyme, one activity. rat/Qdpr, whose actual job is quinonoid
dihydrobiopterin reduction in tetrahydrobiopterin recycling, carries response to aluminum ion, response to lead ion, response to glucagon, cellular response to xenobiotic stimulus and liver development. rat/Hsd11b2 has six
flagged stimulus terms.

What makes these hard is that they are not false. GSTT1 activity is part of
how a cell handles a xenobiotic. The problem is proportion: the annotation set
implies a stimulus-specialist when the biology is one broad-specificity
transferase. GO's own criterion — "required for the response to occur" — is the
right discriminator, and it is exactly the question the induction experiment
does not answer.

A sharp variant is the co-exposure artifact. rat/Gsta4's response to zinc ion and response to herbicide come from the same zinc/paraquat co-exposure
experiment (PMID:20553223); the reviewer noted there is "no mechanistic evidence
that zinc directly modulates GSTA4-4". A factorial exposure design yields one
term per factor regardless of which factor drove the induction.

2. Developmental time-course → tissue-development term

rat/Ckmt2 is the cleanest instance. Sarcomeric mitochondrial creatine kinase
mRNA is undetectable in prenatal heart and rises sharply after birth
(PMID:8086475), so it carries heart development and skeletal muscle tissue development. As the review puts it, that profile "reflects the maturing heart's
increasing metabolic demand for phosphocreatine buffering, not a direct
instructive role of Ckmt2 in heart morphogenesis."

The same shape recurs: rat/Ephx1 and rat/Qdpr → liver development, rat/Gamt →
embryonic liver development, rat/Hmgcs2 → lung development and adipose tissue development, rat/Pgam2 → spermatogenesis. In every case the protein is
a metabolic enzyme whose product the maturing tissue needs more of. The tissue
builds the enzyme; the enzyme does not build the tissue.

This is the mechanism behind the developmental branch's higher flag rate (22.6%
versus 16.8%, though as noted that difference is not statistically separable
from noise on 84 rows): a metabolic enzyme genuinely participates in a stress
response in a way it does not participate in organogenesis. The six genes
above are the evidence; the flag rate is a description of how often the pattern
came up, not a demonstration that it exists.

3. Differential-expression screen batch: one screen, one term, N genes

PMID:21492153 is a 2-D gel proteomics comparison of proliferating versus
differentiated Caco-2 intestinal cells. It reports "53 proteins that were
differently regulated during the differentiation process", 34 of them identified
by MALDI-TOF, and those identifications were curated involved_in
GO:0030855 epithelial cell differentiation with IEP. In
this corpus that single paper is the source for eight genes:
human/ACADVL, human/ACTL8, human/CDK1, human/CPT1A, human/FN3K, human/LGALS3,
human/PGK1 and human/TOLLIP — a very-long-chain acyl-CoA dehydrogenase, a
cyclin-dependent kinase, a glycolytic kinase, a fructosamine kinase, a galectin
and a TLR adaptor. Six of the eight were flagged (five MARK_AS_OVER_ANNOTATED,
human/FN3K REMOVE); the other two, human/LGALS3 and human/TOLLIP, were kept as
non-core with reviewers noting the evidence is "correlative expression-pattern"
and the process "secondary" to the protein's actual role.

human/CDK1 is the case that exposes the underlying logic error. The review notes
that CDK1 "promotes proliferation, which decreases during differentiation" —
CDK1's abundance is anti-correlated with its causal contribution. The paper
itself says as much: "proteins associated with proliferation, cell growth and
cancer were downregulated, reflecting the loss of the tumorigenic phenotype of
the cells." An abundance change carries no sign information about the direction
of the causal role, so "changed during process X" and "promotes process X" are
simply different claims — and here the annotation was made from a change in the
wrong direction.

This pattern is the IEP analogue of the SPKW mapping layer: a single upstream
decision (here, "annotate every hit in this screen") propagates a term across a
functionally unrelated gene set, and the resulting annotations look independent
because they sit on different genes.

At global scale this is the dominant shape of IEP, and eight genes was a small
example.
The distribution of IEP over references is sharply bimodal: 62.4% of
the 11,777 references behind global IEP contribute exactly one annotation,
while the top ten contribute 1,130 between them. Every one of the six largest is
a single screen annotated to a single term:

Reference IEP rows Gene products Terms The one term Group
PMID:20439489 — miRNA 34a/100/137 modulate mouse ESC differentiation 291 291 1 cellular response to leukemia inhibitory factor MGI
PMID:23012479 — Impact of lactobacilli on orally acquired listeriosis 153 153 1 response to bacterium MGI
PMID:11967071 — Over 1000 genes are involved in the DNA damage response of E. coli 152 152 1 DNA damage response EcoliWiki
PMID:25858512 — miR-26a/miR-384-5p required for LTP maintenance 130 130 1 long-term synaptic potentiation MGI
PMID:23646144 — miRNAs in organ-of-Corti degeneration in age-related hearing loss 100 100 1 sensory perception of sound MGI
PMID:11486054 — Patterns of gene expression during Drosophila mesoderm development 72 72 1 mesoderm development FlyBase

Three of those are miRNA-profiling studies annotating differentially expressed
miRNA precursors, so "changed abundance" is the entire observation. And
PMID:11967071 is the clearest single case in all of IEP: a genome-wide E. coli
DNA-damage transcriptome, whose own title concedes that "over 1000 genes are
involved", yielding DNA damage response for 152 genes including the maltoporin
lamB, the maltose transport subunit malF, fumarate reductase frdA,
D-serine deaminase dsdA and asparagine synthetase asnA. Those are not DNA
repair proteins; they are genes whose transcripts moved. This is
pattern 5 executed 152
times from one experiment.

Two review consequences follow. A batch-sourced IEP row should be judged against
its cohort, not on its own, because the cohort reveals the annotation rule that
produced it. And because these batches are concentrated in MGI, which the repo
sampled at 0.05x before this project touched it, the reviewed corpus was blind
to the most extreme form of the pattern; the cohort below is the start of fixing
that, and it moved the MGI row share to 0.73x on five reviews.

Cohort review: 130 miRNAs, one term, one paper

PMID:25858512 is a natural experiment, because the paper itself sorts its miRNAs
into tiers. It detected 372 miRNAs in hippocampal slices, found that only
12 changed during LTP, and functionally validated three — miR-26a,
miR-384-5p and let-7a — by electrophysiology, time-lapse spine imaging and 3'
UTR reporter assays. MGI annotated 130 miRNA precursors to
long-term synaptic potentiation. The paper's own abstract describes what the
rest amount to: "presents a catalogue of candidate 'LTP miRNAs'".

Five members were reviewed, one from each tier, so that the batch is tested with
an internal positive control rather than assumed to be wrong:

Tier miRNA What the paper shows Action
Validated mouse/Mir26a-1 Title miRNA; required for LTP maintenance and spine enlargement via RSK3 ACCEPT
Validated mouse/Mir384 Title miRNA; same experiments ACCEPT
Changed, untested mouse/Mir30e Named as one of the six downregulated among the 12 that changed MARK_AS_OVER_ANNOTATED
Detected only mouse/Mir100 Not mentioned anywhere in the full text REMOVE
Detected only mouse/Mir127 Not mentioned anywhere in the full text REMOVE

The tier predicts the verdict exactly. This matters because it separates two
claims that are easy to conflate: the batch is not wrong because it is a batch,
it is wrong for the members whose only qualification is having been detected.
mouse/Mir30e is the instructive middle case — its expression genuinely changed,
so IEP is the correct evidence code and the observation is sound; what fails is
the leap from "changed" to "acts upstream of or within".

The batch also has a defect visible only from the cohort: it omits let-7a,
one of the three miRNAs the paper establishes as required, while including 127
that it does not. The annotation set is not merely over-inclusive, it is
misaligned with the paper's conclusions at both ends.

Three further batch papers annotate these same five miRNAs, and two produced
sharper findings than the LTP batch itself:

Across the 13 IEP rows in this cohort: 2 ACCEPT, 7 MARK_AS_OVER_ANNOTATED,
4 REMOVE — an 85% flag rate against 16.7% corpus-wide, and 4 REMOVEs added to a
corpus that previously held 7 in total. Targeting batch cohorts rather than
individual rows is therefore a high-yield review strategy, which is the practical
lesson for the remaining candidates.

4. Promiscuous hub inversion: the regulator annotated as a responder

ARATH/PIF3 carries seven IEP rows — response to heat, response to cold,
response to ethylene, response to auxin, response to abscisic acid,
response to salt, response to water-immersion restraint stress — six of them
from a single gene-family expression survey (PMID:23708772). All seven were
flagged.

PIF3 is a phytochrome-interacting transcription factor: it runs the light and
hormone response programmes. Its transcript answering every stimulus is what a
signalling hub's transcript does. Annotating the hub as a responder to each
stimulus inverts the regulator/target relationship and buries the actual
function, which the gene's IMP annotations already carry. The seventh row is a
granularity failure rather than an inversion, and is taken up in
pattern 7.

5. Regulon membership is a property of the promoter

ECOLI/arnF is annotated response to iron(III) ion because the arnBCADTEF
operon is induced by iron through the BasS-BasR two-component system. The
review's verdict
is precise: "the IEP evidence code is technically appropriate... However,
annotating a gene to 'response to iron(III) ion' based solely on transcriptional
induction conflates regulation with function. ArnF is a flippase that
translocates undecaprenyl phosphate-alpha-L-Ara4N; it does not participate in
iron sensing, binding, or detoxification." The iron-responsiveness
belongs to the operon's promoter; the flippase does not sense or handle iron.
(This case also appears in the IBA project,
where the same gene's IBA rows are analysed.)

yeast/THI22 is the same shape with an extra twist: thiamine-dependent regulation
is real, but the paper that establishes the regulation (PMID:10383756) also
establishes that THI22 is not required for thiamine biosynthesis, so the
regulon membership points at a process the gene demonstrably does not carry out.

6. Marker-gene circularity in Dictyostelium

Dictyostelium development is staged, and stage-specific genes are used as
stage markers precisely because their expression tracks the stage. A
developmental transcriptome (PMID:25887420) supplies IEP rows for 17 genes here.

The circularity has to be judged case by case, and this corpus contains both
verdicts. DICDI/cotB is a prespore marker: the review flags
slug development involved in sorocarp development because "there is no
evidence that the SP70 protein participates causally in slug development... the
gene's core function lies in spore coat structure, not in the morphogenesis of
the slug."
DICDI/mhcA gets the same treatment for aggregation involved in sorocarp development. But DICDI/ecmB was accepted as core for culmination involved in sorocarp development — ecmB is a prestalk extracellular-matrix protein, so
the late-development induction the transcriptome records is the production of
the material culmination consumes. The marker is the product.

7. Wrong-granularity terms, in both directions

PIF3's seventh row is the over-specific case: response to water-immersion restraint stress (GO:1990785) is a rodent stress-model term applied to a plant
submergence experiment, and DICDI/acaA's response to imidacloprid names one
insecticide from one exposure.

The commoner error runs the other way. ARATH/CRY1 and ARATH/CRY2 each carry a
generic response to light stimulus IEP row, both MODIFYed — and the reviewers'
stated reason is granularity, not agency: the cryptochromes really are
blue-light photoreceptors, so "responds to light" is true and merely
uninformative next to the response to blue light and blue light signaling pathway terms proposed in its place. CRY2's separate IEP rows on response to blue light and response to low fluence blue light stimulus by blue low-fluence system were both ACCEPTed, which is the same judgment made from the other side:
IEP on a photoreceptor's own stimulus is fine once the term names the stimulus
the protein actually senses. ARATH/SOC1's IEP row on positive regulation of DNA-templated transcription was likewise MODIFYed to the Pol II-specific child.

Neither direction warrants REMOVE. The observation is sound and the term is in
the right lineage; only its level is wrong, which is precisely what MODIFY says.

GORULE:0000006 violations are an ECO-mapping artifact

Globally 1,223 IEP annotations (4.8%) sit on non-BP terms, violating the hard
validation rule: 1,147 cellular-component and 76 molecular-function. They are not
sloppy curation, and the diagnosis is worth recording because it is mechanical
and it accounts for nearly all of them.

"IEP" in a GAF is not one ECO class. Of the 25,401 rows, 23,747 (93.5%) are
literally ECO:0000270; the other 1,654 use a more specific descendant class
that collapses to IEP in the GAF projection. Splitting the aspect violations by
ECO class localises the problem almost perfectly:

ECO class IEP rows BP MF CC
ECO:0000270 (expression pattern) 23,747 23,691 36 20
ECO:0000279 (qualitative western immunoblotting) 1,417 285 22 1,110
all other descendant classes 237 202 18 17

The generic parent class is 99.8% BP — essentially rule-compliant. One descendant
class, ECO:0000279, contributes 1,110 of the 1,147 CC violations in GOA
(96.8%)
, and 1,132 of all 1,223 aspect violations (92.6%), while being 5.6% of
IEP.

The repo's 23 examples are that global picture in miniature (its aspect mix is
proportional to global at 1.01x for BP and 0.80x for CC). Twenty are SynGO
cellular-component annotations (postsynaptic density,
glutamatergic synapse, presynapse, presynaptic active zone) on rat/Hspa8,
rat/Mapk1, mouse/Hspa8, mouse/Casp3, mouse/App, mouse/Notch1, human/APOE and
human/HTT. In the cached GOA rows these carry ECO:0000279, "qualitative
western immunoblotting evidence used in manual assertion" — the evidence class
for detecting a protein in a biochemically fractionated preparation such as a
synaptosome or PSD prep.

ECO:0000279
has two relevant ancestors, and they map to different GAF codes:
ECO:0000314 (direct assay evidence used in manual assertion → IDA) is a direct
parent, and ECO:0000270 (expression pattern evidence used in manual assertion
→ IEP) is reached one step further up, via
ECO:0000284
"protein expression evidence used in manual assertion". Collapsing the specific ECO class down
to a three-letter GAF code forces a choice between them, and the pipeline picks
IEP — dragging a localization assay into a BP-only code. GORULE:0000006 itself
names the correct resolution: "For CC annotations that assess the localization of
a gene product, IDA should be used." The remaining three violations are DisProt
molecular-function rows (human/NFE2L2 ubiquitin protein ligase binding,
human/BAG6 molecular function activator activity, human/ADAM10 protein homodimerization activity); human/BAG6's was reviewed and REMOVEd.

The lesson generalises beyond IEP: where a specific ECO class is multiply
parented across GAF code boundaries, the GAF round-trip is lossy, and a
downstream hard rule then fires on an annotation whose evidence was never
actually expression-pattern evidence.

Where IEP Is Legitimate

IEP is not a code to review adversarially. 22.5% of IEP rows were accepted, and
55 IEP rows ground a term in a gene's core_functions — 50 of them rows the
reviewer both accepted and let reach core, the other 5 riding on a co-annotated
non-IEP row for the same term. The accepted cases share
a single property, and it is the discriminator this project turns on: the
gene's job is the response itself
.

The contrast with the failure cases is not about evidence quality; the
induction experiments behind rat/Gstt1 are as sound as those behind worm/hsp-4.
It is about whether the gene product exists for the annotated process. A
heat-shock chaperone is deployed only during heat shock. A glutathione
transferase does the same chemistry whether or not the animal was dosed with
selenium.

Reviewer Checklist

Before accepting or flagging an IEP row.

Read the experiment, not just the term

Ask the agency question

Check the term and the aspect

Prefer the right action

Recommendations

For reviewers. Treat an IEP row as a statement about regulation and ask
whether it has been mis-phrased as a statement about function. The most common
correct answer is "true, keep it, but it is not what this gene is for."

For curators and GO.

  1. Batch-annotating a screen deserves a second look. The signature is
    mechanical and cheap to detect: one reference, one GO term, many gene products.
    Globally the top six such references produce 72–291 annotations each, and
    reference × term × gene-product-count would flag every one of them at
    submission time rather than at review time. PMID:11967071 is the case for
    doing so — a paper titled "Over 1000 genes are involved in the DNA damage
    response of E. coli" should not yield DNA damage response for a maltoporin.
  2. Cap stimulus clouds. When a gene accumulates many response to X terms
    from independent exposure papers and has a single well-characterised
    activity, the informative annotation is the parent term once, not the
    catalogue. rat/Gstt1's four sibling stimulus terms say less together than
    response to xenobiotic stimulus alone.
  3. Developmental time-courses may need a higher bar than stimulus responses.
    The mechanism is clear enough — metabolic demand tracks tissue maturation
    without any instructive role, as rat/Ckmt2, rat/Ephx1, rat/Qdpr, rat/Hmgcs2,
    rat/Gamt and rat/Pgam2 each show individually. The corpus-level flag-rate gap
    (22.6% versus 16.8%) points the same way but is not statistical support:
    at 19/84 versus 64/381 it is within noise (Fisher p = 0.21). This
    recommendation rests on the worked cases, and testing it properly needs a
    developmental-branch cohort sampled for the purpose.
  4. Fix the ECO→GAF collapse rather than the annotations. The GORULE:0000006
    violations come from a multiply-parented ECO class (ECO:0000279) whose GAF
    projection picks IEP over the equally valid IDA. Choosing the parent by the
    annotation's aspect would clear 1,110 of the 1,147 CC violations in GOA
    in one change, with no curation effort at all.
  5. response to terms could carry the requirement explicitly. GO's
    best-practice text says these terms are for products "required for the
    response to occur", but nothing in the term definitions or the evidence code
    enforces it. An annotation-extension or qualifier distinguishing "acts in the
    response" from "is induced during the response" would let the two claims
    coexist instead of competing.

Action Items


Session Notes

2026-09-13 (review follow-up — reconcile dispositions and branch claims)

2026-09-05 (fourth pass — statistics tightened, review response)

Nothing in the qualitative argument changed; what changed is how confidently the
numbers behind it are stated and whether the committed script regenerates all of
them.

2026-08-02 (third pass — the first batch cohort reviewed)

2026-07-27 (second pass — global denominator)

2026-07-27 (first pass — reviewed corpus)

Slides