aroA

UniProt ID: Q88M05
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: COMPLETE
📝 Provide Detailed Feedback

Gene Description

aroA (PP_1770) of Pseudomonas putida KT2440 is a 746-residue bifunctional cytoplasmic enzyme of aromatic amino acid biosynthesis. Its EPSP synthase module (3-phosphoshikimate 1-carboxyvinyltransferase, EC 2.5.1.19; 5-enolpyruvylshikimate-3-phosphate synthase) catalyzes the penultimate step of the shikimate pathway, transferring the enolpyruvyl moiety of phosphoenolpyruvate to the 5-hydroxyl of shikimate-3-phosphate to yield 5-enolpyruvylshikimate-3-phosphate (EPSP) plus inorganic phosphate; EPSP is then converted to chorismate, the branch-point precursor of phenylalanine, tyrosine, tryptophan, folate, ubiquinone, and other aromatic metabolites. In addition to the canonical EPSP synthase domain (Pfam EPSP_synthase; COG0128; TIGR01356 aroA), the protein carries an N-terminal prephenate/arogenate dehydrogenase (TyrA) module (Pfam PDH_N/PDH_C; COG0287) with a NAD(P)-binding Rossmann fold. UniProt annotates a second catalytic activity for this module, prephenate dehydrogenase (prephenate + NAD+ -> 4-hydroxyphenylpyruvate + CO2 + NADH, EC 1.3.1.12), placing it in the tyrosine-biosynthetic conversion of prephenate to 4-hydroxyphenylpyruvate. The protein is thus a fused EPSP-synthase / prephenate-dehydrogenase enzyme contributing to both chorismate formation and downstream L-tyrosine biosynthesis. The shikimate pathway is absent in animals, making EPSP synthase the molecular target of the herbicide glyphosate, a competitive inhibitor at the PEP site.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0003824 catalytic activity
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: Root-level catalytic activity term; uninformative given the specific enzymatic activities annotated below.
Reason: GO:0003824 is the top-level molecular-function catalytic term and conveys no specific information. The protein has well-supported specific activities (EPSP synthase, EC 2.5.1.19; prephenate dehydrogenase, EC 1.3.1.12) that should be used instead.
GO:0003866 3-phosphoshikimate 1-carboxyvinyltransferase activity
IEA
GO_REF:0000120
ACCEPT
Summary: EPSP synthase activity (EC 2.5.1.19); the canonical, core molecular function of aroA.
Reason: Directly supported by sequence/domain evidence: the EPSP synthase Pfam domain (PF00275), HAMAP rule MF_00210, COG0128, NCBIfam TIGR01356 (aroA), conserved PEP and shikimate-3-phosphate binding residues, and mapping to Rhea:21256 / EC 2.5.1.19. This is the defining function of aroA.
GO:0004665 prephenate dehydrogenase (NADP+) activity
IEA
GO_REF:0000002
KEEP AS NON CORE
Summary: NADP+-dependent prephenate dehydrogenase activity inferred from the fused TyrA domain. The cofactor specificity (NADP+ vs NAD+) is not experimentally established for this protein.
Reason: The protein carries a genuine N-terminal prephenate/arogenate dehydrogenase (TyrA) module (Pfam PDH_N/PDH_C; COG0287), so prephenate dehydrogenase activity is a plausible second function. However, UniProt's curated CATALYTIC ACTIVITY block lists only the NAD+ route (EC 1.3.1.12), and the NADP+ specificity here is purely an InterPro electronic inference (IPR003099) with no cofactor evidence. Retain as a non-core, lower-confidence activity rather than a core function.
GO:0005737 cytoplasm
IEA
GO_REF:0000120
ACCEPT
Summary: Cytoplasmic localization, consistent with a soluble shikimate-pathway metabolic enzyme.
Reason: EPSP synthase is a soluble cytosolic enzyme of central aromatic amino acid biosynthesis; cytoplasmic localization is supported by UniProt-SubCell (SL-0086) and HAMAP rule MF_00210, with no signal/transmembrane features.
GO:0006571 L-tyrosine biosynthetic process
IEA
GO_REF:0000120
KEEP AS NON CORE
Summary: Tyrosine biosynthesis; supported via the fused prephenate dehydrogenase (TyrA) domain that converts prephenate to 4-hydroxyphenylpyruvate.
Reason: The TyrA (prephenate dehydrogenase) module places this protein in the tyrosine-specific branch (prephenate -> 4-hydroxyphenylpyruvate, UniPathway step 1/1 of the NAD+ route). This is a real but secondary process relative to the core EPSP synthase / chorismate-biosynthesis role, hence non-core.
GO:0008652 amino acid biosynthetic process
IEA
GO_REF:0000104
MARK AS OVER ANNOTATED
Summary: General amino acid biosynthetic process; correct but non-specific given the more precise aromatic/chorismate terms.
Reason: True but high-level. The more specific processes (chorismate biosynthetic process, aromatic amino acid biosynthetic process, L-tyrosine biosynthetic process) capture the role precisely, making this generic parent redundant.
GO:0008977 prephenate dehydrogenase (NAD+) activity
IEA
GO_REF:0000120
KEEP AS NON CORE
Summary: NAD+-dependent prephenate dehydrogenase activity from the fused TyrA domain; the second catalytic function of this bifunctional protein.
Reason: Supported by the prephenate/arogenate dehydrogenase domain (residues ~14-302; Pfam PDH_N/PDH_C; COG0287) and by UniProt's curated CATALYTIC ACTIVITY block citing the NAD+ reaction (Rhea:13869, EC 1.3.1.12) and the L-tyrosine biosynthesis (NAD+ route) pathway. This is the better-supported of the two prephenate dehydrogenase cofactor variants but remains the secondary (non-core) function relative to EPSP synthase.
GO:0009073 aromatic amino acid biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: Aromatic amino acid biosynthesis; accurate at the family level for an EPSP-synthase / chorismate-pathway enzyme also feeding tyrosine biosynthesis.
Reason: Both functional modules act within aromatic amino acid biosynthesis: EPSP synthase produces the chorismate precursor common to Phe/Tyr/Trp, and the TyrA domain feeds the tyrosine branch. The term is appropriately specific.
GO:0009423 chorismate biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: Chorismate biosynthesis; the core biological process of the EPSP synthase activity (penultimate shikimate-pathway step).
Reason: EPSP synthase catalyzes step 6/7 of chorismate biosynthesis from D-erythrose-4-phosphate and PEP (UniPathway UPA00053/UER00089). This is the most precise and well-supported biological-process term for the core function. Consistent with KT2440-specific metabolic-engineering evidence that tuning aroA expression contributes to flux through the shikimate pathway toward chorismate-derived products (see aroA-deep-research-falcon.md, citing PMID:41029715).
GO:0016491 oxidoreductase activity
IEA
GO_REF:0000104
MARK AS OVER ANNOTATED
Summary: Generic oxidoreductase parent term covering the prephenate dehydrogenase activity.
Reason: Redundant high-level parent of the specific prephenate dehydrogenase activities (GO:0008977 / GO:0004665) already annotated. Provides no additional information beyond the specific terms.
GO:0016628 oxidoreductase activity, acting on the CH-CH group of donors, NAD or NADP as acceptor
IEA
GO_REF:0000117
MARK AS OVER ANNOTATED
Summary: Intermediate oxidoreductase-class parent term for the prephenate dehydrogenase activity.
Reason: A grouping parent of the specific prephenate dehydrogenase (NAD+/NADP+) activities. The leaf terms GO:0008977 / GO:0004665 are retained, so this mid-level class is redundant over-annotation.
GO:0016765 transferase activity, transferring alkyl or aryl (other than methyl) groups
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: Generic enolpyruvyl/alkyl transferase parent term for the EPSP synthase activity.
Reason: High-level parent of the specific EPSP synthase activity (GO:0003866, enolpyruvyl transferase) which is retained. Redundant given the leaf term.
GO:0070403 NAD+ binding
IEA
GO_REF:0000002
KEEP AS NON CORE
Summary: NAD+ cofactor binding by the Rossmann-fold prephenate dehydrogenase (TyrA) domain.
Reason: Consistent with the NAD(P)-binding Rossmann fold of the fused TyrA domain (InterPro IPR046826) and with the NAD+-dependent prephenate dehydrogenase activity. A supporting cofactor-binding term for the secondary activity, so non-core rather than a primary function.

Core Functions

EPSP synthase (3-phosphoshikimate 1-carboxyvinyltransferase) catalyzing the penultimate step of the shikimate pathway, the chorismate-yielding branch of aromatic amino acid biosynthesis.

Supporting Evidence:
  • file:PSEPK/aroA/aroA-uniprot.txt
    Catalyzes the transfer of the enolpyruvyl moiety of phosphoenolpyruvate (PEP) to the 5-hydroxyl of shikimate-3-phosphate (S3P) to produce enolpyruvyl shikimate-3-phosphate and inorganic phosphate; EC 2.5.1.19; EPSP synthase family; chorismate biosynthesis step 6/7.
  • file:PSEPK/aroA/aroA-deep-research-falcon.md
    aroA is annotated as 3-phosphoshikimate 1-carboxyvinyltransferase / EPSP synthase (EC 2.5.1.19) catalyzing S3P + PEP -> EPSP + Pi, the penultimate step of the shikimate pathway leading to chorismate, the common precursor for Phe/Tyr/Trp biosynthesis.

References

Gene Ontology annotation through association of InterPro records with GO terms
Electronic Gene Ontology annotations created by transferring manual GO annotations between related proteins based on shared sequence features
Electronic Gene Ontology annotations created by ARBA machine learning models
Combined Automated Annotation using Multiple IEA Methods
file:PSEPK/aroA/aroA-deep-research-falcon.md
Deep research report (falcon) for aroA / EPSP synthase (Q88M05) in P. putida KT2440
  • Synthesizes EPSPS mechanism, shikimate-pathway/chorismate role, cytosolic localization inference, and KT2440-specific engineering evidence for aroA.
Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.
  • Genome sequence of P. putida KT2440 in which PP_1770 (aroA, Q88M05) is annotated, establishing the locus and organism context for this gene.
Combinatorial engineering pinpoints shikimate pathway bottlenecks in para-aminobenzoic acid production in Pseudomonas putida.
  • In P. putida KT2440 metabolic engineering for para-aminobenzoic acid (a chorismate-derived product), aroA (EPSPS) was tuned among shikimate-pathway genes; reducing aroA expression to native levels lowered product titer, indicating aroA expression contributes to aromatic-pathway flux in vivo.

Suggested Questions for Experts

Q: Is the prephenate dehydrogenase (TyrA) module of P. putida AroA catalytically active in vivo, and does it prefer NAD+ or NADP+ as cofactor?

Q: Does the AroA-TyrA domain fusion form a substrate channel or otherwise functionally couple chorismate biosynthesis with the tyrosine branch in P. putida?

Suggested Experiments

Experiment: Heterologously express and purify Q88M05 and assay both EPSP synthase (S3P + PEP) and prephenate dehydrogenase (prephenate + NAD+/NADP+) activities to confirm bifunctionality and determine cofactor preference.

Experiment: Construct an aroA deletion/complementation in P. putida KT2440 and test for aromatic amino acid (and specifically tyrosine) auxotrophy to establish in vivo requirement of each catalytic module.

Deep Research

Asta

(aroA-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:20:02.871083

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.766)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.720)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[3] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.711)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[4] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[5] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[6] GOnet: a tool for interactive Gene Ontology analysis

  • Authors: M. Pomaznoy, Brendan Ha, Bjoern Peters
  • Year: 2018
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  • DOI: 10.1186/s12859-018-2533-3
  • PMID: 30526489
  • PMCID: 6286514
  • Citations: 247
  • Influential citations: 17
  • Summary: The open-source GOnet web-application is created, which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis and provides insight into the functional interconnection of the submitted entries.
  • Evidence snippets:
  • Snippet 1 (score: 0.701)
    > In a basic workflow, the GOnet application receives a list of gene symbols, protein symbols, or protein IDs (UniProt IDs) as an input, and outputs a graph (an example given in Fig. 1). There are various input parameters which will affect the actual structure of the graph visualized and its appearance. The first main user choice is which GO terms the genes are annotated against:
    > 1. GO terms statistically significantly over-represented in the gene list submitted. 2. A predefined subset (also known as 'GO slim'), or a user-supplied list of terms.
    > In the first case the analysis will be referred to as an 'enrichment' analysis, in the second as an 'annotation' analysis.
    > Input parameters 1) Gene list. A mandatory input parameter containing the genes/proteins of interest. Currently human and mouse data is supported. An example of a human gene list might look like this:
    > Fig. 1 Sample network output generated by GOnet application. Gene differentially expressed in CD4 Bulk Memory T cells in Latent TB patients compared to healthy controls were used as an example [22] The gene list can also be accompanied with a contrast value. For example, This contrast value can be any decimal number, such as the log-fold change of gene expression between two conditions. This is merely a visualization enhancement. If the value is supplied it can be used later to differentially color specific genes in the graph (note different colors of gene nodes in Fig. 1), and visually indicate up-or down-regulation of specific genes and gene clusters.
    > The application can process common gene symbols (like in the example above), UniProt IDs, and MGI Accession IDs (mouse only). The former type of ID (gene symbols), although is the most human friendly, can unfortunately be ambiguous. For example, AIM1 can mean 'absent in melanoma' (also called CRYBG1) or 'Aurora and Ipl1-like midbody-associated protein' (also known as AURKB). Due to this ambiguity UniProt IDs or MGI accession IDs (for mouse) are preferred.
    > 2) GO namespace. Can be any of 'biological process', 'molecular function' or 'cellular component'.

[7] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.699)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[8] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.694)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[9] Functional annotation of parasitic worm genomes, by assigning protein names and GO terms

  • Authors: Avril Coghlan, M. Berriman
  • Year: 2018
  • Venue: Unknown venue
  • URL: https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  • DOI: 10.1038/protex.2018.055
  • Citations: 1
  • Summary: A computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and Go terms from orthologues in other species.
  • Evidence snippets:
  • Snippet 1 (score: 0.693)
    > Given a set of predicted protein-coding genes for a newly sequenced genome, functional annotation involves assigning putative functions to the predicted genes. Two ways in which this can be done are assigning protein names and Gene Ontology (GO;Gene Ontology Consortium, 2010) terms to the predicted proteins. Here we describe a computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and GO terms from orthologues in other species.
    > When assigning protein names, UniProt protein naming rules (www.uniprot.org/docs/nameprot) are followed where possible. This recommends that a good and stable name for a protein is "as neutral as possible"; that a protein name "should be, as far as possible, unique and attributed to all orthologs"; and a protein name "should not contain a specific characteristic of the protein, and in particular it should not reflect the function or role of the protein, nor its subcellular location, its domain structure, its tissue specificity, its molecular weight or its species of origin".
    > In our protocol, a protein name is assigned to each predicted protein based on curated names in UniProt (Bairoch & Apweiler, 2000) for human, zebrafish, Drosophila melanogaster, Caenorhabditis elegans, and Schistosoma mansoni orthologues identified from a database of gene families (e.g. built using Ensembl Compara; Vilella et al. 2009), or (if no information is found from orthologues) based on InterPro (Hunter et al. 2012) domains. Figure 1 shows an example of using our protein naming pipeline for four Strongyloides ratti genes that belong to the tubulin polyglutamylase family (underlined in pink), where four different protein names were assigned to them (in pink), based on names of their C. elegans or human orthologues.
    > Since each of the S. ratti genes belonged to a different subfamily of the tubulin polyglutamylase family, they were assigned different names.

[10] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[11] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.678)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[12] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.677)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[13] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.676)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[14] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.664)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[15] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.663)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[16] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[17] Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser

  • Authors: Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała
  • Year: 2025
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  • DOI: 10.1186/s12864-025-12329-8
  • PMID: 41408139
  • PMCID: 12822150
  • Citations: 2
  • Summary: An R script is developed that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database and provides a single list of updated symbols with annotation about their ambiguity.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > Gene nomenclature contains current official symbols and various numbers of synonyms, which pose a challenge to integrating genomic data and increase the probability that different genes share the same symbol. Therefore, we retrieved identifiers assigned to all protein-coding genes in human, mouse, rat and pig genomes that are available in the Ensembl genome browser (release 113) to assess the number of genes, compare species and identify ambiguous symbols. Results: Our analysis revealed that the total number of symbols, both official symbols and synonyms, used to identify protein-coding genes ranges from 16,600 in pigs to 64,580 in mice. Furthermore, the gene nomenclature is not complete because there are also genes without an assigned symbol, which indicates gaps in understanding protein-coding genes, especially in pigs. We also found a large number of gene symbols that map to more than one gene. These symbols might complicate the identification of about 10% of rat and mouse genes and 18% of human protein-coding genes. A simple solution for this problem is the usage of stable gene IDs assigned by scientific institutions and committees (Ensembl, NCBI, RGD, HGNC and VGNC) provided that the genomic information associated with these IDs is retrieved directly from proprietary databases containing the most accurate data. Finally, although gene symbols may pose a problem with unequivocal identification of genes, there are instances when no other identifiers are available in the literature. Therefore, we have developed an R script performing search of the Ensembl database and integrating data to provide a single list of updated symbols with annotation about their ambiguity. Conclusions: Gene symbols are not always reliable and should be reported together with stable IDs to enable unequivocal identification of genes. Therefore, data containing only gene symbols should be used cautiously to avoid misidentification of genes. A solution for this problem is our R script REgeness that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database.

[18] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[19] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.659)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[20] AgAnimalGenomes: browsers for viewing and manually annotating farm animal genomes

  • Authors: D. Triant, Amy T. Walsh, Gabrielle Hartley, B. Petry, Morgan R. Stegemiller et al.
  • Year: 2023
  • Venue: Mammalian Genome
  • URL: https://www.semanticscholar.org/paper/38a969fd5641e503106cb215010f84ea0a271f99
  • DOI: 10.1007/s00335-023-10008-1
  • PMID: 37460664
  • PMCID: 10382368
  • Citations: 5
  • Summary: This work presents genome visualization and annotation tools to support seven livestock species, available in a new resource called AgAnimalGenomes, and describes the data and search methods available and how to use the provided tools to edit and create new gene models.
  • Evidence snippets:
  • Snippet 1 (score: 0.655)
    > As previously described (Triant et al. 2020), once a proteincoding gene annotation is complete, each new or modified isoform should be compared to a well-curated protein sequence database to check for congruency with known proteins. The sequence of an annotation is obtained by right clicking it and selecting Get Sequence. The first choice of database to search is the well-curated UniProtKB/Swissprot database using BLAST at either the UniProt (https:// www. unipr ot. org/ blast) or NCBI website (https:// blast. ncbi. nlm. nih. gov/ Blast. cgi) (Sayers et al. 2023a;UniProt Consortium 2023). If there is no match with a significant e-value (< 1e−05) in UniProtKB/Swissprot, the next database to try is the Model Organisms (landmark) database at NCBI. If that fails, select the RefSeq Proteins database and exclude your organism of interest from the search. Although RefSeq includes computationally predicted and hypothetical proteins, an alignment to a homologous protein from another organism provides support for the annotation. An alignment that covers the full length of both the annotated protein and the database protein sequence suggests the annotation is correct. An alignment that encompasses the full length of an annotated protein sequence but only part of a database protein suggests that the annotation is truncated. You may be able to correct the annotation with additional evidence, but if there is not sufficient evidence the issue can be noted in the Annotation Information Panel under the Comment tab. A partial alignment of an annotated protein to a database protein suggests the annotation has a reading frame shift or was extended incorrectly. Aligning the coding sequence (CDS) to the protein database will reveal whether the problem is due to a reading frame shift. Further annotation editing should be performed to correct the reading frame. If an incorrect extension was due to the merging of two genes, you should edit or redo the annotation. Any unresolved issues should be entered in the Comment section of the Annotation Information Panel.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  3. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  4. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  5. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  6. M. Pomaznoy, Brendan Ha, Bjoern Peters (2018). GOnet: a tool for interactive Gene Ontology analysis. BMC Bioinformatics. https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  7. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  8. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  9. Avril Coghlan, M. Berriman (2018). Functional annotation of parasitic worm genomes, by assigning protein names and GO terms. https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  10. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  11. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  12. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  13. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  14. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  15. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  16. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  17. Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała (2025). Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser. BMC Genomics. https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  18. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  19. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  20. D. Triant, Amy T. Walsh, Gabrielle Hartley, B. Petry, Morgan R. Stegemiller et al. (2023). AgAnimalGenomes: browsers for viewing and manually annotating farm animal genomes. Mammalian Genome. https://www.semanticscholar.org/paper/38a969fd5641e503106cb215010f84ea0a271f99

Falcon

(aroA-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 29 citations 1 artifacts 2026-06-11T20:52:12.952603

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: Functional annotation of aroA / EPSP synthase (UniProt Q88M05) in Pseudomonas putida KT2440

Executive summary

The UniProt target Q88M05 is annotated as 3‑phosphoshikimate 1‑carboxyvinyltransferase (EC 2.5.1.19), also called 5‑enolpyruvylshikimate‑3‑phosphate synthase (EPSPS), classically encoded by aroA. EPSPS catalyzes the transfer of the enolpyruvyl group from phosphoenolpyruvate (PEP) to shikimate‑3‑phosphate (S3P) to form 5‑enolpyruvylshikimate‑3‑phosphate (EPSP) and inorganic phosphate, a penultimate step in the shikimate pathway leading to chorismate and aromatic amino acids. A key caveat for P. putida KT2440 is that historical pathway depictions contain an annotation inconsistency in which PP1770 is labeled as “TyrA” and simultaneously described with an EPSPS-like name; this should not be conflated with the well-established bacterial meaning of aroA = EPSPS. The most direct KT2440-specific functional evidence recovered here is pathway-engineering phenotypes: tuning aroA expression affected flux to the aromatic-derived product p‑aminobenzoic acid (pABA).

Target verification and ambiguity handling (critical)

Verified target (user-supplied UniProt context): UniProt Q88M05, gene name aroA, ordered locus PP_1770, organism Pseudomonas putida KT2440.

Detected ambiguity in KT2440 literature: In a KT2440 aromatic-pathway analysis, PP1770 was presented in a pathway context as “PP1770 or TyrA” and described with dual functional labels including “prephenate dehydrogenase, putative/3‑phosphoshikimate 1‑carboxyvinyltransferase.” (molinahenares2009functionalanalysisof pages 2-4). This conflicts with the conventional assignment of tyrA to prephenate dehydrogenase and aroA to EPSPS, and it implies historical misannotation or figure-level conflation. Accordingly, this report treats Q88M05 as aroA/EPSPS (per UniProt target identity) and uses KT2440 papers only for statements they explicitly support.

Category Details Quantitative data Key sources (year; URL) Notes
Verified target identity UniProt Q88M05; gene aroA; ordered locus PP_1770; organism Pseudomonas putida KT2440 Molina-Henares et al. 2009; https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4) Literature for PP1770 in KT2440 exists, but direct biochemical characterization of Q88M05 in the retrieved sources is limited.
Core enzymatic function 3-phosphoshikimate 1-carboxyvinyltransferase / 5-enolpyruvylshikimate-3-phosphate synthase (EPSPS; EC 2.5.1.19) catalyzes shikimate-3-phosphate (S3P) + phosphoenolpyruvate (PEP) → 5-enolpyruvylshikimate-3-phosphate (EPSP) + inorganic phosphate Reaction stoichiometry shown; glyphosate can inhibit by occupying the PEP site Shende et al. 2024; https://doi.org/10.1039/d3np00037k (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 50-64) Current mechanistic understanding places EPSPS as an enolpyruvyl transferase acting through a tetrahedral intermediate; glyphosate is a competitive PEP-site inhibitor.
Pathway context EPSPS performs the penultimate step of the shikimate pathway, leading to chorismate, the common precursor for phenylalanine, tyrosine, and tryptophan biosynthesis Shende et al. 2024; https://doi.org/10.1039/d3np00037k (shende2024theshikimatepathway pages 10-11); Molina-Henares et al. 2009; https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4) In bacteria this is a cytosolic metabolic enzyme in central aromatic amino-acid biosynthesis, inferred from pathway/structural context (shende2024theshikimatepathway pages 50-64).
Annotation inconsistency to flag PP1770 was reported in one KT2440 pathway source as “PP1770 or TyrA” with dual/ambiguous labeling including “prephenate dehydrogenase, putative/3-phosphoshikimate 1-carboxyvinyltransferase” Molina-Henares et al. 2009; https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4) This is the main inconsistency that requires caution; the user-supplied UniProt entry specifically identifies Q88M05 as aroA/EPSPS, so literature must not be conflated with true tyrA/prephenate dehydrogenase studies.
P. putida KT2440 engineering relevance In KT2440 pABA pathway optimization, aroA was included among shikimate-pathway genes tuned by combinatorial expression to improve production Best strain produced 185.4 mg/L pABA; lowering aroA/aroK/aroQ/aroGD146N expression to native levels caused a 39.9% decrease in pABA in top strain S12 Campos-Magaña et al. 2025; https://doi.org/10.1186/s13036-025-00553-5 (camposmagana2025combinatorialengineeringpinpoints pages 2-4, camposmagana2025combinatorialengineeringpinpoints pages 4-6, camposmagana2025combinatorialengineeringpinpoints pages 8-9) Evidence supports aroA as a practical flux-control point in aromatic-pathway engineering in P. putida, although aroB was highlighted as the stronger bottleneck in that study.
Recent EPSPS developments (general) Directed evolution platforms are being used to obtain EPSPS variants with both catalytic competence and glyphosate tolerance One evolved EPSPS variant reached Ki ≈ 1 mM for glyphosate and ~2.5-fold improved enzymatic efficiency versus the starting enzyme Reed et al. 2024; https://doi.org/10.1073/pnas.2317027121 (reed2024evolvingdualtraitepsp pages 1-2) This is not P. putida-specific, but it is highly relevant to modern functional interpretation and real-world use of EPSPS enzymes.
Recent mechanistic expansion (general) A 2024 study showed MurA can also catalyze S3P + PEP → EPSP + Pi in bryophytes, revealing an alternative route to EPSP formation MurA activity was ~100-fold lower than EPSPS; MurA activity on S3P/PEP was ~8-fold higher than on its canonical substrate pair Caygill et al. 2024; https://doi.org/10.1073/pnas.2412997121 (caygill2024muracatalyzedsynthesisof pages 1-2, caygill2024muracatalyzedsynthesisof pages 6-7) Important for interpreting glyphosate tolerance biology broadly; not evidence that KT2440 uses MurA for this role.
Glyphosate resistance relevance In bacteria, resistance can arise through target-site aroA mutations, EPSPS overproduction/gene amplification, transport/efflux changes, or glyphosate degradation/detoxification Example selection range for Salmonella target-site mutants: 0.35–2 g/L glyphosate Hertel et al. 2021; https://doi.org/10.1111/1462-2920.15534 (hertel2021molecularmechanismsunderlying pages 1-5, hertel2021molecularmechanismsunderlying pages 24-27, hertel2021molecularmechanismsunderlying pages 5-8, hertel2021molecularmechanismsunderlying pages 12-15) These mechanisms frame how aroA function is exploited or bypassed under herbicide pressure.

Table: This table summarizes the verified identity, biochemical function, pathway role, annotation caveats, and applied relevance of the target protein UniProt Q88M05 / aroA / PP_1770 from Pseudomonas putida KT2440. It also includes recent quantitative findings useful for interpreting EPSPS function and engineering significance.

1) Key concepts and definitions (current understanding)

1.1 Enzyme name and EC definition

EPSP synthase (EPSPS; EC 2.5.1.19) is an enolpyruvyl transferase in the shikimate pathway that catalyzes:

S3P + PEP → EPSP + Pi

This penultimate step installs a second PEP-derived unit onto the shikimate scaffold to form EPSP, which is then converted to chorismate in the final shikimate-pathway step (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 50-64).

1.2 Mechanism and inhibitor biology

A 2024 authoritative review describes EPSPS as an “alkyl transferase-type enzyme” operating through a tetrahedral intermediate, and emphasizes that EPSPS catalysis involves C–O bond cleavage of PEP (unusual among many PEP-utilizing enzymes, which often cleave P–O bonds) (shende2024theshikimatepathway pages 10-11).

EPSPS is also the canonical molecular target of glyphosate, which competitively occupies the PEP binding site, thereby preventing normal turnover and starving the cell of downstream aromatic amino acids (shende2024theshikimatepathway pages 10-11, hertel2021molecularmechanismsunderlying pages 1-5).

1.3 Enzyme classes (sequence/phenotype concepts)

The same 2024 review summarizes three EPSPS “classes”: Class I (often glyphosate-sensitive; found in plants and some bacteria), Class II (microbial; variable sensitivity), and Class III (microbial; low identity to E. coli EPSPS) (shende2024theshikimatepathway pages 10-11). This classification is widely used to interpret glyphosate sensitivity and potential resistance routes.

2) Biological role and pathway context in bacteria (with localization inference)

2.1 Role in the shikimate pathway and aromatic amino acid synthesis

EPSPS catalyzes the penultimate reaction of the shikimate pathway, which produces chorismate, the common precursor for phenylalanine, tyrosine, and tryptophan biosynthesis (shende2024theshikimatepathway pages 10-11, molinahenares2009functionalanalysisof pages 2-4). Because these amino acids are foundational building blocks and also feed numerous downstream aromatic metabolites, aroA/EPSPS is typically central to anabolic metabolism in bacteria.

2.2 Cellular localization

The retrieved sources do not explicitly state subcellular localization for bacterial EPSPS. However, EPSPS is treated as a soluble metabolic enzyme in core carbon/anabolic metabolism, and bacterial structural context (e.g., E. coli EPSPS with S3P and glyphosate bound) supports the interpretation that it functions in the cytosol rather than in membranes or secretion pathways (shende2024theshikimatepathway pages 50-64).

3) Pseudomonas putida KT2440-specific evidence (genetics, phenotypes, applications)

3.1 KT2440 genetics/auxotrophy context available in retrieved literature

A genome-wide mutant-library screen in KT2440 identified many conditionally essential genes for growth on glucose minimal medium and recovered multiple aromatic amino-acid auxotrophs, including mutants in tryptophan biosynthesis genes (trpA/D/C/E/G/F) and downstream aromatic genes such as pheA and tyrA (molina‐henares2010identificationofconditionally pages 2-3, molina‐henares2010identificationofconditionally pages 6-7). In the retrieved text segments, aroA/EPSPS itself is not explicitly reported as an identified conditionally essential locus (molina‐henares2010identificationofconditionally pages 2-3, molina‐henares2010identificationofconditionally pages 6-7). This absence could reflect library coverage, essentiality preventing recovery, annotation differences, or that aroA is discussed elsewhere (e.g., supplement) not retrieved here.

Separately, an aromatic biosynthesis functional study reports targeted phenotypes for pheA and tyrA in the PP1769–PP1770 region and documents aromatic amino-acid rescue patterns, but it likewise does not provide direct aroA/EPSPS phenotypes in the supplied pages (molinahenares2009functionalanalysisof pages 6-7).

3.2 Real-world implementation: KT2440 metabolic engineering leveraging aroA

A 2025 P. putida study optimizing production of the aromatic-derived compound p‑aminobenzoic acid (pABA) explicitly defines aroA as EPSPS (“3‑phosphoshikimate‑1‑carboxylvinyl transferase”) and places it in the shikimate pathway step converting S3P → EPSP (with PEP as donor) (camposmagana2025combinatorialengineeringpinpoints pages 2-4).

Using a Design-of-Experiments (Plackett–Burman) combinatorial expression approach across multiple shikimate-pathway genes, pABA titers ranged from ~2 mg/L to 186.2 mg/L in the initial screen (camposmagana2025combinatorialengineeringpinpoints pages 4-6). In their top strain (S12), pABA reached 185.40 mg/L, and reducing expression of aroA (together with aroK, aroQ, and aroGD146N) back to native levels caused a 39.9% decrease in pABA production (p = 0.001) (camposmagana2025combinatorialengineeringpinpoints pages 8-9). This provides KT2440-specific functional evidence that aroA expression level contributes measurably to aromatic-pathway flux toward a chorismate-derived product under engineered conditions.

4) Recent developments and latest research (prioritizing 2023–2024)

4.1 2024: Directed evolution and structure-guided improvement of EPSPS function under glyphosate

A 2024 PNAS study developed a synthetic yeast selection system that enables simultaneous selection for glyphosate tolerance and retained/improved catalytic efficiency of EPSPS variants (reed2024evolvingdualtraitepsp pages 1-2). The study reports recovery of a mutant enzyme with Ki near 1 mM for glyphosate and approximately 2.5-fold improved enzymatic efficiency relative to the starting enzyme (reed2024evolvingdualtraitepsp pages 1-2). This work illustrates a modern trend: treating EPSPS as an engineerable biocatalyst where the classic “resistance vs activity” tradeoff can be mitigated with selection design and structural interpretation (reed2024evolvingdualtraitepsp pages 1-2, reed2024evolvingdualtraitepsp pages 5-6).

4.2 2024: Alternative enzymology for EPSP formation (MurA promiscuity) and glyphosate tolerance

A second 2024 PNAS study reports that MurA (canonically involved in peptidoglycan biosynthesis) can catalyze the same net reaction as EPSPS (S3P + PEP → EPSP + Pi) in the bryophyte Marchantia polymorpha (caygill2024muracatalyzedsynthesisof pages 1-2). Enzyme assays showed MurA activity on S3P/PEP was ~100-fold lower than EPSPS, but ~8-fold higher than MurA’s activity on its canonical UDP-GlcNAc/PEP substrate pair (caygill2024muracatalyzedsynthesisof pages 6-7). Genetic and heterologous-expression evidence linked this alternative activity to glyphosate tolerance (caygill2024muracatalyzedsynthesisof pages 1-2). Although this is not a KT2440 result, it expands the conceptual landscape for “EPSP-forming enzymes,” which is relevant when interpreting resistance and evolutionary possibilities.

4.3 2024: Shikimate pathway review consolidating current mechanistic understanding

A 2024 Natural Product Reports review synthesizes EPSPS mechanism, inhibitor binding, structural information (including E. coli EPSPS complex with glyphosate; PDB 1G6S), and classification of enzyme classes relevant to predicting glyphosate sensitivity in different organisms (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 50-64). This is a high-authority source for current definitions and mechanistic consensus.

5) Expert opinions/authoritative synthesis and resistance framework

A domain-leading 2021 Environmental Microbiology review frames bacterial glyphosate resistance as arising via four broad mechanisms: (i) reduced EPSPS sensitivity or increased EPSPS production, (ii) degradation of glyphosate, (iii) detoxification/modification, and (iv) altered transport (reduced uptake/increased export) (hertel2021molecularmechanismsunderlying pages 1-5). This review provides concrete examples of aroA/EPSPS-associated target-site resistance (e.g., Pro101Ser; Gly96Ala) and emphasizes that overproduction (including amplification or promoter-up changes) can effectively titrate glyphosate (hertel2021molecularmechanismsunderlying pages 5-8). It also highlights non-target routes including transport and enzymatic modification (e.g., N-acetylation) that prevent EPSPS inhibition (hertel2021molecularmechanismsunderlying pages 12-15).

Quantitatively, the review reports laboratory selection of Salmonella aroA mutants under 0.35–2 g/L glyphosate, illustrating the selection pressures under which target-site changes can arise (hertel2021molecularmechanismsunderlying pages 5-8).

6) Applications and real-world implementations

  1. Industrial/biotech strain engineering of aromatic products: In P. putida KT2440, adjusting expression of shikimate-pathway genes including aroA is used to improve yields of chorismate-derived products such as pABA (camposmagana2025combinatorialengineeringpinpoints pages 2-4, camposmagana2025combinatorialengineeringpinpoints pages 8-9). This exemplifies aroA’s practical role as a flux-controlling node in aromatic anabolic pathways.

  2. Glyphosate tolerance engineering: EPSPS variants (often bacterial-derived such as Agrobacterium CP4 EPSPS) are historically foundational for glyphosate-tolerant crops; modern 2024 work continues to refine EPSPS variants to improve both resistance and activity using high-throughput selection and structural analysis (hertel2021molecularmechanismsunderlying pages 5-8, reed2024evolvingdualtraitepsp pages 1-2).

  3. Environmental/clinical microbiology implications: Because glyphosate targets EPSPS, environmental exposures can select for bacterial resistance via multiple mechanisms (target-site, transport, detoxification), potentially intersecting with broader stress-adaptation and resistance landscapes (hertel2021molecularmechanismsunderlying pages 1-5).

7) Limitations of the current evidence base for this specific UniProt protein

Despite extensive general knowledge on aroA/EPSPS, direct experimental characterization of UniProt Q88M05 in P. putida KT2440 (purified enzyme kinetics, substrate specificity beyond the canonical S3P/PEP reaction, structure, or explicit knockout essentiality) was not found in the retrieved KT2440 primary literature segments. The report therefore relies on: (i) authoritative EPSPS mechanism reviews (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 50-64), and (ii) KT2440-specific engineering phenotypes involving aroA expression (camposmagana2025combinatorialengineeringpinpoints pages 8-9), while explicitly flagging KT2440 annotation inconsistencies (molinahenares2009functionalanalysisof pages 2-4).

Key cited sources (with publication dates and URLs)

  • Shende VV, Bauman KD, Moore BS. The shikimate pathway: gateway to metabolic diversity. Natural Product Reports. Jan 2024. https://doi.org/10.1039/d3np00037k (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 50-64)
  • Reed KB et al. Evolving dual-trait EPSP synthase variants using a synthetic yeast selection system. PNAS. Aug 2024. https://doi.org/10.1073/pnas.2317027121 (reed2024evolvingdualtraitepsp pages 1-2, reed2024evolvingdualtraitepsp pages 5-6)
  • Caygill S et al. MurA-catalyzed synthesis of EPSP confers glyphosate tolerance in bryophytes. PNAS. Nov 2024. https://doi.org/10.1073/pnas.2412997121 (caygill2024muracatalyzedsynthesisof pages 1-2, caygill2024muracatalyzedsynthesisof pages 6-7)
  • Campos‑Magaña MA et al. Combinatorial engineering pinpoints shikimate pathway bottlenecks in pABA production in Pseudomonas putida. Journal of Biological Engineering. Sep 2025. https://doi.org/10.1186/s13036-025-00553-5 (camposmagana2025combinatorialengineeringpinpoints pages 2-4, camposmagana2025combinatorialengineeringpinpoints pages 8-9)
  • Hertel R et al. Molecular mechanisms underlying glyphosate resistance in bacteria. Environmental Microbiology. Jun 2021. https://doi.org/10.1111/1462-2920.15534 (hertel2021molecularmechanismsunderlying pages 1-5, hertel2021molecularmechanismsunderlying pages 5-8, hertel2021molecularmechanismsunderlying pages 12-15)
  • Molina‑Henares MA et al. Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440. Microbial Biotechnology. Dec 2009. https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 6-7)
  • Molina‑Henares MA et al. Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium… Environmental Microbiology. Jun 2010. https://doi.org/10.1111/j.1462-2920.2010.02166.x (molina‐henares2010identificationofconditionally pages 2-3, molina‐henares2010identificationofconditionally pages 6-7)

References

  1. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  2. (shende2024theshikimatepathway pages 10-11): Vikram V. Shende, Katherine D. Bauman, and Bradley S. Moore. The shikimate pathway: gateway to metabolic diversity. Natural product reports, 41:604-648, Jan 2024. URL: https://doi.org/10.1039/d3np00037k, doi:10.1039/d3np00037k. This article has 173 citations and is from a peer-reviewed journal.

  3. (shende2024theshikimatepathway pages 50-64): Vikram V. Shende, Katherine D. Bauman, and Bradley S. Moore. The shikimate pathway: gateway to metabolic diversity. Natural product reports, 41:604-648, Jan 2024. URL: https://doi.org/10.1039/d3np00037k, doi:10.1039/d3np00037k. This article has 173 citations and is from a peer-reviewed journal.

  4. (camposmagana2025combinatorialengineeringpinpoints pages 2-4): Marco A Campos-Magaña, Sara Moreno-Paz, Maria Martin-Pascual, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering pinpoints shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. Journal of Biological Engineering, Sep 2025. URL: https://doi.org/10.1186/s13036-025-00553-5, doi:10.1186/s13036-025-00553-5. This article has 0 citations and is from a peer-reviewed journal.

  5. (camposmagana2025combinatorialengineeringpinpoints pages 4-6): Marco A Campos-Magaña, Sara Moreno-Paz, Maria Martin-Pascual, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering pinpoints shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. Journal of Biological Engineering, Sep 2025. URL: https://doi.org/10.1186/s13036-025-00553-5, doi:10.1186/s13036-025-00553-5. This article has 0 citations and is from a peer-reviewed journal.

  6. (camposmagana2025combinatorialengineeringpinpoints pages 8-9): Marco A Campos-Magaña, Sara Moreno-Paz, Maria Martin-Pascual, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering pinpoints shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. Journal of Biological Engineering, Sep 2025. URL: https://doi.org/10.1186/s13036-025-00553-5, doi:10.1186/s13036-025-00553-5. This article has 0 citations and is from a peer-reviewed journal.

  7. (reed2024evolvingdualtraitepsp pages 1-2): Kevin B. Reed, Wantae Kim, Hongyuan Lu, Clayton T. Larue, Shirley Guo, Sierra M. Brooks, Michael R. Montez, James M. Wagner, Y. Jessie Zhang, and Hal S. Alper. Evolving dual-trait epsp synthase variants using a synthetic yeast selection system. Proceedings of the National Academy of Sciences of the United States of America, Aug 2024. URL: https://doi.org/10.1073/pnas.2317027121, doi:10.1073/pnas.2317027121. This article has 7 citations and is from a highest quality peer-reviewed journal.

  8. (caygill2024muracatalyzedsynthesisof pages 1-2): Samuel Caygill, Thomas Köcher, and Liam Dolan. Mura-catalyzed synthesis of 5-enolpyruvylshikimate-3-phosphate confers glyphosate tolerance in bryophytes. Proceedings of the National Academy of Sciences of the United States of America, Nov 2024. URL: https://doi.org/10.1073/pnas.2412997121, doi:10.1073/pnas.2412997121. This article has 10 citations and is from a highest quality peer-reviewed journal.

  9. (caygill2024muracatalyzedsynthesisof pages 6-7): Samuel Caygill, Thomas Köcher, and Liam Dolan. Mura-catalyzed synthesis of 5-enolpyruvylshikimate-3-phosphate confers glyphosate tolerance in bryophytes. Proceedings of the National Academy of Sciences of the United States of America, Nov 2024. URL: https://doi.org/10.1073/pnas.2412997121, doi:10.1073/pnas.2412997121. This article has 10 citations and is from a highest quality peer-reviewed journal.

  10. (hertel2021molecularmechanismsunderlying pages 1-5): Robert Hertel, Johannes Gibhardt, Marion Martienssen, Ramona Kuhn, and Fabian M. Commichau. Molecular mechanisms underlying glyphosate resistance in bacteria. Jun 2021. URL: https://doi.org/10.1111/1462-2920.15534, doi:10.1111/1462-2920.15534. This article has 67 citations and is from a domain leading peer-reviewed journal.

  11. (hertel2021molecularmechanismsunderlying pages 24-27): Robert Hertel, Johannes Gibhardt, Marion Martienssen, Ramona Kuhn, and Fabian M. Commichau. Molecular mechanisms underlying glyphosate resistance in bacteria. Jun 2021. URL: https://doi.org/10.1111/1462-2920.15534, doi:10.1111/1462-2920.15534. This article has 67 citations and is from a domain leading peer-reviewed journal.

  12. (hertel2021molecularmechanismsunderlying pages 5-8): Robert Hertel, Johannes Gibhardt, Marion Martienssen, Ramona Kuhn, and Fabian M. Commichau. Molecular mechanisms underlying glyphosate resistance in bacteria. Jun 2021. URL: https://doi.org/10.1111/1462-2920.15534, doi:10.1111/1462-2920.15534. This article has 67 citations and is from a domain leading peer-reviewed journal.

  13. (hertel2021molecularmechanismsunderlying pages 12-15): Robert Hertel, Johannes Gibhardt, Marion Martienssen, Ramona Kuhn, and Fabian M. Commichau. Molecular mechanisms underlying glyphosate resistance in bacteria. Jun 2021. URL: https://doi.org/10.1111/1462-2920.15534, doi:10.1111/1462-2920.15534. This article has 67 citations and is from a domain leading peer-reviewed journal.

  14. (molina‐henares2010identificationofconditionally pages 2-3): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

  15. (molina‐henares2010identificationofconditionally pages 6-7): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

  16. (molinahenares2009functionalanalysisof pages 6-7): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  17. (reed2024evolvingdualtraitepsp pages 5-6): Kevin B. Reed, Wantae Kim, Hongyuan Lu, Clayton T. Larue, Shirley Guo, Sierra M. Brooks, Michael R. Montez, James M. Wagner, Y. Jessie Zhang, and Hal S. Alper. Evolving dual-trait epsp synthase variants using a synthetic yeast selection system. Proceedings of the National Academy of Sciences of the United States of America, Aug 2024. URL: https://doi.org/10.1073/pnas.2317027121, doi:10.1073/pnas.2317027121. This article has 7 citations and is from a highest quality peer-reviewed journal.

Artifacts

Citations

  1. molinahenares2009functionalanalysisof pages 2-4
  2. shende2024theshikimatepathway pages 10-11
  3. shende2024theshikimatepathway pages 50-64
  4. reed2024evolvingdualtraitepsp pages 1-2
  5. molinahenares2009functionalanalysisof pages 6-7
  6. camposmagana2025combinatorialengineeringpinpoints pages 2-4
  7. camposmagana2025combinatorialengineeringpinpoints pages 4-6
  8. camposmagana2025combinatorialengineeringpinpoints pages 8-9
  9. caygill2024muracatalyzedsynthesisof pages 1-2
  10. caygill2024muracatalyzedsynthesisof pages 6-7
  11. hertel2021molecularmechanismsunderlying pages 1-5
  12. hertel2021molecularmechanismsunderlying pages 5-8
  13. hertel2021molecularmechanismsunderlying pages 12-15
  14. hertel2021molecularmechanismsunderlying pages 24-27
  15. reed2024evolvingdualtraitepsp pages 5-6
  16. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  17. https://doi.org/10.1039/d3np00037k
  18. https://doi.org/10.1186/s13036-025-00553-5
  19. https://doi.org/10.1073/pnas.2317027121
  20. https://doi.org/10.1073/pnas.2412997121
  21. https://doi.org/10.1111/1462-2920.15534
  22. https://doi.org/10.1111/j.1462-2920.2010.02166.x
  23. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  24. https://doi.org/10.1039/d3np00037k,
  25. https://doi.org/10.1186/s13036-025-00553-5,
  26. https://doi.org/10.1073/pnas.2317027121,
  27. https://doi.org/10.1073/pnas.2412997121,
  28. https://doi.org/10.1111/1462-2920.15534,
  29. https://doi.org/10.1111/j.1462-2920.2010.02166.x,

📄 View Raw YAML

id: Q88M05
gene_symbol: aroA
product_type: PROTEIN
status: COMPLETE
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: >-
  aroA (PP_1770) of Pseudomonas putida KT2440 is a 746-residue bifunctional
  cytoplasmic enzyme of aromatic amino acid biosynthesis. Its EPSP synthase
  module (3-phosphoshikimate 1-carboxyvinyltransferase, EC 2.5.1.19;
  5-enolpyruvylshikimate-3-phosphate synthase) catalyzes the penultimate step of
  the shikimate pathway, transferring the enolpyruvyl moiety of
  phosphoenolpyruvate to the 5-hydroxyl of shikimate-3-phosphate to yield
  5-enolpyruvylshikimate-3-phosphate (EPSP) plus inorganic phosphate; EPSP is
  then converted to chorismate, the branch-point precursor of phenylalanine,
  tyrosine, tryptophan, folate, ubiquinone, and other aromatic metabolites. In
  addition to the canonical EPSP synthase domain (Pfam EPSP_synthase; COG0128;
  TIGR01356 aroA), the protein carries an N-terminal prephenate/arogenate
  dehydrogenase (TyrA) module (Pfam PDH_N/PDH_C; COG0287) with a NAD(P)-binding
  Rossmann fold. UniProt annotates a second catalytic activity for this module,
  prephenate dehydrogenase (prephenate + NAD+ -> 4-hydroxyphenylpyruvate + CO2 +
  NADH, EC 1.3.1.12), placing it in the tyrosine-biosynthetic conversion of
  prephenate to 4-hydroxyphenylpyruvate. The protein is thus a fused
  EPSP-synthase / prephenate-dehydrogenase enzyme contributing to both chorismate
  formation and downstream L-tyrosine biosynthesis. The shikimate pathway is
  absent in animals, making EPSP synthase the molecular target of the herbicide
  glyphosate, a competitive inhibitor at the PEP site.
existing_annotations:
- term:
    id: GO:0003824
    label: catalytic activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: Root-level catalytic activity term; uninformative given the specific enzymatic activities annotated below.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      GO:0003824 is the top-level molecular-function catalytic term and conveys
      no specific information. The protein has well-supported specific activities
      (EPSP synthase, EC 2.5.1.19; prephenate dehydrogenase, EC 1.3.1.12) that
      should be used instead.
- term:
    id: GO:0003866
    label: 3-phosphoshikimate 1-carboxyvinyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: EPSP synthase activity (EC 2.5.1.19); the canonical, core molecular function of aroA.
    action: ACCEPT
    reason: >-
      Directly supported by sequence/domain evidence: the EPSP synthase Pfam
      domain (PF00275), HAMAP rule MF_00210, COG0128, NCBIfam TIGR01356 (aroA),
      conserved PEP and shikimate-3-phosphate binding residues, and mapping to
      Rhea:21256 / EC 2.5.1.19. This is the defining function of aroA.
- term:
    id: GO:0004665
    label: prephenate dehydrogenase (NADP+) activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: NADP+-dependent prephenate dehydrogenase activity inferred from the fused TyrA domain. The cofactor specificity (NADP+ vs NAD+) is not experimentally established for this protein.
    action: KEEP_AS_NON_CORE
    reason: >-
      The protein carries a genuine N-terminal prephenate/arogenate dehydrogenase
      (TyrA) module (Pfam PDH_N/PDH_C; COG0287), so prephenate dehydrogenase
      activity is a plausible second function. However, UniProt's curated
      CATALYTIC ACTIVITY block lists only the NAD+ route (EC 1.3.1.12), and the
      NADP+ specificity here is purely an InterPro electronic inference
      (IPR003099) with no cofactor evidence. Retain as a non-core, lower-confidence
      activity rather than a core function.
- term:
    id: GO:0005737
    label: cytoplasm
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: located_in
  review:
    summary: Cytoplasmic localization, consistent with a soluble shikimate-pathway metabolic enzyme.
    action: ACCEPT
    reason: >-
      EPSP synthase is a soluble cytosolic enzyme of central aromatic amino acid
      biosynthesis; cytoplasmic localization is supported by UniProt-SubCell
      (SL-0086) and HAMAP rule MF_00210, with no signal/transmembrane features.
- term:
    id: GO:0006571
    label: L-tyrosine biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: Tyrosine biosynthesis; supported via the fused prephenate dehydrogenase (TyrA) domain that converts prephenate to 4-hydroxyphenylpyruvate.
    action: KEEP_AS_NON_CORE
    reason: >-
      The TyrA (prephenate dehydrogenase) module places this protein in the
      tyrosine-specific branch (prephenate -> 4-hydroxyphenylpyruvate, UniPathway
      step 1/1 of the NAD+ route). This is a real but secondary process relative
      to the core EPSP synthase / chorismate-biosynthesis role, hence non-core.
- term:
    id: GO:0008652
    label: amino acid biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000104
  qualifier: involved_in
  review:
    summary: General amino acid biosynthetic process; correct but non-specific given the more precise aromatic/chorismate terms.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      True but high-level. The more specific processes (chorismate biosynthetic
      process, aromatic amino acid biosynthetic process, L-tyrosine biosynthetic
      process) capture the role precisely, making this generic parent redundant.
- term:
    id: GO:0008977
    label: prephenate dehydrogenase (NAD+) activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: NAD+-dependent prephenate dehydrogenase activity from the fused TyrA domain; the second catalytic function of this bifunctional protein.
    action: KEEP_AS_NON_CORE
    reason: >-
      Supported by the prephenate/arogenate dehydrogenase domain (residues
      ~14-302; Pfam PDH_N/PDH_C; COG0287) and by UniProt's curated CATALYTIC
      ACTIVITY block citing the NAD+ reaction (Rhea:13869, EC 1.3.1.12) and the
      L-tyrosine biosynthesis (NAD+ route) pathway. This is the better-supported
      of the two prephenate dehydrogenase cofactor variants but remains the
      secondary (non-core) function relative to EPSP synthase.
- term:
    id: GO:0009073
    label: aromatic amino acid biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: Aromatic amino acid biosynthesis; accurate at the family level for an EPSP-synthase / chorismate-pathway enzyme also feeding tyrosine biosynthesis.
    action: ACCEPT
    reason: >-
      Both functional modules act within aromatic amino acid biosynthesis: EPSP
      synthase produces the chorismate precursor common to Phe/Tyr/Trp, and the
      TyrA domain feeds the tyrosine branch. The term is appropriately specific.
- term:
    id: GO:0009423
    label: chorismate biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: Chorismate biosynthesis; the core biological process of the EPSP synthase activity (penultimate shikimate-pathway step).
    action: ACCEPT
    reason: >-
      EPSP synthase catalyzes step 6/7 of chorismate biosynthesis from
      D-erythrose-4-phosphate and PEP (UniPathway UPA00053/UER00089). This is the
      most precise and well-supported biological-process term for the core
      function. Consistent with KT2440-specific metabolic-engineering evidence that
      tuning aroA expression contributes to flux through the shikimate pathway
      toward chorismate-derived products (see aroA-deep-research-falcon.md,
      citing PMID:41029715).
- term:
    id: GO:0016491
    label: oxidoreductase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000104
  qualifier: enables
  review:
    summary: Generic oxidoreductase parent term covering the prephenate dehydrogenase activity.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      Redundant high-level parent of the specific prephenate dehydrogenase
      activities (GO:0008977 / GO:0004665) already annotated. Provides no
      additional information beyond the specific terms.
- term:
    id: GO:0016628
    label: oxidoreductase activity, acting on the CH-CH group of donors, NAD or NADP as acceptor
  evidence_type: IEA
  original_reference_id: GO_REF:0000117
  qualifier: enables
  review:
    summary: Intermediate oxidoreductase-class parent term for the prephenate dehydrogenase activity.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      A grouping parent of the specific prephenate dehydrogenase (NAD+/NADP+)
      activities. The leaf terms GO:0008977 / GO:0004665 are retained, so this
      mid-level class is redundant over-annotation.
- term:
    id: GO:0016765
    label: transferase activity, transferring alkyl or aryl (other than methyl) groups
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: Generic enolpyruvyl/alkyl transferase parent term for the EPSP synthase activity.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      High-level parent of the specific EPSP synthase activity (GO:0003866,
      enolpyruvyl transferase) which is retained. Redundant given the leaf term.
- term:
    id: GO:0070403
    label: NAD+ binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: NAD+ cofactor binding by the Rossmann-fold prephenate dehydrogenase (TyrA) domain.
    action: KEEP_AS_NON_CORE
    reason: >-
      Consistent with the NAD(P)-binding Rossmann fold of the fused TyrA domain
      (InterPro IPR046826) and with the NAD+-dependent prephenate dehydrogenase
      activity. A supporting cofactor-binding term for the secondary activity, so
      non-core rather than a primary function.
core_functions:
- description: >-
    EPSP synthase (3-phosphoshikimate 1-carboxyvinyltransferase) catalyzing the
    penultimate step of the shikimate pathway, the chorismate-yielding branch of
    aromatic amino acid biosynthesis.
  molecular_function:
    id: GO:0003866
    label: 3-phosphoshikimate 1-carboxyvinyltransferase activity
  supported_by:
  - reference_id: file:PSEPK/aroA/aroA-uniprot.txt
    supporting_text: >-
      Catalyzes the transfer of the enolpyruvyl moiety of phosphoenolpyruvate
      (PEP) to the 5-hydroxyl of shikimate-3-phosphate (S3P) to produce
      enolpyruvyl shikimate-3-phosphate and inorganic phosphate; EC 2.5.1.19;
      EPSP synthase family; chorismate biosynthesis step 6/7.
  - reference_id: file:PSEPK/aroA/aroA-deep-research-falcon.md
    supporting_text: >-
      aroA is annotated as 3-phosphoshikimate 1-carboxyvinyltransferase / EPSP
      synthase (EC 2.5.1.19) catalyzing S3P + PEP -> EPSP + Pi, the penultimate
      step of the shikimate pathway leading to chorismate, the common precursor
      for Phe/Tyr/Trp biosynthesis.
  directly_involved_in:
  - id: GO:0009423
    label: chorismate biosynthetic process
# NOTE: The fused N-terminal prephenate/arogenate dehydrogenase (TyrA) module
# (prephenate dehydrogenase NAD+ activity, GO:0008977; L-tyrosine biosynthetic
# process, GO:0006571) is a genuine but secondary, non-core activity of this
# bifunctional protein and is reviewed as KEEP_AS_NON_CORE in existing_annotations
# rather than listed here as a core function.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000104
  title: Electronic Gene Ontology annotations created by transferring manual GO annotations between related proteins based on shared sequence features
  findings: []
- id: GO_REF:0000117
  title: Electronic Gene Ontology annotations created by ARBA machine learning models
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: file:PSEPK/aroA/aroA-deep-research-falcon.md
  title: Deep research report (falcon) for aroA / EPSP synthase (Q88M05) in P. putida KT2440
  findings:
  - statement: Synthesizes EPSPS mechanism, shikimate-pathway/chorismate role, cytosolic localization inference, and KT2440-specific engineering evidence for aroA.
- id: PMID:12534463
  title: Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.
  findings:
  - statement: Genome sequence of P. putida KT2440 in which PP_1770 (aroA, Q88M05) is annotated, establishing the locus and organism context for this gene.
    reference_section_type: RESULTS
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: >-
      PubMed-verified genome paper (Nelson et al., Environ Microbiol 2002) cited
      in the UniProt entry as the source for the PP_1770 locus. Establishes
      organism/locus, not direct enzymatic characterization.
- id: PMID:41029715
  title: Combinatorial engineering pinpoints shikimate pathway bottlenecks in para-aminobenzoic acid production in Pseudomonas putida.
  findings:
  - statement: In P. putida KT2440 metabolic engineering for para-aminobenzoic acid (a chorismate-derived product), aroA (EPSPS) was tuned among shikimate-pathway genes; reducing aroA expression to native levels lowered product titer, indicating aroA expression contributes to aromatic-pathway flux in vivo.
    reference_section_type: RESULTS
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: >-
      PMID confirmed via PubMed search (single match to title keywords +
      organism). KT2440-specific engineering study; supports aroA as a
      flux-relevant shikimate-pathway node rather than providing direct
      purified-enzyme kinetics. Surfaced through the falcon deep-research summary.
suggested_questions:
- question: Is the prephenate dehydrogenase (TyrA) module of P. putida AroA catalytically active in vivo, and does it prefer NAD+ or NADP+ as cofactor?
- question: Does the AroA-TyrA domain fusion form a substrate channel or otherwise functionally couple chorismate biosynthesis with the tyrosine branch in P. putida?
suggested_experiments:
- description: Heterologously express and purify Q88M05 and assay both EPSP synthase (S3P + PEP) and prephenate dehydrogenase (prephenate + NAD+/NADP+) activities to confirm bifunctionality and determine cofactor preference.
- description: Construct an aroA deletion/complementation in P. putida KT2440 and test for aromatic amino acid (and specifically tyrosine) auxotrophy to establish in vivo requirement of each catalytic module.