trpE

UniProt ID: Q88QS1
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

Anthranilate synthase component I (TrpE, locus PP_0417), the large alpha subunit of anthranilate synthase (EC 4.1.3.27). Together with the glutamine amidotransferase beta subunit TrpG, it forms a heterotetrameric complex that catalyzes the first committed step of L-tryptophan biosynthesis, the conversion of chorismate to anthranilate. TrpE binds chorismate and performs the amination/lyase chemistry, using ammonia supplied by TrpG from hydrolysis of L-glutamine; the products are anthranilate, pyruvate and L-glutamate. In the absence of TrpG, TrpE alone can produce anthranilate directly from chorismate when free ammonia is abundant. The enzyme requires Mg2+ and is a soluble, cytoplasmic enzyme of aromatic amino acid primary metabolism. Anthranilate synthase is typically feedback-inhibited by L-tryptophan, the pathway end product. In P. putida KT2440, loss of trpE causes tryptophan auxotrophy that is rescued by anthranilate, indole or tryptophan, confirming its placement at the anthranilate-forming step.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000162 L-tryptophan biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: TrpE catalyzes the first committed step (chorismate to anthranilate) of de novo L-tryptophan biosynthesis. This is the correct, well-supported biological process for this enzyme.
Reason: Anthranilate synthase component I is the entry enzyme of the tryptophan branch of aromatic amino acid biosynthesis. The IEA assignment (InterPro IPR005256, UniPathway UPA00035) is corroborated by experimental genetics in KT2440, where a trpE (PP_0417) insertion mutant is a tryptophan auxotroph rescued by anthranilate, indole or tryptophan (PMID:21261884; see file:PSEPK/trpE/trpE-deep-research-falcon.md).
GO:0004049 anthranilate synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: TrpE is the synthase (alpha) component of anthranilate synthase (EC 4.1.3.27), catalyzing chorismate + L-glutamine to anthranilate + pyruvate + L-glutamate (RHEA:21732). This is the core molecular function and is correct.
Reason: Assigned from InterPro IPR005256, RHEA:21732 and EC 4.1.3.27, consistent with the protein family (anthranilate synthase component I), the Pfam chorismate-binding domain, and the UniProt catalytic activity annotation. GO:0004049 is the precise molecular function term.

Core Functions

Anthranilate synthase component I activity - binds chorismate and, using ammonia supplied by the TrpG glutaminase subunit (or free ammonia at high concentration), converts chorismate to anthranilate with release of pyruvate, the first committed step of L-tryptophan biosynthesis.

Supporting Evidence:
  • GO_REF:0000120
    Anthranilate synthase component 1; EC 4.1.3.27; Reaction=chorismate + L-glutamine = anthranilate + pyruvate + L-glutamate + H(+); Rhea:RHEA:21732.
  • PMID:21261884
    A trpE mutant did not grow on minimal medium but growth was restored by anthranilate, indole, or tryptophan supplementation, placing TrpE (PP_0417) at the anthranilate-forming step of tryptophan biosynthesis.

References

Combined Automated Annotation using Multiple IEA Methods
Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440
  • A mini-Tn5 insertion in trpE (PP_0417) produces a tryptophan auxotroph in KT2440 rescued by anthranilate, indole or tryptophan, placing TrpE at the anthranilate-forming step of tryptophan biosynthesis; trpE is a monocistronic unit separate from the trpGDC operon.
    "A trpE mutant did not grow on minimal medium but growth was restored by anthranilate, indole, or tryptophan supplementation."

Deep Research

Asta

(trpE-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-05T20:16:38.578088

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.776)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.685)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[2] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.769)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[3] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.732)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[4] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.726)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[5] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.724)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[6] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.720)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[7] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[8] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.704)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[9] The alpha-ketoacid dehydrogenase complexes of Drosophila melanogaster.

  • Authors: Steven J. Marygold
  • Year: 2024
  • Venue: microPublication Biology
  • URL: https://www.semanticscholar.org/paper/50942e603e0e14ee9195c0d7cb52db11a521f964
  • DOI: 10.17912/micropub.biology.001209
  • PMID: 38741935
  • PMCID: 11089389
  • Citations: 2
  • Summary: This work identifies and classify the genes encoding all Drosophila AKDHC subunits, update their functional annotations and integrate this work into the FlyBase database.
  • Evidence snippets:
  • Snippet 1 (score: 0.701)
    > Symbol: gene symbol in FlyBase -asterisk (*) indicates a gene with testis-specific expression. CG#: gene annotation ID in FlyBase. Function: component and associated EC number (where available/applicable). Key reference(s) for initial identification or genetic characterization (in a metabolic context): 1. Gruntenko et al. 1998;2. Chen et al. 2008;3. Yoon et al. 2017;4. Yap et al. 2021a;5. Yap et al. 2021b;6. Whittle et al. 2023;7. González Morales et al. 2023;8. Homem et al. 2014;9. Bonnay et al. 2020;10. Ivanova et al. 2004;11. Boyko et al. 2020;12. Liu et al. 2017;13. Li et al. 2020;14. Devilliers et al. 2021;15. Goyal et al. 2022;16. Huang et al. 2022;17. Plaçais et al. 2017;18. Dung et al. 2018;19. Rabah et al. 2023;20. Klenz et al. 1995;21. Katsube et al. 1997;22. Gándara et al. 2019;23. Lambrechts et al. 2019;24. Lee et al. 2022;25. Chen et al. 2006;26. Kim et al. 2023;27. Tsai et al. 2020. Human ortholog: gene symbol at HGNC, with % amino acid identity between the encoded protein and the Drosophila protein. Human disease: OMIM symbol for disease(s) associated with the human gene (Amberger et al. 2019) -see Extended Data File 1 for details.

[10] RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications

  • Authors: Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al.
  • Year: 2025
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  • DOI: 10.1093/nar/gkaf1206
  • PMID: 41206768
  • PMCID: 12807602
  • Citations: 2
  • Summary: The new RM2Target v2.0 will serve as a foundational resource for exploring RNA epitranscriptomic regulation, enabling investigations into cross-talk among modifications, underlying molecular mechanisms, and disease connections, thereby facilitating both basic research and translational applications in RNA epigenetics.
  • Evidence snippets:
  • Snippet 1 (score: 0.698)
    > To obtain basic information on WERs and their target genes, such as official gene symbols, gene IDs, gene types, and genomic locations, gene annotations were downloaded from the GENCODE project [ 44 ] for human and mouse, and from NCBI [ 45 ] and Ensembl [ 46 ] for the other species. Genomic locations were extracted from the corresponding GTF annotation files. Gene symbols were primarily standardized based on the NCBI Gene database [ 45 ] for mRNAs and lncRNAs, GtR-NAdb [ 47 ] for tRNAs, miRbase [ 48 ] for microRNAs, and cir-cBase [ 49 ] for circRNAs. Deprecated or substituted versions of genes were filtered out. The LiftOver [ 50 ] program was employed to convert and unify genomic coordinates across different genome assembly versions.
    > The functional descriptions of WERs were compiled based on the UniProt database [ 51 ] and further supplemented with evidence from relevant publications, with particular emphasis on their functions as RNA modification regulatory proteins.

[11] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.696)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[12] Integrated genomic-transcriptomic analysis of clavulanic acid production in differentially productive Streptomyces clavuligerus strains

  • Authors: J. Gong, Jeong Sang Yi, Seungchan An, Hang Su Cho, Chang Hun Shin et al.
  • Year: 2025
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/b4903d3729bba93d1d47e38f3353a26f3530a8dd
  • DOI: 10.1038/s41598-025-29509-x
  • PMID: 41310174
  • PMCID: 12749726
  • Citations: 1
  • Summary: Findings include large plasmid deletions, an enrichment of mutations in secondary metabolite biosynthesis and regulatory genes, and metabolic shifts redirecting amino acid and carbon flux toward CA biosynthetic pathways.
  • Evidence snippets:
  • Snippet 1 (score: 0.692)
    > Gene annotation was primarily derived from the S. clavuligerus reference genome in the NCBI database and was annotated using the NCBI Prokaryotic Genome Annotation Pipeline 59 . However, several CA biosynthetic genes were manually corrected based on published literature 9 . For instance, two loci were originally annotated as clavaminate synthase 1 (cas1), but one of these loci is located near the cephamycin C biosynthetic cluster, indicating it was actually cas2. Following this correction, the RefSeq accession numbers of all genes in the reference genome were cross-referenced with the UniProt database to obtain additional annotations 60 . For the mutated genes identified through ICA, protein existence levels were manually assigned based on the UniProt data, including protein existence status, annotation score, similar proteins, and relevant publications.

[13] The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize

  • Authors: Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al.
  • Year: 2016
  • Venue: BMC Systems Biology
  • URL: https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  • DOI: 10.1186/s12918-016-0369-x
  • PMID: 27899149
  • PMCID: 5129634
  • Citations: 11
  • Influential citations: 1
  • Summary: CornCyc’s computational predictions are more accurate than those in MaizeCyc when compared to experimentally determined function assignments, demonstrating the relative strength of the enzymatic function assignment pipeline used to generate CornCyc.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > A gold standard set of protein functional annotations was generated by extracting data from UniProt [16] and BRENDA [17]. We extracted all protein sequence and annotation data from UniProt (release 2016_05) for the organism Zea mays, keeping the EC annotations only from the manually reviewed component of UniProt, while removing those annotations that had not undergone manual review. We also extracted experimentally verified protein annotations for Zea mays from BRENDA (release 2016.1). The UniProt and BRENDA annotations were then merged by matching proteins based on the database crosslinks provided by BRENDA, resulting in the union of the reviewed annotations from UniProt and the experimentally verified annotations of BRENDA with duplicates removed. The merged protein annotations were then matched to the B73 RefGen_v2 translated gene models using BLASTP based on a sequence identity cutoff of 96% and an e-value cutoff of 1e-20. We selected the top scoring hit for each protein which resulted in matches to 1,815 unique maize proteins. EC annotations for alternate isoforms were consolidated at the gene level, resulting in 1,475 experimentally verified or manually reviewed protein functional annotations across 1,450 maize genes.

[14] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.687)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[15] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.682)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[16] Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser

  • Authors: Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała
  • Year: 2025
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  • DOI: 10.1186/s12864-025-12329-8
  • PMID: 41408139
  • PMCID: 12822150
  • Citations: 2
  • Summary: An R script is developed that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database and provides a single list of updated symbols with annotation about their ambiguity.
  • Evidence snippets:
  • Snippet 1 (score: 0.679)
    > Gene nomenclature contains current official symbols and various numbers of synonyms, which pose a challenge to integrating genomic data and increase the probability that different genes share the same symbol. Therefore, we retrieved identifiers assigned to all protein-coding genes in human, mouse, rat and pig genomes that are available in the Ensembl genome browser (release 113) to assess the number of genes, compare species and identify ambiguous symbols. Results: Our analysis revealed that the total number of symbols, both official symbols and synonyms, used to identify protein-coding genes ranges from 16,600 in pigs to 64,580 in mice. Furthermore, the gene nomenclature is not complete because there are also genes without an assigned symbol, which indicates gaps in understanding protein-coding genes, especially in pigs. We also found a large number of gene symbols that map to more than one gene. These symbols might complicate the identification of about 10% of rat and mouse genes and 18% of human protein-coding genes. A simple solution for this problem is the usage of stable gene IDs assigned by scientific institutions and committees (Ensembl, NCBI, RGD, HGNC and VGNC) provided that the genomic information associated with these IDs is retrieved directly from proprietary databases containing the most accurate data. Finally, although gene symbols may pose a problem with unequivocal identification of genes, there are instances when no other identifiers are available in the literature. Therefore, we have developed an R script performing search of the Ensembl database and integrating data to provide a single list of updated symbols with annotation about their ambiguity. Conclusions: Gene symbols are not always reliable and should be reported together with stable IDs to enable unequivocal identification of genes. Therefore, data containing only gene symbols should be used cautiously to avoid misidentification of genes. A solution for this problem is our R script REgeness that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database.

[17] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.678)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[18] Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

  • Authors: Kristian Barrett, Cameron J. Hunt, L. Lange, I. Grigoriev, A. Meyer
  • Year: 2023
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/cf508bb4b0c60e0806ee7b9af7440d14c1d31ef2
  • DOI: 10.1093/nar/gkad385
  • PMID: 37216585
  • PMCID: 10320097
  • Citations: 8
  • Summary: The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources My coCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes, allowing users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences.
  • Evidence snippets:
  • Snippet 1 (score: 0.677)
    > To inspect the transcriptomics result of individual genes, proteins originating from JGI have a hyperlink to a summary page which links to a 'Genome browser' page that displays the predicted gene splicing, including which regions have RNA support (Figure 1 D).
    > Hence, in the protein specific site in the JGI w e bsite under 'To Genome Browser', the current protein (GeneCatalog) can be seen together with se v eral other alternati v e predictions of the gene splicing, which is essential to have correct, for the protein to function naturally (Figure 2 ).
    > As the RNA coverage supports the exon / intron splicing, it is possible to infer whether a particular gene, in this case Table 1. Comparison between the dbCAN webserver, the eggNOG webserver for both family and functional annotation of CAZymes and the updated CUPP-w e bserver using the recommended significance cut-off at 5. The column 'CAZy -All characterized' encompasses all 10784 characterized proteins in the CAZy database used for the training, whereas the 'CAZy -Newly characterized ' a gene such as the one selected in Figure 2 , is more likely to work after heterologous gene expression.
    > To further improve the enzyme selection, all NCBI Genbank accessions were mapped to Uniprot ID to link to the specific Uniprot accession page including the predicted Al-phaFold structures, Go annotations and InterPro annotations and more ( 19 ).

[19] Protocol for gene annotation, prediction, and validation of genomic gene expansion

  • Authors: Quanwei Zhang, Zhengdong D. Zhang
  • Year: 2022
  • Venue: STAR Protocols
  • URL: https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3
  • DOI: 10.1016/j.xpro.2022.101692
  • PMID: 36125934
  • PMCID: 9494284
  • Citations: 1
  • Summary: A detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation is described and steps to discover functionality of each copy of replicated genes are detailed.
  • Evidence snippets:
  • Snippet 1 (score: 0.670)
    > 3. Gene annotation and functional annotation. a. Gene structure annotation.
    > In addition to gene prediction models, evidence from orthologous protein sequences and transcriptome assembly could be used to improve annotation quality. Protein sequences of orthologous genes can be obtained from UniProt (The UniProt, 2017). Ones from Swiss-Port have been reviewed and thus are of higher quality. Transcriptome assembly may be available from previous studies or can be assembled de novo from RNA-seq reads by Trinity (Haas et al., 2013). High quality transcriptome assembly can be selected as described in (Zhang et al., 2021). Note: Details about gene structure annotation (Holt and Yandell, 2011) can be found at http:// gmod.org/wiki/MAKER_Tutorial, https://darencard.net/blog/2017-05-16-maker-genomeannotation/, and the protocol (Campbell et al., 2014).
    > b. Quality measurement and functional annotation.
    > For each predicted gene, Maker2 provides the annotation edit distance (AED) score, which measures the goodness of fit between its predicted gene structure and its evidence support. The lower the score, the more accurate the prediction. If more than 90% genes with AED scores lower than 0.5, the genome can be considered well annotated. In addition to the AED score, a high proportion of recognizable domains contained in predicted protein -e.g., higher than 50% -also indicates a good annotation. Recognizable protein domains can by scanned by InterProScan (Jones et al., 2014), assigning potential function to predicted genes.
    > Note: Besides the aforementioned quality measurement, we strongly recommend measuring the completeness of the genome assembly and annotation by checking the existence of a set of Benchmarking Universal Single-Copy Orthologs (BUSCO) (Simao et al., 2015). A high-level completeness of genome assembly and annotation is imperative for a better identification of gene expansion. Based on the result of this analysis, researchers can decide whether they need to further improve the genome assembly before predicting gene expansion. A detailed protocol of BUSCO is available at

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  2. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  3. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  4. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  5. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  6. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  7. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  8. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  9. Steven J. Marygold (2024). The alpha-ketoacid dehydrogenase complexes of Drosophila melanogaster.. microPublication Biology. https://www.semanticscholar.org/paper/50942e603e0e14ee9195c0d7cb52db11a521f964
  10. Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al. (2025). RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications. Nucleic Acids Research. https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  11. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  12. J. Gong, Jeong Sang Yi, Seungchan An, Hang Su Cho, Chang Hun Shin et al. (2025). Integrated genomic-transcriptomic analysis of clavulanic acid production in differentially productive Streptomyces clavuligerus strains. Scientific Reports. https://www.semanticscholar.org/paper/b4903d3729bba93d1d47e38f3353a26f3530a8dd
  13. Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al. (2016). The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize. BMC Systems Biology. https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  14. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  15. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  16. Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała (2025). Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser. BMC Genomics. https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  17. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  18. Kristian Barrett, Cameron J. Hunt, L. Lange, I. Grigoriev, A. Meyer (2023). Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes. Nucleic Acids Research. https://www.semanticscholar.org/paper/cf508bb4b0c60e0806ee7b9af7440d14c1d31ef2
  19. Quanwei Zhang, Zhengdong D. Zhang (2022). Protocol for gene annotation, prediction, and validation of genomic gene expansion. STAR Protocols. https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3

Falcon

(trpE-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 31 citations 2 artifacts 2026-06-11T21:42:19.033265

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: functional annotation of trpE (UniProt Q88QS1; locus PP_0417) in Pseudomonas putida KT2440

0) Target verification (mandatory disambiguation)

The symbol trpE is used across many bacteria, but in Pseudomonas putida KT2440 it specifically corresponds to PP_0417, which is annotated and experimentally supported as anthranilate synthase component I (TrpE), required for tryptophan prototrophy (molinahenares2009functionalanalysisof pages 2-4). A mini‑Tn5 insertion in PP_0417 causes a tryptophan auxotrophy, linking this KT2440 locus unambiguously to TrpE function in tryptophan biosynthesis (molinahenares2009functionalanalysisof pages 2-4).

1) Key concepts and definitions (current understanding)

1.1 Anthranilate synthase (AS): committed step enzyme in tryptophan biosynthesis

Anthranilate synthase (EC 4.1.3.27) catalyzes the first committed step of L‑tryptophan biosynthesis: conversion of chorismate → anthranilate (parthasarathy2018athreeringcircus pages 5-6, naz2023insightintoderegulation pages 11-14). In most bacteria, AS is a two-subunit glutamine amidotransferase complex comprising:
- TrpE: the synthase/“large” (α) subunit, which binds chorismate and performs the amination/lyase chemistry leading to anthranilate and pyruvate formation (parthasarathy2018athreeringcircus pages 5-6, naz2023insightintoderegulation pages 11-14).
- TrpG: the glutaminase/“small” (β) subunit, which hydrolyzes glutamine to produce ammonia, which is then transferred to the TrpE active site (parthasarathy2018athreeringcircus pages 5-6, funke2024validationofaminodeoxychorismate pages 2-5).

1.2 Reaction chemistry, substrates, and products

Mechanistic descriptions in reviews and structural analyses support a pathway in which TrpE/TrpG carries out an amidation/lyase transformation at chorismate:
- Primary substrates/co-substrates: chorismate + glutamine (as ammonia source via TrpG) (parthasarathy2018athreeringcircus pages 5-6, funke2024validationofaminodeoxychorismate pages 2-5).
- Products: anthranilate (precursor of tryptophan), pyruvate (released from chorismate side chain), and glutamate (from glutamine hydrolysis) (parthasarathy2018athreeringcircus pages 5-6, funke2024validationofaminodeoxychorismate pages 2-5).
Two operational modes are described: a glutamine-dependent reaction requiring both subunits and an ammonia-dependent reaction in which the α subunit can use free ammonia under high ammonium availability (naz2023insightintoderegulation pages 14-15).

1.3 Regulation: feedback inhibition by L‑tryptophan

AS is commonly allosterically feedback-inhibited by L‑tryptophan, which reduces flux through the committed step of the pathway (niraula2025aromaticaminoacids pages 13-14, naz2023insightintoderegulation pages 11-14). Structural-level discussions place important determinants of inhibition in/near TrpE domain interfaces and binding crevices coupled to the TrpG interaction region (naz2023insightintoderegulation pages 11-14, naz2023insightintoderegulation pages 14-15).

2) Gene/protein function in Pseudomonas putida KT2440

2.1 Pathway placement and biological role

In KT2440, TrpE (PP_0417) functions upstream of anthranilate in the (essentially one-way) chorismate→tryptophan biosynthetic route (molinahenares2009functionalanalysisof pages 1-2). This is experimentally supported by precursor feeding experiments:
- A trpE mutant did not grow on minimal medium but growth was restored by anthranilate, indole, or tryptophan supplementation (molinahenares2009functionalanalysisof pages 4-6).
This rescue pattern is consistent with TrpE acting at the anthranilate-forming step (and thus upstream of indole and tryptophan) (molinahenares2009functionalanalysisof pages 4-6).

2.2 Operon/transcription organization and genomic context

In KT2440, tryptophan biosynthesis genes are split across loci. In the cluster containing PP_0417:
- trpE (PP_0417) is expressed as a single monocistronic transcript.
- trpG-trpD-trpC form a separate operon (tight spacing/overlaps and RT-PCR confirmation) (molinahenares2009functionalanalysisof pages 2-4).
Other trp genes are in distinct regions: trpA-trpB form an operon transcribed divergently from the repressor gene trpI, and trpF is an unlinked monocistronic unit (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6). This gene organization is shown in the KT2440 trp gene map and RT-PCR schematic (molinahenares2009functionalanalysisof media 744c52d9).

2.3 Protein–protein interactions: TrpE–TrpG interface is essential

AS activity depends on formation of a functional TrpG–TrpE complex, in which ammonia generated by TrpG is channeled to TrpE (funke2024validationofaminodeoxychorismate pages 2-5). A recent experimental validation study (in E. coli with cross-reference to Pseudomonas sequence conservation) identifies a conserved synthase aspartate (ecTrpE D367) interacting with glutaminase residues (ecTrpG Y132, S173) as critical for complex stability and function, indicating the TrpE–TrpG interface is a mechanistically and evolutionarily constrained interaction surface (funke2024validationofaminodeoxychorismate pages 2-5).

2.4 Cellular localization

No KT2440-specific subcellular localization experiments were identified in the retrieved sources. However, anthranilate synthase is described as a soluble glutamine amidotransferase enzyme complex of primary metabolism, consistent with cytosolic localization in bacteria (parthasarathy2018athreeringcircus pages 5-6, funke2024validationofaminodeoxychorismate pages 2-5). This statement is therefore best treated as a strong inference from enzyme class/function rather than a strain-specific localization measurement.

3) Recent developments and latest research (prioritizing 2023–2024)

3.1 2023: structure-informed understanding of feedback inhibition and engineering targets

A 2023 review focusing on deregulation of amino-acid feedback inhibition summarizes structural features and residue-level determinants for AS (TrpE/TrpG), emphasizing: (i) TrpE as the chorismate-binding/anthranilate-forming subunit; (ii) TrpG as the glutamine amidotransferase subunit; and (iii) the importance of conformational states and TrpE residues in tryptophan binding and feedback control (naz2023insightintoderegulation pages 11-14, naz2023insightintoderegulation pages 14-15). This supports current “expert consensus” that TrpE is a key lever for tuning tryptophan/anthranilate flux in strain engineering.

3.2 2024: anthranilate synthase interface as a novel antibiotic target concept

A 2024 Applied and Environmental Microbiology study proposes a therapeutic strategy distinct from classic active-site inhibitors: blocking conserved protein–protein interactions required for assembly of glutamine amidotransferase complexes. The authors validate that disrupting conserved TrpG–TrpE interface hot spots is strongly growth-limiting in vivo on minimal medium due to tryptophan deficiency, and note high conservation of these residues across a non-redundant set of 695 AS/ADCS sequences (funke2024validationofaminodeoxychorismate pages 2-5). This work positions AS (TrpE/TrpG) not only as a metabolic enzyme but also as a potentially “high-robustness” antimicrobial target because interface disruption affects essential complex formation (funke2024validationofaminodeoxychorismate pages 2-5).

3.3 2024: functional genomics context in P. putida KT2440

A 2024 mSystems paper applies independent component analysis to a large RB‑TnSeq fitness compendium (179 conditions) to define functional gene modules (fModules) in KT2440. It reports a specific “tryptophan biosynthesis” fModule (9 genes) explaining 8.47% of dataset variance, indicating that tryptophan biosynthesis genes show a coherent, condition-dependent fitness signature in KT2440 (borchert2024machinelearninganalysis pages 4-6). (The excerpted text does not list whether PP_0417/trpE is among those 9 genes; gene-level membership is referenced via an external resource.)

4) Current applications and real-world implementations

4.1 Metabolic engineering in P. putida KT2440: anthranilate production from glucose

Anthranilate is both a pathway intermediate and an industrially relevant compound. In a KT2440 engineering study, a markerless deletion of trpDC (downstream of anthranilate) enabled anthranilate accumulation, facilitated by the fact that in KT2440 (unlike E. coli) trpEG and trpDC are encoded by separate ORFs (kuepper2015metabolicengineeringof pages 2-3). The authors further used expression constructs including trpES40FG (a feedback-insensitive anthranilate synthase variant) for improved production (kuepper2015metabolicengineeringof pages 2-3).

Quantitative performance: Under tryptophan-limited fed-batch conditions, the best engineered strain achieved 1.54 ± 0.3 g/L anthranilate (11.23 mM) from glucose (kuepper2015metabolicengineeringof pages 2-3). This demonstrates that TrpE-mediated flux control at the committed step is directly exploitable for biomanufacturing.

5) Statistics and data points (recent and/or authoritative)

  • Screen scale and genetics (KT2440): nearly 150,000 mini‑Tn5 mutants screened to isolate aromatic amino-acid auxotrophs; the trpE auxotroph had insertion in the 57th codon of PP_0417 (molinahenares2009functionalanalysisof pages 2-4).
  • Pathway genetics (KT2440): RT‑PCR evidence supports monocistronic trpE and trpGDC operon organization (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof media 744c52d9).
  • Bioprocess metric: 1.54 ± 0.3 g/L (11.23 mM) anthranilate achieved in engineered KT2440 (kuepper2015metabolicengineeringof pages 2-3).
  • Antibiotic-target conservation analysis: interface hot-spot residues conserved across 695 AS/ADCS sequences (non-redundant filtered set) in a 2024 validation study (funke2024validationofaminodeoxychorismate pages 2-5).
  • Systems-level functional genomics: tryptophan biosynthesis module explains 8.47% of variance in RB‑TnSeq dataset (KT2440) (borchert2024machinelearninganalysis pages 4-6).

6) Visual evidence (pathway and gene organization)

Molina‑Henares et al. provide (i) a pathway schematic detailing the enzymatic steps from chorismate to tryptophan and (ii) a KT2440 gene-organization diagram showing trpE as a separate unit from the trpGDC operon; these figures support both the biochemical pathway position and the transcriptional organization used in this annotation (molinahenares2009functionalanalysisof media 89069567, molinahenares2009functionalanalysisof media 744c52d9).

7) Conclusions for functional annotation of UniProt Q88QS1 (KT2440 TrpE)

Primary molecular function: chorismate→anthranilate synthase activity as the TrpE (anthranilate synthase component I) subunit of the TrpE/TrpG glutamine amidotransferase complex (EC 4.1.3.27), implementing the first committed step in L‑tryptophan biosynthesis (parthasarathy2018athreeringcircus pages 5-6, molinahenares2009functionalanalysisof pages 2-4).

Biological process: de novo L‑tryptophan biosynthesis from chorismate via anthranilate; genetic and feeding tests in KT2440 place trpE upstream of anthranilate/indole/tryptophan (molinahenares2009functionalanalysisof pages 4-6).

Complex/partners: requires interaction with TrpG to use glutamine as nitrogen source; the TrpG–TrpE interface is essential and evolutionarily conserved, now also investigated as a PPI-inhibitor target class (funke2024validationofaminodeoxychorismate pages 2-5).

Gene context: monocistronic trpE (PP_0417) separated from trpGDC operon; other trp genes are in additional loci (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof media 744c52d9).

Applied relevance: TrpE is a major control node for redirecting chorismate flux to anthranilate in P. putida KT2440, enabling gram‑per‑liter anthranilate production in engineered strains (kuepper2015metabolicengineeringof pages 2-3).


Summary table

Category Key points Best supporting citations (pqac IDs) URLs / publication dates
Identity UniProt Q88QS1 corresponds to trpE / PP_0417 in Pseudomonas putida KT2440 and encodes anthranilate synthase component I (large/synthase subunit) in tryptophan biosynthesis; KT2440 studies explicitly annotate PP0417 as TrpE and show loss of function causes tryptophan auxotrophy. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2) Molina-Henares et al., Microbial Biotechnology (Dec 2009): https://doi.org/10.1111/j.1751-7915.2008.00062.x
Reaction TrpE is the synthase/alpha component of anthranilate synthase (EC 4.1.3.27), catalyzing the first committed step of L-tryptophan biosynthesis by converting chorismate to anthranilate in concert with TrpG. Mechanistically, TrpE forms/acts on the aminated chorismate intermediate and supports pyruvate elimination. (parthasarathy2018athreeringcircus pages 5-6, naz2023insightintoderegulation pages 11-14) Parthasarathy et al., Frontiers in Molecular Biosciences (Apr 2018): https://doi.org/10.3389/fmolb.2018.00029; Naz et al., Microbial Cell Factories (Aug 2023): https://doi.org/10.1186/s12934-023-02178-z
Substrates / products Canonical glutamine-dependent reaction uses chorismate plus ammonia derived from glutamine; products are anthranilate, pyruvate, and glutamate. Under some conditions, the alpha subunit can use free ammonia instead of glutamine-derived ammonia. (parthasarathy2018athreeringcircus pages 5-6, naz2023insightintoderegulation pages 14-15) Parthasarathy et al. (Apr 2018): https://doi.org/10.3389/fmolb.2018.00029; Naz et al. (Aug 2023): https://doi.org/10.1186/s12934-023-02178-z
Pathway role In KT2440, trpE functions at or before anthranilate formation in the one-way pathway from chorismate to tryptophan; precursor-feeding experiments showed trpE mutants are rescued by anthranilate, indole, or tryptophan, placing TrpE upstream of these intermediates. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2, molinahenares2009functionalanalysisof media 89069567) Molina-Henares et al. (Dec 2009): https://doi.org/10.1111/j.1751-7915.2008.00062.x
Complex / partner TrpE acts with TrpG (anthranilate synthase component II / glutamine amidotransferase) as a two-subunit enzyme, commonly organized as αβ or α2β2 assemblies; TrpG supplies ammonia from glutamine to the TrpE active site. Conserved subunit interfaces are essential for activity. (parthasarathy2018athreeringcircus pages 5-6, naz2023insightintoderegulation pages 11-14, funke2024validationofaminodeoxychorismate pages 1-2, funke2024validationofaminodeoxychorismate pages 2-5) Parthasarathy et al. (Apr 2018): https://doi.org/10.3389/fmolb.2018.00029; Naz et al. (Aug 2023): https://doi.org/10.1186/s12934-023-02178-z; Funke et al., Applied and Environmental Microbiology (May 2024): https://doi.org/10.1128/aem.00572-24
Regulation Anthranilate synthase is feedback-inhibited by L-tryptophan. Structural analyses place inhibitory residues in/near the chorismate-binding domain and domain interfaces; elevated Trp prevents the conformational state needed for efficient catalysis/ammonia transfer. Feedback-insensitive trpE alleles are widely exploited in engineering. (niraula2025aromaticaminoacids pages 13-14, naz2023insightintoderegulation pages 11-14, naz2023insightintoderegulation pages 14-15, ramosvaldovinos2024optimizingfermentationstrategies pages 7-8) Niraula et al. (Jan 2025): https://doi.org/10.3390/biotech14010006; Naz et al. (Aug 2023): https://doi.org/10.1186/s12934-023-02178-z; Ramos-Valdovinos & Martínez-Antonio, Processes (Nov 2024): https://doi.org/10.3390/pr12112422
Gene organization In P. putida KT2440, trpE is a monocistronic transcription unit separate from the trpGDC operon; trp genes are split across multiple chromosomal loci, with trpBA and trpI elsewhere. Engineering literature also notes trpEG and trpDC are encoded by separate open reading frames in KT2440. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4, kuepper2015metabolicengineeringof pages 2-3, molinahenares2009functionalanalysisof media 89069567) Molina-Henares et al. (Dec 2009): https://doi.org/10.1111/j.1751-7915.2008.00062.x; Kuepper et al., Frontiers in Microbiology (Nov 2015): https://doi.org/10.3389/fmicb.2015.01310
Localization TrpE is a cytosolic enzyme inferred from its role in soluble primary metabolism and from its bacterial anthranilate synthase family organization; no evidence in the gathered KT2440 sources suggests membrane or extracellular localization. (parthasarathy2018athreeringcircus pages 5-6, naz2023insightintoderegulation pages 11-14) Parthasarathy et al. (Apr 2018): https://doi.org/10.3389/fmolb.2018.00029; Naz et al. (Aug 2023): https://doi.org/10.1186/s12934-023-02178-z
Phenotypes A mini-Tn5 insertion in trpE produced a tryptophan auxotroph in KT2440. Rescue by anthranilate/indole/tryptophan demonstrates specific impairment of the first committed step of Trp biosynthesis rather than a broad growth defect. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2) Molina-Henares et al. (Dec 2009): https://doi.org/10.1111/j.1751-7915.2008.00062.x
Applications / engineering KT2440 has been engineered to accumulate anthranilate by deleting trpDC and overexpressing feedback-insensitive trpE/trpG variants. In fed-batch, the best reported strain reached 1.54 ± 0.3 g/L anthranilate (11.23 mM) from glucose. TrpE is thus a practical flux-control node for aromatic biomanufacturing. (kuepper2015metabolicengineeringof pages 2-3) Kuepper et al., Frontiers in Microbiology (Nov 2015): https://doi.org/10.3389/fmicb.2015.01310
Recent developments (2023–2024) Recent work emphasizes (i) structure-guided deregulation of anthranilate synthase feedback inhibition, including residue-level analysis of Trp and chorismate binding; (ii) PPI-targeted antibiotic strategies aimed at the conserved TrpG–TrpE interface; and (iii) renewed interest in aromatic-pathway rewiring for high-yield production platforms. (naz2023insightintoderegulation pages 11-14, naz2023insightintoderegulation pages 14-15, funke2024validationofaminodeoxychorismate pages 1-2, funke2024validationofaminodeoxychorismate pages 2-5) Naz et al. (Aug 2023): https://doi.org/10.1186/s12934-023-02178-z; Funke et al. (May 2024): https://doi.org/10.1128/aem.00572-24
Statistics / data Quantitative findings include: 1.54 ± 0.3 g/L (11.23 mM) anthranilate in engineered KT2440; Funke et al. analyzed conservation across ~695 bacterial AS/ADCS sequences to identify interface hotspots; Molina-Henares et al. isolated four tryptophan auxotrophs in KT2440 mutagenesis, including a trpE mutant. (kuepper2015metabolicengineeringof pages 2-3, funke2024validationofaminodeoxychorismate pages 2-5, molinahenares2009functionalanalysisof pages 1-2) Kuepper et al. (Nov 2015): https://doi.org/10.3389/fmicb.2015.01310; Funke et al. (May 2024): https://doi.org/10.1128/aem.00572-24; Molina-Henares et al. (Dec 2009): https://doi.org/10.1111/j.1751-7915.2008.00062.x

Table: This table summarizes the validated functional annotation of Pseudomonas putida KT2440 trpE (UniProt Q88QS1 / PP_0417), covering biochemical function, gene organization, phenotypes, engineering uses, and recent 2023–2024 developments with direct context-ID citations.

References

  1. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  2. (parthasarathy2018athreeringcircus pages 5-6): Anutthaman Parthasarathy, Penelope J. Cross, Renwick C. J. Dobson, Lily E. Adams, Michael A. Savka, and André O. Hudson. A three-ring circus: metabolism of the three proteogenic aromatic amino acids and their role in the health of plants and animals. Frontiers in Molecular Biosciences, Apr 2018. URL: https://doi.org/10.3389/fmolb.2018.00029, doi:10.3389/fmolb.2018.00029. This article has 423 citations.

  3. (naz2023insightintoderegulation pages 11-14): Sadia Naz, Pi Liu, Umar Farooq, and Hongwu Ma. Insight into de-regulation of amino acid feedback inhibition: a focus on structure analysis method. Microbial Cell Factories, Aug 2023. URL: https://doi.org/10.1186/s12934-023-02178-z, doi:10.1186/s12934-023-02178-z. This article has 23 citations and is from a peer-reviewed journal.

  4. (funke2024validationofaminodeoxychorismate pages 2-5): Franziska Jasmin Funke, Sandra Schlee, and Reinhard Sterner. Validation of aminodeoxychorismate synthase and anthranilate synthase as novel targets for bispecific antibiotics inhibiting conserved protein-protein interactions. Applied and Environmental Microbiology, May 2024. URL: https://doi.org/10.1128/aem.00572-24, doi:10.1128/aem.00572-24. This article has 5 citations and is from a peer-reviewed journal.

  5. (naz2023insightintoderegulation pages 14-15): Sadia Naz, Pi Liu, Umar Farooq, and Hongwu Ma. Insight into de-regulation of amino acid feedback inhibition: a focus on structure analysis method. Microbial Cell Factories, Aug 2023. URL: https://doi.org/10.1186/s12934-023-02178-z, doi:10.1186/s12934-023-02178-z. This article has 23 citations and is from a peer-reviewed journal.

  6. (niraula2025aromaticaminoacids pages 13-14): Archana Niraula, Amir Danesh, Natacha Merindol, Fatma Meddeb-Mouelhi, and Isabel Desgagné-Penix. Aromatic amino acids: exploring microalgae as a potential biofactory. BioTech, 14:6, Jan 2025. URL: https://doi.org/10.3390/biotech14010006, doi:10.3390/biotech14010006. This article has 9 citations.

  7. (molinahenares2009functionalanalysisof pages 1-2): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  8. (molinahenares2009functionalanalysisof pages 4-6): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  9. (molinahenares2009functionalanalysisof media 744c52d9): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  10. (borchert2024machinelearninganalysis pages 4-6): Andrew J. Borchert, Alissa C. Bleem, Hyun Gyu Lim, Kevin Rychel, Keven D. Dooley, Zoe A. Kellermyer, Tracy L. Hodges, Bernhard O. Palsson, and Gregg T. Beckham. Machine learning analysis of rb-tnseq fitness data predicts functional gene modules in pseudomonas putida kt2440. Mar 2024. URL: https://doi.org/10.1128/msystems.00942-23, doi:10.1128/msystems.00942-23. This article has 13 citations and is from a peer-reviewed journal.

  11. (kuepper2015metabolicengineeringof pages 2-3): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

  12. (molinahenares2009functionalanalysisof media 89069567): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  13. (funke2024validationofaminodeoxychorismate pages 1-2): Franziska Jasmin Funke, Sandra Schlee, and Reinhard Sterner. Validation of aminodeoxychorismate synthase and anthranilate synthase as novel targets for bispecific antibiotics inhibiting conserved protein-protein interactions. Applied and Environmental Microbiology, May 2024. URL: https://doi.org/10.1128/aem.00572-24, doi:10.1128/aem.00572-24. This article has 5 citations and is from a peer-reviewed journal.

  14. (ramosvaldovinos2024optimizingfermentationstrategies pages 7-8): Miguel Angel Ramos-Valdovinos and Agustino Martínez-Antonio. Optimizing fermentation strategies for enhanced tryptophan production in escherichia coli: integrating genetic and environmental controls for industrial applications. Processes, Nov 2024. URL: https://doi.org/10.3390/pr12112422, doi:10.3390/pr12112422. This article has 10 citations.

Artifacts

Citations

  1. molinahenares2009functionalanalysisof pages 2-4
  2. naz2023insightintoderegulation pages 14-15
  3. molinahenares2009functionalanalysisof pages 1-2
  4. molinahenares2009functionalanalysisof pages 4-6
  5. funke2024validationofaminodeoxychorismate pages 2-5
  6. borchert2024machinelearninganalysis pages 4-6
  7. kuepper2015metabolicengineeringof pages 2-3
  8. parthasarathy2018athreeringcircus pages 5-6
  9. naz2023insightintoderegulation pages 11-14
  10. niraula2025aromaticaminoacids pages 13-14
  11. funke2024validationofaminodeoxychorismate pages 1-2
  12. ramosvaldovinos2024optimizingfermentationstrategies pages 7-8
  13. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  14. https://doi.org/10.3389/fmolb.2018.00029;
  15. https://doi.org/10.1186/s12934-023-02178-z
  16. https://doi.org/10.1186/s12934-023-02178-z;
  17. https://doi.org/10.1128/aem.00572-24
  18. https://doi.org/10.3390/biotech14010006;
  19. https://doi.org/10.3390/pr12112422
  20. https://doi.org/10.1111/j.1751-7915.2008.00062.x;
  21. https://doi.org/10.3389/fmicb.2015.01310
  22. https://doi.org/10.3389/fmicb.2015.01310;
  23. https://doi.org/10.1128/aem.00572-24;
  24. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  25. https://doi.org/10.3389/fmolb.2018.00029,
  26. https://doi.org/10.1186/s12934-023-02178-z,
  27. https://doi.org/10.1128/aem.00572-24,
  28. https://doi.org/10.3390/biotech14010006,
  29. https://doi.org/10.1128/msystems.00942-23,
  30. https://doi.org/10.3389/fmicb.2015.01310,
  31. https://doi.org/10.3390/pr12112422,

📄 View Raw YAML

id: Q88QS1
gene_symbol: trpE
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: >-
  Anthranilate synthase component I (TrpE, locus PP_0417), the large alpha
  subunit of anthranilate synthase (EC 4.1.3.27). Together with the glutamine
  amidotransferase beta subunit TrpG, it forms a heterotetrameric complex that
  catalyzes the first committed step of L-tryptophan biosynthesis, the conversion
  of chorismate to anthranilate. TrpE binds chorismate and performs the
  amination/lyase chemistry, using ammonia supplied by TrpG from hydrolysis of
  L-glutamine; the products are anthranilate, pyruvate and L-glutamate. In the
  absence of TrpG, TrpE alone can produce anthranilate directly from chorismate
  when free ammonia is abundant. The enzyme requires Mg2+ and is a soluble,
  cytoplasmic enzyme of aromatic amino acid primary metabolism. Anthranilate
  synthase is typically feedback-inhibited by L-tryptophan, the pathway end
  product. In P. putida KT2440, loss of trpE causes tryptophan auxotrophy that is
  rescued by anthranilate, indole or tryptophan, confirming its placement at the
  anthranilate-forming step.
existing_annotations:
- term:
    id: GO:0000162
    label: L-tryptophan biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: >-
      TrpE catalyzes the first committed step (chorismate to anthranilate) of de
      novo L-tryptophan biosynthesis. This is the correct, well-supported
      biological process for this enzyme.
    action: ACCEPT
    reason: >-
      Anthranilate synthase component I is the entry enzyme of the tryptophan
      branch of aromatic amino acid biosynthesis. The IEA assignment (InterPro
      IPR005256, UniPathway UPA00035) is corroborated by experimental genetics in
      KT2440, where a trpE (PP_0417) insertion mutant is a tryptophan auxotroph
      rescued by anthranilate, indole or tryptophan (PMID:21261884; see
      file:PSEPK/trpE/trpE-deep-research-falcon.md).
- term:
    id: GO:0004049
    label: anthranilate synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: >-
      TrpE is the synthase (alpha) component of anthranilate synthase
      (EC 4.1.3.27), catalyzing chorismate + L-glutamine to anthranilate +
      pyruvate + L-glutamate (RHEA:21732). This is the core molecular function
      and is correct.
    action: ACCEPT
    reason: >-
      Assigned from InterPro IPR005256, RHEA:21732 and EC 4.1.3.27, consistent
      with the protein family (anthranilate synthase component I), the Pfam
      chorismate-binding domain, and the UniProt catalytic activity annotation.
      GO:0004049 is the precise molecular function term.
core_functions:
- description: >-
    Anthranilate synthase component I activity - binds chorismate and, using
    ammonia supplied by the TrpG glutaminase subunit (or free ammonia at high
    concentration), converts chorismate to anthranilate with release of pyruvate,
    the first committed step of L-tryptophan biosynthesis.
  molecular_function:
    id: GO:0004049
    label: anthranilate synthase activity
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: >-
      Anthranilate synthase component 1; EC 4.1.3.27; Reaction=chorismate +
      L-glutamine = anthranilate + pyruvate + L-glutamate + H(+); Rhea:RHEA:21732.
    full_text_unavailable: true
  - reference_id: PMID:21261884
    supporting_text: >-
      A trpE mutant did not grow on minimal medium but growth was restored by
      anthranilate, indole, or tryptophan supplementation, placing TrpE
      (PP_0417) at the anthranilate-forming step of tryptophan biosynthesis.
    full_text_unavailable: true
  directly_involved_in:
  - id: GO:0000162
    label: L-tryptophan biosynthetic process
  substrates:
  - id: CHEBI:29748
    label: chorismate
  - id: CHEBI:58359
    label: L-glutamine
references:
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:21261884
  title: Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440
  findings:
  - statement: >-
      A mini-Tn5 insertion in trpE (PP_0417) produces a tryptophan auxotroph in
      KT2440 rescued by anthranilate, indole or tryptophan, placing TrpE at the
      anthranilate-forming step of tryptophan biosynthesis; trpE is a
      monocistronic unit separate from the trpGDC operon.
    supporting_text: >-
      A trpE mutant did not grow on minimal medium but growth was restored by
      anthranilate, indole, or tryptophan supplementation.
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: >-
      PubMed-verified: PMID:21261884 = Molina-Henares et al., Microbial
      Biotechnology 2009, 2(1):91-100, doi:10.1111/j.1751-7915.2008.00062.x. The
      abstract confirms KT2440 has a single chorismate-to-tryptophan pathway with
      trpE as a separate transcriptional unit, isolated via mini-Tn5 auxotroph
      screening; supports the tryptophan-auxotrophy phenotype for PP_0417/trpE.
proposed_new_terms: []
suggested_questions: []
suggested_experiments: []