trpC

UniProt ID: Q88QR6
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

Indole-3-glycerol phosphate synthase (IGPS; TrpC; EC 4.1.1.48), the enzyme catalyzing the fourth step of L-tryptophan biosynthesis. It converts 1-(2-carboxyphenylamino)-1-deoxy-D-ribulose 5-phosphate (CdRP) into (1S,2R)-1-C-(indol-3-yl)glycerol 3-phosphate (indole-3-glycerol phosphate, IGP), releasing CO2 and water in an irreversible decarboxylative ring-closure reaction. This indole-ring-forming step lies downstream of anthranilate phosphoribosyltransferase (TrpD) and upstream of the tryptophan synthase subunits (TrpA/TrpB). The protein adopts a classic (beta/alpha)8 TIM-barrel (ribulose-phosphate-binding barrel) fold and acts as a soluble cytoplasmic metabolic enzyme. In Pseudomonas putida KT2440 the gene (PP_0422) lies in the trpGDC operon, and loss-of-function insertions cause tryptophan auxotrophy. The enzyme family is highly conserved across bacteria and has no human homolog.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000162 L-tryptophan biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: TrpC/IGPS catalyzes the fourth step of L-tryptophan biosynthesis; this process annotation is correct and represents a core function of the gene.
Reason: Matches the UniProt-curated pathway annotation (L-tryptophan from chorismate, step 4/5) and is supported by tryptophan-auxotrophy phenotypes of PP_0422 insertion mutants in P. putida KT2440. This is a core function.
GO:0004425 indole-3-glycerol-phosphate synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: This is the canonical molecular function of TrpC, matching the UniProt RecName (Indole-3-glycerol phosphate synthase, EC 4.1.1.48) and the curated catalytic activity (CdRP to indole-3-glycerol phosphate + CO2 + H2O).
Reason: Directly supported by UniProt HAMAP-Rule MF_00134, InterPro IGPS domain signatures (IPR001468/IPR013798/IPR045186), EC 4.1.1.48 and RHEA:23476. Core molecular function.
GO:0004640 phosphoribosylanthranilate isomerase activity
IEA
GO_REF:0000118
REMOVE
Summary: Phosphoribosylanthranilate isomerase (PRAI, TrpF, EC 5.3.1.24) catalyzes the third step of tryptophan biosynthesis and is a distinct activity from IGPS. This activity is NOT supported for P. putida KT2440 TrpC by UniProt, which assigns only IGPS (EC 4.1.1.48). The annotation is a TreeGrafter over-propagation arising because some bacteria (e.g. E. coli) have a bifunctional TrpC(F) protein, whereas in Pseudomonas TrpF is a separate gene.
Reason: UniProt Q88QR6 (HAMAP MF_00134) annotates only the monofunctional IGPS activity (277 aa, single IGPS domain), with no PRAI/TrpF domain. The IEA TreeGrafter inference reflects the bifunctional TrpCF architecture of some lineages and is not applicable to this monofunctional Pseudomonas enzyme. This is an electronic (IEA) prediction argued against on biological/domain grounds, not the second-guessing of an experimental annotation.

Core Functions

Catalyzes the fourth step of L-tryptophan biosynthesis, the decarboxylative ring closure of CdRP to indole-3-glycerol phosphate.

Supporting Evidence:
  • GO_REF:0000120
    UniProt RecName Indole-3-glycerol phosphate synthase, EC 4.1.1.48; catalytic activity CdRP to (indol-3-yl)glycerol 3-phosphate + CO2 + H2O (RHEA:23476).

References

TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
file:PSEPK/trpC/trpC-deep-research-falcon.md
Deep research report on P. putida KT2440 trpC (PP_0422)
  • PP_0422 (trpC) lies in the trpGDC operon; a mini-Tn5 insertion in PP_0422 yields tryptophan auxotrophy in KT2440, confirming its essential role in endogenous tryptophan biosynthesis.
    "Molina-Henares et al. 2009 (Microbial Biotechnology 2:91-100) operon mapping, RT-PCR cotranscription, and auxotrophy screen."
Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440

Suggested Questions for Experts

Q: Has the IGPS activity of P. putida KT2440 TrpC (PP_0422) been biochemically characterized (kcat/KM), or is the assignment based solely on homology and the auxotrophy phenotype?

Suggested Experiments

Experiment: Purify recombinant PP_0422 and measure IGPS activity (CdRP to IGP) to confirm the monofunctional assignment and exclude any cryptic PRAI activity.

Deep Research

Asta

(trpC-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-05T20:18:20.319503

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.740)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[2] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.735)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[3] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.731)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.669)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[4] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.723)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[5] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.720)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[6] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.706)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[7] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.690)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[8] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.683)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[9] Undergraduate Bioinformatics Conceptualizing Form and Function on a Molecular Scale

  • Authors: P. Kramer, Jack Treml
  • Year: 2022
  • Venue: Midwestern Journal of Undergraduate Sciences
  • URL: https://www.semanticscholar.org/paper/58a6af68fc2fd768735d429724ffdd52e16dfb27
  • DOI: 10.17161/mjusc.v1i1.18565
  • Summary: The following is a walkthrough of a project designed to overcome the lack of sense for proteins as real objects.
  • Evidence snippets:
  • Snippet 1 (score: 0.679)
    > i. Click "See more" to view a bar chart containing data on where in the body's tissues the gene is expressed (as determined by RNA sequencing). Save and include this bar chart as the deliverable for this step.
    > II. Universal Protein Research Knowledgebase (UniProtKB) 8 6. UniProt Entry Number
    > i. Follow the UniProt link in the Resources then search for the protein using the NCBI Gene ID ii. Carefully select the result that best matches the gene and organism of interest by clicking on the blue entry number. iii. This page will be used later to gather further details about the protein.
    > III. RCSB Protein Data Bank (PDB) 9 7. RCSB PDB Solved Structure Identifi er i. Follow the RCSB PDB link in the Resources and search for the protein by either the common name or the NCBI Gene ID, making sure to select the organism of interest on the left. ii. You must ensure that your chosen protein has an existing solved structure in this data bank in order to do a mutational analysis in later parts of this exercise.
    > IV. NCBI GenBank 10 8. AA Protein Sequence i. From the NCBI Gene page, go to the "Genomic regions, transcripts, and products" section and then click "GenBank" on the right. Scroll down to the fi rst Coding Sequence "CDS" section and look directly after "/translation=" for the full protein sequence. ii. Sequence needs to be in FASTA Format consisting of '>' followed by a simple name, a return, and then the sequence in one continuous line of text. See "FASTA Formatting" link in Resources.

[10] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.676)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[11] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[12] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.669)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[13] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[14] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[15] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.658)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[16] FlyBase: establishing a Gene Group resource for Drosophila melanogaster

  • Authors: H. Attrill, K. Falls, J. Goodman, G. Millburn, Giulia Antonazzo et al.
  • Year: 2015
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/0d58fdc5f1dfcb2910f7cc003338a82ef827f85c
  • DOI: 10.1093/nar/gkv1046
  • PMID: 26467478
  • PMCID: 4702782
  • Citations: 326
  • Influential citations: 28
  • Summary: FlyBase (flybase.org), the MOD for Drosophila melanogaster, has established a ‘Gene Group’ resource: high-quality sets of genes derived from the published literature and organized into individual report pages, to enable researchers with diverse backgrounds and interests to easily view and analyse acknowledged D. melanogsaster gene sets and compare them with those of other species.
  • Evidence snippets:
  • Snippet 1 (score: 0.657)
    > Given the wealth of post-genomic data, it should be straightforward to query a biological database and obtain a robust list of genes related by the shared attributes of their products, such as actins, protein kinases or subunits of the proteasome. However, this is often not the case. For evolutionary-related genes, BLAST (1) or domain-based searches may yield a good preliminary list, but without further analysis, particularly when inferring function from se-quence, it is hard to distinguish false positives. Additionally, there may be false negatives because a gene may be somewhat atypical or fails to score above a given threshold. Gene Ontology (GO) annotations can be used to search for gene products that are related by common biological attributes, but annotation is not exhaustive and expressing features that pertain to sequence does not fall within its scope (2,3).
    > An alternative approach to finding related genes (at least in species with well-characterized genomes such as humans and model organisms) can be to search for gene symbols that share a common prefix. This can be effective for databases such as WormBase (4) or the HUGO Gene Nomenclature Committee (HGNC) (5) that assign unifying and systematic gene symbols to nematode and human genes, respectively, based on shared structures, functions or phenotypes. However, this strategy is not generally applicable to Drosophila melanogaster genes in FlyBase (6) as these are traditionally named on a gene-by-gene basis by the authors who first publish on the gene, often reflecting a specific mutant phenotype.
    > A third method for acquiring a set of related genes is to directly consult relevant research or review articles. The advantage of this approach is that the list is compiled directly from an expert and peer-reviewed source, and as such will be robust and clearly attributable. However, it can be time-consuming to seek out and then extract a set of genes from individual publications (or, often, their supplementary data) and lists obtained in this manner are inherently uncoupled from the relevant species database, meaning that the listed gene symbols/IDs may become stale over time.
    > Several databases have addressed these issues by providing explicit sets of related genes.

[17] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[18] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.654)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[19] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.653)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  2. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  3. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  4. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  5. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  6. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  7. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  8. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  9. P. Kramer, Jack Treml (2022). Undergraduate Bioinformatics Conceptualizing Form and Function on a Molecular Scale. Midwestern Journal of Undergraduate Sciences. https://www.semanticscholar.org/paper/58a6af68fc2fd768735d429724ffdd52e16dfb27
  10. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  11. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  12. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  13. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  14. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  15. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  16. H. Attrill, K. Falls, J. Goodman, G. Millburn, Giulia Antonazzo et al. (2015). FlyBase: establishing a Gene Group resource for Drosophila melanogaster. Nucleic Acids Research. https://www.semanticscholar.org/paper/0d58fdc5f1dfcb2910f7cc003338a82ef827f85c
  17. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  18. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  19. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76

Falcon

(trpC-deep-research-falcon.md)
Functional annotation research report: **trpC (PP_0422; UniProt Q88QR6)** in *Pseudomonas putida* KT2440 Falcon Edison Scientific Literature 27 citations 2 artifacts 2026-06-11T21:27:32.022254

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Functional annotation research report: trpC (PP_0422; UniProt Q88QR6) in Pseudomonas putida KT2440

0) Scope and identity verification (critical disambiguation)

Target identity (confirmed): The gene symbol trpC in this report refers specifically to PP_0422 from Pseudomonas putida strain KT2440 (ATCC 47054 / DSM 6125 / etc.), annotated in UniProt as indole-3-glycerol phosphate synthase (IGPS; EC 4.1.1.48). Multiple independent lines of KT2440-specific evidence support this assignment: PP_0422 lies in the trpGDC tryptophan-biosynthesis operon, shows high similarity to characterized TrpC from E. coli, and loss-of-function insertions in PP_0422 yield tryptophan auxotrophy. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2)

Why ambiguity matters: “trpC” is a widely used bacterial symbol; some taxa have fused or multifunctional proteins combining TrpC with other activities, or gene fusions in related pathways. Therefore, all organism-specific conclusions below are restricted to KT2440 PP_0422/Q88QR6. (barona‐gomez2003occurrenceofa pages 1-2, molinahenares2009functionalanalysisof pages 2-4)

1) Key concepts and definitions (current understanding)

1.1 Canonical function of TrpC/IGPS

Indole-3-glycerol phosphate synthase (IGPS; TrpC; EC 4.1.1.48) catalyzes the indole-ring-forming step of microbial tryptophan biosynthesis by converting 1-(o-carboxyphenylamino)-1-deoxyribulose 5′-phosphate (CdRP) into indole-3-glycerol phosphate (IGP; also abbreviated InGP). (esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3, molinahenares2009functionalanalysisof pages 4-6)

In P. putida KT2440, pathway mapping places TrpC downstream of TrpD (anthranilate phosphoribosyltransferase) and upstream of TrpA/TrpB (tryptophan synthase subunits), consistent with the standard chorismate→anthranilate→tryptophan route. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media a5651e48)

1.2 Pathway context in KT2440

A KT2440 genetic and biochemical reconstruction supports a single route from chorismate to tryptophan, with the trp genes arranged in operons and single-gene transcription units characteristic of many bacteria. (molinahenares2009functionalanalysisof pages 1-2, molinahenares2009functionalanalysisof pages 4-6)

Figure evidence: the tryptophan pathway and trp gene organization in KT2440 (including trpGDC with trpC/PP_0422) are shown in extracted figures from Molina-Henares et al. (2009). (molinahenares2009functionalanalysisof media a5651e48, molinahenares2009functionalanalysisof media 77d9c78a)

2) Molecular function: reaction chemistry, substrate specificity, and mechanism

2.1 Reaction and substrate specificity

IGPS/TrpC is functionally defined by its specificity for the pathway intermediate CdRP, producing IGP. This is the enzyme’s core substrate/product relationship in bacteria. (esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3)

In KT2440, TrpC is explicitly identified as indoleglycerol phosphate synthase within the reconstructed pathway and gene cluster, supporting that PP_0422 participates in this exact chemistry rather than a divergent TrpC-like function. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4)

2.2 Enzyme fold and catalytic mechanism (authoritative mechanistic synthesis)

A detailed minireview of bacterial IGPS (focused on Mycobacterium tuberculosis IGPS but comparing across homologs) describes IGPS as a (β/α)8 TIM-barrel enzyme (a classic aldolase/TIM-barrel scaffold). (esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3, esposito2022indole‐3‐glycerolphosphatesynthase pages 4-6)

Mechanistic consensus (3-step sequence): cyclization to an intermediate, irreversible decarboxylation, and dehydration to yield IGP. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4)

Catalytic residue-level insights (from structural/kinetic studies of bacterial IGPS): conserved Lys/Glu residues participate in proton transfers and dehydration; structural data identify phosphate-binding and indole-stabilizing interactions in the active site (e.g., H-bonding and π-cation interactions) as part of the conserved catalytic architecture. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4, esposito2022indole‐3‐glycerolphosphatesynthase pages 4-6)

Quantitative kinetic data (example bacterial IGPS, not KT2440-specific): for MtIGPS, reported KM values vary substantially across studies (e.g., 55 μM vs. 500 μM vs. 1.13 mM) and a reported kcat ≈ 0.16 s−1 (25 °C); pH dependence indicates apparent pKa values around ~6–6.8 for catalytic parameters. These figures illustrate that IGPS kinetics are sensitive to assay and substrate preparation, and should not be transferred quantitatively to KT2440 without direct measurement. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4, esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3)

3) KT2440-specific genomic context, regulation, and phenotype evidence

3.1 Operon context and transcriptional organization

In P. putida KT2440, trpC = PP_0422 is within a compact operon trpG–trpD–trpC (trpGDC), with very short intergenic distances (including a 6-nt overlap between trpD and trpC) and RT-PCR evidence for cotranscription of trpGDC. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6)

This organization supports functional coupling with upstream tryptophan-branch reactions and is consistent with tight pathway control at the transcriptional/operon level. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media 77d9c78a)

3.2 Loss-of-function phenotypes (auxotrophy)

A genome-wide mini-Tn5 screen in KT2440 identified tryptophan auxotrophs with insertions in multiple trp genes, including a PP_0422 (trpC) insertion at the 32nd codon leading to tryptophan requirement. This is direct genetic evidence that PP_0422 is essential for endogenous tryptophan biosynthesis under the tested minimal conditions. (molinahenares2009functionalanalysisof pages 1-2, molinahenares2009functionalanalysisof pages 2-4)

3.3 Conservation across Pseudomonas

Tryptophan-pathway enzymes in Pseudomonas were reported as highly conserved (71–97% identity range across species for pathway enzymes), supporting functional conservation of TrpC/IGPS across the genus and lending additional confidence to annotation transfer. (molinahenares2009functionalanalysisof pages 4-6)

4) Cellular localization

No direct experimental subcellular localization (e.g., fractionation, microscopy, localization tags) for KT2440 TrpC (PP_0422/Q88QR6) was found in the retrieved full-text evidence. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4)

Given its role in central amino-acid biosynthesis and the fact that bacterial tryptophan-biosynthesis enzymes are generally soluble metabolic enzymes, the most parsimonious functional inference is that TrpC acts in the cytosol; however, this remains an inference rather than a demonstrated localization for KT2440 in the currently retrieved literature. (molinahenares2009functionalanalysisof pages 4-6)

5) Current applications and real-world implementations

5.1 Metabolic engineering in P. putida KT2440: blocking TrpC to accumulate anthranilate

A practical, KT2440-specific application of trpC knowledge is pathway rerouting at the chorismate→anthranilate→tryptophan node.

Kuepper et al. (2015) engineered P. putida KT2440 for anthranilate (o-aminobenzoate; oAB) production from glucose by deleting the trpDC operon, explicitly described as encoding TrpD (anthranilate phosphoribosyltransferase) and TrpC (IGPS). This deletion blocks anthranilate consumption toward tryptophan. (kuepper2015metabolicengineeringof pages 1-2)

Quantitative outcome: with additional pathway deregulation (feedback-insensitive aroG and modified trpE/trpS context as described), the best engineered strain achieved 1.54 ± 0.3 g/L (11.23 mM) anthranilate in tryptophan-limited fed-batch culture. This is a direct, real-world strain-engineering use-case for trpC functional annotation in KT2440. (kuepper2015metabolicengineeringof pages 1-2)

6) Recent developments and latest research (prioritizing 2023–2024)

6.1 2024: Shikimate pathway rewiring in P. putida enables near-theoretical aromatic yields (context for tryptophan-branch control)

A 2024 preprint/report describes reprogramming P. putida catabolism to rely on the shikimate pathway (normally anabolic) as a primary route for growth-associated pyruvate supply (“shikimate pathway-dependent catabolism”, SDC). The study combined metabolic modeling, rational engineering, and adaptive laboratory evolution (ALE) with biosensor-based selection, achieving ~89% of the pathway’s maximum theoretical yield for an aromatic product (4-hydroxybenzoate). (bruinsma2024shikimatepathwaydependentcatabolism pages 1-4, santos2024shikimatepathwaydependentcatabolism pages 1-5)

This work is not a direct study of trpC, but it is highly relevant to the tryptophan branch because it demonstrates that modern strategies can drive very high flux through chorismate/shikimate-derived nodes in P. putida. It also highlights design constraints: deleting certain chorismate-derived reactions (e.g., anthranilate synthase routes) is infeasible because it would create auxotrophies, underscoring the physiological coupling of aromatic product formation to essential metabolites like tryptophan. (bruinsma2024shikimatepathwaydependentcatabolism pages 4-6)

6.2 2023: Engineering strategies to manage anthranilate/tryptophan-branch byproducts (transferable principles)

A 2023 metabolic engineering study in Corynebacterium glutamicum optimized shikimate-derived production and explicitly addressed anthranilate accumulation (a tryptophan-branch metabolite) by reducing anthranilate synthase activity via targeted mutagenesis while avoiding tryptophan auxotrophy; the study achieved 661 mg/L (4 mM) p-coumaric acid and demonstrated co-culture conversion to 31.2 mg/L resveratrol. While not in P. putida, these results reflect current engineering practice at the aromatic branchpoint where trp genes (including trpC) are often manipulated to tune flux. (mutz2023microbialsynthesisof pages 1-2)

7) Expert opinions and analysis (authoritative sources)

7.1 IGPS as a structurally conserved TIM-barrel enzyme and potential target

A focused minireview (ChemBioChem, 2022) frames IGPS as a highly conserved TIM-barrel enzyme, summarizes evidence for its mechanism and structural determinants, and argues that bacterial IGPS can be a useful antibacterial target because some pathogens rely on tryptophan biosynthesis during host immune pressure and because the enzyme lacks a human homolog. Although discussed in a tuberculosis context, these points represent authoritative expert synthesis of IGPS enzymology and structure–function relationships. (esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3)

7.2 Genome-context + mutant phenotypes as strong functional evidence in KT2440

Molina-Henares et al. (2009) provide a KT2440-specific, experimentally grounded approach to functional annotation that combines operon mapping (RT-PCR), mutant phenotypes, and pathway logic. This type of evidence is particularly strong for essential metabolic enzymes like TrpC even when purified-enzyme kinetics for the exact strain are not available. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2)

8) Evidence map (compact summary)

Feature Key finding Best citation IDs to support it Source (author/year/doi URL)
Enzyme name trpC / PP_0422 / UniProt Q88QR6 in Pseudomonas putida KT2440 is assigned as indole-3-glycerol phosphate synthase (IGPS) based on sequence similarity to E. coli TrpC and placement in the tryptophan biosynthesis locus. (molinahenares2009functionalanalysisof pages 2-4) Molina-Henares et al., 2009, Microbial Biotechnology, https://doi.org/10.1111/j.1751-7915.2008.00062.x
EC number The gene product corresponds to EC 4.1.1.48, the canonical IGPS carboxy-lyase in tryptophan biosynthesis. This matches the functional assignment used in comparative and pathway literature. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4, barona‐gomez2003occurrenceofa pages 1-2) Esposito et al., 2022, ChemBioChem, https://doi.org/10.1002/cbic.202100314; Barona-Gómez & Hodgson, 2003, EMBO Reports, https://doi.org/10.1038/sj.embor.embor771
Reaction IGPS catalyzes conversion of 1-(o-carboxyphenylamino)-1-deoxyribulose-5-phosphate (CdRP) to indole-3-glycerol phosphate (IGP/InGP), the indole-ring-forming step of the pathway. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4, esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3) Esposito et al., 2022, ChemBioChem, https://doi.org/10.1002/cbic.202100314
Substrate/product The substrate is CdRP and the product is IGP/InGP; in the KT2440 pathway map, TrpC is explicitly positioned at the step producing indoleglycerol phosphate from the upstream anthranilate-derived intermediate. (molinahenares2009functionalanalysisof pages 4-6, esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3) Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x; Esposito et al., 2022, https://doi.org/10.1002/cbic.202100314
Pathway step TrpC functions in the tryptophan biosynthetic pathway downstream of TrpD and upstream of TrpA/TrpB, converting the anthranilate branch intermediate into IGP before terminal tryptophan synthase steps. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media a5651e48) Molina-Henares et al., 2009, Microbial Biotechnology, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Operon context In KT2440, trpC is part of the trpGDC operon; RT-PCR showed cotranscription of trpG-trpD-trpC, and trpD overlaps trpC by 6 nt, indicating tight transcriptional/functional coupling. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6) Molina-Henares et al., 2009, Microbial Biotechnology, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Evidence type Functional annotation is supported by genome context, RT-PCR operon mapping, sequence similarity, pathway reconstruction, and mutant phenotype analysis rather than direct biochemical characterization of the purified KT2440 enzyme. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 7-8) Molina-Henares et al., 2009, Microbial Biotechnology, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Phenotype Loss of trpC causes tryptophan auxotrophy in KT2440: a mini-Tn5 insertion in PP_0422 yielded an auxotrophic mutant in the aromatic-pathway screen. (molinahenares2009functionalanalysisof pages 1-2) Molina-Henares et al., 2009, Microbial Biotechnology, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Applications / metabolic engineering Deleting trpDC (including trpC) in KT2440 was used to block anthranilate consumption and redirect flux for product formation; the best engineered strain produced 1.54 ± 0.3 g/L (11.23 mM) anthranilate in tryptophan-limited fed-batch culture. (kuepper2015metabolicengineeringof pages 1-2) Kuepper et al., 2015, Frontiers in Microbiology, https://doi.org/10.3389/fmicb.2015.01310
Structure / domains TrpC/IGPS is a (β/α)8 TIM-barrel enzyme with a conserved catalytic architecture. The UniProt/domain assignment for Q88QR6 is consistent with this family-level structural model, although no KT2440-specific structure was identified here. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4, esposito2022indole‐3‐glycerolphosphatesynthase pages 4-6, kursula2003crystallographicstudieson pages 28-30) Esposito et al., 2022, ChemBioChem, https://doi.org/10.1002/cbic.202100314; Kursula, 2003
Localization No direct KT2440 localization experiment was found in the retrieved evidence. Given its role in central amino-acid biosynthesis and lack of membrane/export features in the discussed sources, TrpC is best inferred to function in the cytosol. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4) Molina-Henares et al., 2009, Microbial Biotechnology, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Recent 2023–2024 relevance Recent work did not directly recharacterize KT2440 TrpC, but 2024 shikimate-pathway rewiring in P. putida showed that flux through chorismate-derived aromatics can reach 89% of theoretical yield, highlighting the modern engineering context in which trpC-associated branch control remains important. (santos2024shikimatepathwaydependentcatabolism pages 1-5, bruinsma2024shikimatepathwaydependentcatabolism pages 1-4) dos Santos et al., 2024, https://doi.org/10.21203/rs.3.rs-4761679/v1; Bruinsma et al., 2024, bioRxiv, https://doi.org/10.1101/2024.07.06.602327

Table: This table summarizes the experimentally supported functional annotation of Pseudomonas putida KT2440 trpC (PP_0422; UniProt Q88QR6), including pathway role, operon context, phenotype, structural inference, and engineering relevance. It is useful as a compact evidence map for gene-function annotation.

9) Limitations and open gaps for KT2440 PP_0422/Q88QR6

  • Direct biochemical characterization of KT2440 TrpC (KM/kcat, substrate analog specificity, inhibition) was not located in the retrieved texts; mechanistic/kinetic details cited here derive from other bacterial IGPS homologs and should be treated as family-level inference rather than KT2440-specific parameters. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4, esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3)
  • Experimental cellular localization for KT2440 TrpC was not found; cytosolic localization is inferred from pathway role and typical bacterial metabolism, but not demonstrated here. (molinahenares2009functionalanalysisof pages 4-6)

Key sources (with URLs and publication dates)

  1. Molina-Henares MA et al. Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440. Microbial Biotechnology. Dec 2009. https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof media a5651e48)
  2. Kuepper J et al. Metabolic Engineering of Pseudomonas putida KT2440 to Produce Anthranilate from Glucose. Frontiers in Microbiology. Nov 2015. https://doi.org/10.3389/fmicb.2015.01310 (kuepper2015metabolicengineeringof pages 1-2)
  3. Esposito N et al. Indole-3-Glycerol Phosphate Synthase from Mycobacterium tuberculosis: A Potential New Drug Target. ChemBioChem. Sep 2022. https://doi.org/10.1002/cbic.202100314 (esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3, esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4)
  4. Bruinsma L et al. Shikimate pathway-Dependent Catabolism: enabling near-to-maximum production yield of aromatics. bioRxiv preprint. Jul 2024. https://doi.org/10.1101/2024.07.06.602327 (bruinsma2024shikimatepathwaydependentcatabolism pages 1-4, bruinsma2024shikimatepathwaydependentcatabolism pages 4-6)
  5. dos Santos VMPM et al. Shikimate pathway-Dependent Catabolism: enabling near-to-maximum production yield of aromatics. (preprint/posted version). Aug 2024. https://doi.org/10.21203/rs.3.rs-4761679/v1 (santos2024shikimatepathwaydependentcatabolism pages 1-5)
  6. Mutz M et al. Microbial synthesis of the plant natural product precursor p-coumaric acid with Corynebacterium glutamicum. Microbial Cell Factories. Oct 2023. https://doi.org/10.1186/s12934-023-02222-y (mutz2023microbialsynthesisof pages 1-2)

References

  1. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  2. (molinahenares2009functionalanalysisof pages 1-2): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  3. (barona‐gomez2003occurrenceofa pages 1-2): Francisco Barona‐Gómez and David A Hodgson. Occurrence of a putative ancient‐like isomerase involved in histidine and tryptophan biosynthesis. EMBO reports, 4:296-300, Mar 2003. URL: https://doi.org/10.1038/sj.embor.embor771, doi:10.1038/sj.embor.embor771. This article has 119 citations and is from a highest quality peer-reviewed journal.

  4. (esposito2022indole‐3‐glycerolphosphatesynthase pages 1-3): Nikolas Esposito, David W. Konas, and Nina M. Goodey. Indole‐3‐glycerol phosphate synthase from mycobacterium tuberculosis: a potential new drug target. Sep 2022. URL: https://doi.org/10.1002/cbic.202100314, doi:10.1002/cbic.202100314. This article has 5 citations and is from a peer-reviewed journal.

  5. (molinahenares2009functionalanalysisof pages 4-6): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  6. (molinahenares2009functionalanalysisof media a5651e48): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  7. (molinahenares2009functionalanalysisof media 77d9c78a): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  8. (esposito2022indole‐3‐glycerolphosphatesynthase pages 4-6): Nikolas Esposito, David W. Konas, and Nina M. Goodey. Indole‐3‐glycerol phosphate synthase from mycobacterium tuberculosis: a potential new drug target. Sep 2022. URL: https://doi.org/10.1002/cbic.202100314, doi:10.1002/cbic.202100314. This article has 5 citations and is from a peer-reviewed journal.

  9. (esposito2022indole‐3‐glycerolphosphatesynthase pages 3-4): Nikolas Esposito, David W. Konas, and Nina M. Goodey. Indole‐3‐glycerol phosphate synthase from mycobacterium tuberculosis: a potential new drug target. Sep 2022. URL: https://doi.org/10.1002/cbic.202100314, doi:10.1002/cbic.202100314. This article has 5 citations and is from a peer-reviewed journal.

  10. (kuepper2015metabolicengineeringof pages 1-2): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

  11. (bruinsma2024shikimatepathwaydependentcatabolism pages 1-4): Lyon Bruinsma, Christos Batianis, Sara Moreno Paz, Kesi Kurnia, Job. J Dirkmaat, Alexandra Müller, Jose Juncosa Nunez, Ruud A. Weusthuis, and Vitor A. P. Martins dos Santos. Shikimate pathway-dependent catabolism: enabling near-to-maximum production yield of aromatics. BioRxiv, Jul 2024. URL: https://doi.org/10.1101/2024.07.06.602327, doi:10.1101/2024.07.06.602327. This article has 0 citations.

  12. (santos2024shikimatepathwaydependentcatabolism pages 1-5): Vitor Martins dos Santos, Lyon Bruinsma, Christos Batianis, Sara Moreno-Paz, Kesi Kurnia, Job Dirkmaat, Alexandra Müller, Jose Juncosa Nuñez, and Ruud Weusthuis. Shikimate pathway-dependent catabolism: enabling near-to-maximum production yield of aromatics. Unknown journal, Aug 2024. URL: https://doi.org/10.21203/rs.3.rs-4761679/v1, doi:10.21203/rs.3.rs-4761679/v1.

  13. (bruinsma2024shikimatepathwaydependentcatabolism pages 4-6): Lyon Bruinsma, Christos Batianis, Sara Moreno Paz, Kesi Kurnia, Job. J Dirkmaat, Alexandra Müller, Jose Juncosa Nunez, Ruud A. Weusthuis, and Vitor A. P. Martins dos Santos. Shikimate pathway-dependent catabolism: enabling near-to-maximum production yield of aromatics. BioRxiv, Jul 2024. URL: https://doi.org/10.1101/2024.07.06.602327, doi:10.1101/2024.07.06.602327. This article has 0 citations.

  14. (mutz2023microbialsynthesisof pages 1-2): Mario Mutz, Dominic Kösters, Benedikt Wynands, Nick Wierckx, and Jan Marienhagen. Microbial synthesis of the plant natural product precursor p-coumaric acid with corynebacterium glutamicum. Microbial Cell Factories, Oct 2023. URL: https://doi.org/10.1186/s12934-023-02222-y, doi:10.1186/s12934-023-02222-y. This article has 24 citations and is from a peer-reviewed journal.

  15. (molinahenares2009functionalanalysisof pages 7-8): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  16. (kursula2003crystallographicstudieson pages 28-30): I Kursula. Crystallographic studies on the structure-function relationships in triosephosphate isomerase. Unknown journal, 2003.

Artifacts

Citations

  1. molinahenares2009functionalanalysisof pages 4-6
  2. kuepper2015metabolicengineeringof pages 1-2
  3. bruinsma2024shikimatepathwaydependentcatabolism pages 4-6
  4. mutz2023microbialsynthesisof pages 1-2
  5. molinahenares2009functionalanalysisof pages 2-4
  6. molinahenares2009functionalanalysisof pages 1-2
  7. santos2024shikimatepathwaydependentcatabolism pages 1-5
  8. bruinsma2024shikimatepathwaydependentcatabolism pages 1-4
  9. molinahenares2009functionalanalysisof pages 7-8
  10. kursula2003crystallographicstudieson pages 28-30
  11. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  12. https://doi.org/10.1002/cbic.202100314;
  13. https://doi.org/10.1038/sj.embor.embor771
  14. https://doi.org/10.1002/cbic.202100314
  15. https://doi.org/10.1111/j.1751-7915.2008.00062.x;
  16. https://doi.org/10.3389/fmicb.2015.01310
  17. https://doi.org/10.21203/rs.3.rs-4761679/v1;
  18. https://doi.org/10.1101/2024.07.06.602327
  19. https://doi.org/10.21203/rs.3.rs-4761679/v1
  20. https://doi.org/10.1186/s12934-023-02222-y
  21. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  22. https://doi.org/10.1038/sj.embor.embor771,
  23. https://doi.org/10.1002/cbic.202100314,
  24. https://doi.org/10.3389/fmicb.2015.01310,
  25. https://doi.org/10.1101/2024.07.06.602327,
  26. https://doi.org/10.21203/rs.3.rs-4761679/v1,
  27. https://doi.org/10.1186/s12934-023-02222-y,

📄 View Raw YAML

id: Q88QR6
gene_symbol: trpC
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: >-
  Indole-3-glycerol phosphate synthase (IGPS; TrpC; EC 4.1.1.48), the enzyme
  catalyzing the fourth step of L-tryptophan biosynthesis. It converts
  1-(2-carboxyphenylamino)-1-deoxy-D-ribulose 5-phosphate (CdRP) into
  (1S,2R)-1-C-(indol-3-yl)glycerol 3-phosphate (indole-3-glycerol phosphate, IGP),
  releasing CO2 and water in an irreversible decarboxylative ring-closure reaction.
  This indole-ring-forming step lies downstream of anthranilate
  phosphoribosyltransferase (TrpD) and upstream of the tryptophan synthase subunits
  (TrpA/TrpB). The protein adopts a classic (beta/alpha)8 TIM-barrel
  (ribulose-phosphate-binding barrel) fold and acts as a soluble cytoplasmic
  metabolic enzyme. In Pseudomonas putida KT2440 the gene (PP_0422) lies in the
  trpGDC operon, and loss-of-function insertions cause tryptophan auxotrophy. The
  enzyme family is highly conserved across bacteria and has no human homolog.
existing_annotations:
- term:
    id: GO:0000162
    label: L-tryptophan biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: TrpC/IGPS catalyzes the fourth step of L-tryptophan biosynthesis; this process annotation is correct and represents a core function of the gene.
    action: ACCEPT
    reason: Matches the UniProt-curated pathway annotation (L-tryptophan from chorismate, step 4/5) and is supported by tryptophan-auxotrophy phenotypes of PP_0422 insertion mutants in P. putida KT2440. This is a core function.
- term:
    id: GO:0004425
    label: indole-3-glycerol-phosphate synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: This is the canonical molecular function of TrpC, matching the UniProt RecName (Indole-3-glycerol phosphate synthase, EC 4.1.1.48) and the curated catalytic activity (CdRP to indole-3-glycerol phosphate + CO2 + H2O).
    action: ACCEPT
    reason: Directly supported by UniProt HAMAP-Rule MF_00134, InterPro IGPS domain signatures (IPR001468/IPR013798/IPR045186), EC 4.1.1.48 and RHEA:23476. Core molecular function.
- term:
    id: GO:0004640
    label: phosphoribosylanthranilate isomerase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: enables
  review:
    summary: Phosphoribosylanthranilate isomerase (PRAI, TrpF, EC 5.3.1.24) catalyzes the third step of tryptophan biosynthesis and is a distinct activity from IGPS. This activity is NOT supported for P. putida KT2440 TrpC by UniProt, which assigns only IGPS (EC 4.1.1.48). The annotation is a TreeGrafter over-propagation arising because some bacteria (e.g. E. coli) have a bifunctional TrpC(F) protein, whereas in Pseudomonas TrpF is a separate gene.
    action: REMOVE
    reason: UniProt Q88QR6 (HAMAP MF_00134) annotates only the monofunctional IGPS activity (277 aa, single IGPS domain), with no PRAI/TrpF domain. The IEA TreeGrafter inference reflects the bifunctional TrpCF architecture of some lineages and is not applicable to this monofunctional Pseudomonas enzyme. This is an electronic (IEA) prediction argued against on biological/domain grounds, not the second-guessing of an experimental annotation.
core_functions:
- description: Catalyzes the fourth step of L-tryptophan biosynthesis, the decarboxylative ring closure of CdRP to indole-3-glycerol phosphate.
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: UniProt RecName Indole-3-glycerol phosphate synthase, EC 4.1.1.48; catalytic activity CdRP to (indol-3-yl)glycerol 3-phosphate + CO2 + H2O (RHEA:23476).
  molecular_function:
    id: GO:0004425
    label: indole-3-glycerol-phosphate synthase activity
  directly_involved_in:
  - id: GO:0000162
    label: L-tryptophan biosynthetic process
references:
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
  reference_review:
    relevance: LOW
    correctness: VERIFIED
    review_notes: TreeGrafter pipeline reference; correctly identifies the source of the PRAI over-propagated annotation, which is not applicable to this monofunctional enzyme.
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: Source of the IGPS and tryptophan-biosynthesis annotations, both consistent with UniProt HAMAP-Rule MF_00134.
- id: file:PSEPK/trpC/trpC-deep-research-falcon.md
  title: Deep research report on P. putida KT2440 trpC (PP_0422)
  findings:
  - statement: PP_0422 (trpC) lies in the trpGDC operon; a mini-Tn5 insertion in PP_0422 yields tryptophan auxotrophy in KT2440, confirming its essential role in endogenous tryptophan biosynthesis.
    supporting_text: Molina-Henares et al. 2009 (Microbial Biotechnology 2:91-100) operon mapping, RT-PCR cotranscription, and auxotrophy screen.
  reference_review:
    relevance: HIGH
    correctness: UNVERIFIED
    review_notes: Deep-research summary citing Molina-Henares et al. 2009 (doi 10.1111/j.1751-7915.2008.00062.x) and Esposito et al. 2022 IGPS minireview. Underlying PMIDs not independently verified (PubMed lookup unavailable); conclusions are consistent with UniProt curation.
- id: PMID:12534463
  title: Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440
  findings: []
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: KT2440 genome reference (Nelson et al. 2002, Environ Microbiol) establishing the locus/gene assignment.
suggested_questions:
- question: Has the IGPS activity of P. putida KT2440 TrpC (PP_0422) been biochemically characterized (kcat/KM), or is the assignment based solely on homology and the auxotrophy phenotype?
suggested_experiments:
- description: Purify recombinant PP_0422 and measure IGPS activity (CdRP to IGP) to confirm the monofunctional assignment and exclude any cryptic PRAI activity.