trpD

UniProt ID: Q88QR7
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

Anthranilate phosphoribosyltransferase (TrpD, EC 2.4.2.18), a cytosolic Mg2+-dependent enzyme that catalyzes the second committed step of L-tryptophan biosynthesis. It transfers the phosphoribosyl group of 5-phospho-alpha-D-ribose 1-diphosphate (PRPP) to anthranilate, producing N-(5-phospho-beta-D-ribosyl)-anthranilate (PRA) and diphosphate. In Pseudomonas putida KT2440 the enzyme is a monofunctional anthranilate phosphoribosyltransferase (the glutamine amidotransferase component of anthranilate synthase is encoded by a separate gene, trpG), belongs to the anthranilate phosphoribosyltransferase family (HAMAP MF_00211), and acts as a homodimer binding two magnesium ions per monomer. Loss of trpD function causes tryptophan auxotrophy that is rescued by L-tryptophan or indole, consistent with its position upstream of the indole/tryptophan branch in the chorismate-derived aromatic amino-acid biosynthetic route.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000162 L-tryptophan biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: TrpD catalyzes the second step of tryptophan biosynthesis (anthranilate to PRA). This is the core biological process for this gene, supported by the conserved pathway role and by experimental tryptophan auxotrophy of P. putida KT2440 trpD mutants.
Reason: Correct core biological process. Consistent with the UniProt pathway annotation (L-tryptophan from chorismate, step 2/5) and with experimental auxotrophy/rescue evidence in KT2440.
GO:0000287 magnesium ion binding
IEA
GO_REF:0000104
ACCEPT
Summary: The enzyme binds two Mg2+ ions per monomer, which assist PRPP binding. UniProt documents specific Mg2+ binding residues (positions 94, 227, 228).
Reason: Magnesium is a required cofactor for this PRPP-dependent phosphoribosyltransferase; supported by HAMAP rule and modeled metal-binding sites.
GO:0004048 anthranilate phosphoribosyltransferase activity
IEA
GO_REF:0000120
ACCEPT
Summary: This is the specific, defining molecular function of TrpD (EC 2.4.2.18), transferring the phosphoribosyl group of PRPP to anthranilate to form PRA.
Reason: Core molecular function. Maps directly to EC 2.4.2.18 / RHEA:11768 and the anthranilate phosphoribosyltransferase family assignment (HAMAP MF_00211).
GO:0005829 cytosol
IEA
GO_REF:0000118
ACCEPT
Summary: As a soluble biosynthetic enzyme in the tryptophan pathway, TrpD acts in the cytosol. No experimental localization in KT2440, but this is the parsimonious and family-consistent compartment.
Reason: Appropriate cellular component for a cytoplasmic amino-acid biosynthetic enzyme; consistent with TreeGrafter inference across the family.
GO:0016757 glycosyltransferase activity
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: This is a high-level grouping term derived from the InterPro "Glycosyl transferase family 3" domain. The actual activity is the more specific anthranilate phosphoribosyltransferase (a pentosyltransferase, GO:0016763 branch), already captured by GO:0004048.
Reason: Uninformative parent term that does not accurately describe the enzyme's function; the specific activity (GO:0004048) and its proper parent pentosyltransferase (GO:0016763) are already annotated. Retaining glycosyltransferase activity is redundant and potentially misleading.
GO:0016763 pentosyltransferase activity
IEA
GO_REF:0000117
KEEP AS NON CORE
Summary: Correct but general parent of the specific anthranilate phosphoribosyltransferase activity. EC 2.4.2.18 is a pentosyltransferase, so this term is accurate.
Reason: True but uninformative relative to the specific GO:0004048 annotation; retain as a correct higher-level grouping rather than the core function.

Core Functions

Catalyzes the Mg2+-dependent transfer of the phosphoribosyl group of PRPP to anthranilate, producing N-(5-phospho-beta-D-ribosyl)-anthranilate (PRA), the second committed step of L-tryptophan biosynthesis.

Supporting Evidence:

References

Gene Ontology annotation through association of InterPro records with GO terms
Electronic Gene Ontology annotations created by transferring manual GO annotations between related proteins based on shared sequence features
Electronic Gene Ontology annotations created by ARBA machine learning models
TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440.
  • PP_0421 is trpD encoding anthranilate phosphoribosyltransferase; a transposon insertion at codon 184 of trpD causes tryptophan auxotrophy, and trpD lies in a trpGDC operon. The trpD mutant is rescued by L-tryptophan and indole.
Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium through the screening of a genome-wide mutant library.
  • A trpD mutant is auxotrophic on M9 minimal medium with growth restored by L-tryptophan, confirming trpD is required for endogenous tryptophan biosynthesis.

Deep Research

Asta

(trpD-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-05T20:17:53.988721

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.731)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.662)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[2] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.718)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[3] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.714)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[4] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.690)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[5] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.685)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[6] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.679)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[7] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.676)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[8] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.669)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[9] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.667)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[10] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[11] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.660)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[12] Characterization of holins, the membrane proteins of coliphage ASEC2201: a genomewide in silico approach

  • Authors: Humaira Saeed, Sudhaker Padmesh, Aditi Singh, S. Singh, Mohammed Haris Siddiqui et al.
  • Year: 2025
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/a39392e12bf3bda67bdfe600053e8403deb3b887
  • DOI: 10.3389/fmicb.2025.1550594
  • PMID: 40703241
  • PMCID: 12283622
  • Citations: 3
  • Summary: In silico identification of cell-penetrating peptide motifs within the holin sequences suggests potential for enhanced intracellular delivery in CPP-fusion therapeutic constructs and demonstrates the potential of integrative in silico approaches in developing a comprehensive foundation for future experimental validation for proteins with no prior functional annotation.
  • Evidence snippets:
  • Snippet 1 (score: 0.654)
    > Protein-coding gene annotation is typically a two-step process. Initially, Prodigal is employed to identify open reading frames (ORFs) by locating gene coordinates, but it does not infer gene function. To assign putative functions, Prokka performs hierarchical annotation by comparing candidate genes to curated protein databases. It begins with a user-supplied, high-confidence protein set, using BLAST+ for sequence similarity searches. If no match is found, it progresses to UniProt's verified bacterial proteins, covering \~ 16,000 sequences, and then optionally to RefSeq proteins specific to the organism's genuscapturing nomenclature consistency. When sequence-based annotation fails, Prokka applies profile-based searches using HMMER's hmmscan to query against Pfam and TIGRFAMs databases. An e-value threshold of 10 −6 is consistently applied to ensure significance. If no reliable match is found across all levels, the gene is designated as a "hypothetical protein. " This layered strategy maximizes annotation accuracy and functional insight across diverse bacterial genomes (Seemann, 2014).
    > The genome of coliphage ASEC2201 has been analyzed and three holin protein coding genes were selected. The sequences of all three holin proteins were retrieved from NCBI using accession no. SRX17770782 in the FASTA format. The sequence similarity search was performed via BLAST against the non-redundant database (Altschul et al., 1990).

[13] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.653)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[14] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.646)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[15] AgAnimalGenomes: browsers for viewing and manually annotating farm animal genomes

  • Authors: D. Triant, Amy T. Walsh, Gabrielle Hartley, B. Petry, Morgan R. Stegemiller et al.
  • Year: 2023
  • Venue: Mammalian Genome
  • URL: https://www.semanticscholar.org/paper/38a969fd5641e503106cb215010f84ea0a271f99
  • DOI: 10.1007/s00335-023-10008-1
  • PMID: 37460664
  • PMCID: 10382368
  • Citations: 5
  • Summary: This work presents genome visualization and annotation tools to support seven livestock species, available in a new resource called AgAnimalGenomes, and describes the data and search methods available and how to use the provided tools to edit and create new gene models.
  • Evidence snippets:
  • Snippet 1 (score: 0.645)
    > As previously described (Triant et al. 2020), once a proteincoding gene annotation is complete, each new or modified isoform should be compared to a well-curated protein sequence database to check for congruency with known proteins. The sequence of an annotation is obtained by right clicking it and selecting Get Sequence. The first choice of database to search is the well-curated UniProtKB/Swissprot database using BLAST at either the UniProt (https:// www. unipr ot. org/ blast) or NCBI website (https:// blast. ncbi. nlm. nih. gov/ Blast. cgi) (Sayers et al. 2023a;UniProt Consortium 2023). If there is no match with a significant e-value (< 1e−05) in UniProtKB/Swissprot, the next database to try is the Model Organisms (landmark) database at NCBI. If that fails, select the RefSeq Proteins database and exclude your organism of interest from the search. Although RefSeq includes computationally predicted and hypothetical proteins, an alignment to a homologous protein from another organism provides support for the annotation. An alignment that covers the full length of both the annotated protein and the database protein sequence suggests the annotation is correct. An alignment that encompasses the full length of an annotated protein sequence but only part of a database protein suggests that the annotation is truncated. You may be able to correct the annotation with additional evidence, but if there is not sufficient evidence the issue can be noted in the Annotation Information Panel under the Comment tab. A partial alignment of an annotated protein to a database protein suggests the annotation has a reading frame shift or was extended incorrectly. Aligning the coding sequence (CDS) to the protein database will reveal whether the problem is due to a reading frame shift. Further annotation editing should be performed to correct the reading frame. If an incorrect extension was due to the merging of two genes, you should edit or redo the annotation. Any unresolved issues should be entered in the Comment section of the Annotation Information Panel.

[16] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.643)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[17] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.640)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[18] A Genome-Wide Association Study Identifying Novel Genetic Markers of Response to Treatment with Interleukin-23 Inhibitors in Psoriasis

  • Authors: Sophia Zachari, K. Liadaki, Angeliki Planaki, E. Zafiriou, Olga Kouvarou et al.
  • Year: 2025
  • Venue: Genes
  • URL: https://www.semanticscholar.org/paper/d5f656311b54e222e7487ea32a061869b30178a1
  • DOI: 10.3390/genes16101195
  • PMID: 41153410
  • PMCID: 12564705
  • Summary: These findings provide promising pharmacogenetic markers which, upon validation in larger, independent cohorts, will enable the translation of a patient’s genotype into a response phenotype, thereby guiding clinical decisions and improving drug effectiveness.
  • Evidence snippets:
  • Snippet 1 (score: 0.638)
    > The UniProt knowledgebase (www.uniprot.org/uniprotkb/), (accessed on 20 June 2025), the central hub for the collection of functional information on proteins, with accurate and rich annotation [33], was used to retrieve the approved human gene and protein names and symbols.

[19] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.637)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  2. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  3. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  4. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  5. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  6. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  7. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  8. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  9. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  10. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  11. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  12. Humaira Saeed, Sudhaker Padmesh, Aditi Singh, S. Singh, Mohammed Haris Siddiqui et al. (2025). Characterization of holins, the membrane proteins of coliphage ASEC2201: a genomewide in silico approach. Frontiers in Microbiology. https://www.semanticscholar.org/paper/a39392e12bf3bda67bdfe600053e8403deb3b887
  13. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  14. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  15. D. Triant, Amy T. Walsh, Gabrielle Hartley, B. Petry, Morgan R. Stegemiller et al. (2023). AgAnimalGenomes: browsers for viewing and manually annotating farm animal genomes. Mammalian Genome. https://www.semanticscholar.org/paper/38a969fd5641e503106cb215010f84ea0a271f99
  16. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  17. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  18. Sophia Zachari, K. Liadaki, Angeliki Planaki, E. Zafiriou, Olga Kouvarou et al. (2025). A Genome-Wide Association Study Identifying Novel Genetic Markers of Response to Treatment with Interleukin-23 Inhibitors in Psoriasis. Genes. https://www.semanticscholar.org/paper/d5f656311b54e222e7487ea32a061869b30178a1
  19. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508

Falcon

(trpD-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 29 citations 2 artifacts 2026-06-11T21:28:34.856602

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Comprehensive research report: trpD (UniProt Q88QR7; locus PP_0421) in Pseudomonas putida KT2440

0) Executive summary (identity verified)

The target protein TrpD in Pseudomonas putida KT2440 (UniProt Q88QR7, ordered locus PP_0421) is anthranilate phosphoribosyltransferase (EC 2.4.2.18), a cytosolic PRPP-dependent phosphoribosyltransferase in the tryptophan de novo biosynthesis pathway. In KT2440, trpD is genetically and transcriptionally linked to trpG and trpC in a trpGDC operon, and loss-of-function causes tryptophan auxotrophy that can be rescued by pathway end products/intermediates (L-tryptophan; indole), consistent with the canonical pathway position. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molina‐henares2010identificationofconditionally pages 6-7)


1) Key concepts and definitions (current understanding)

1.1 Definition of TrpD activity and reaction

Anthranilate phosphoribosyltransferase (TrpD; EC 2.4.2.18) catalyzes transfer of a phosphoribosyl group from PRPP (5-phospho-α-D-ribose-1-diphosphate) to anthranilate, yielding N-(5-phospho-β-D-ribosyl)-anthranilate (PRA) and pyrophosphate (PPi). (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8)

This step is an early committed reaction in the tryptophan biosynthesis route from chorismate: anthranilate is produced by TrpE/TrpG (anthranilate synthase components), and TrpD converts anthranilate into the ribosylated intermediate that proceeds through subsequent transformations toward the indole ring and ultimately tryptophan. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8)

1.2 Enzyme family, structural features, and mechanism (inference framework)

A widely cited enzymology review of PRPP-dependent enzymes describes TrpD enzymes as typically homodimeric proteins with a two-domain architecture (N-terminal helical domain and larger C-terminal α/β domain) that creates an active-site cleft for binding PRPP and anthranilate. Divalent cations (Mg2+ or Mn2+) promote PRPP binding, and conserved motifs (including a glycine-rich region involved in phosphate binding) are characteristic. Structural studies in bacteria have supported a model involving substrate capture and movement of anthranilate through multiple binding sites/a channel toward the catalytic configuration that enables nucleophilic attack on PRPP. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8, hovejensen2017phosphoribosyldiphosphate(prpp) pages 78-78)

These mechanistic and fold-level features are not specific measurements on P. putida KT2440 TrpD in the retrieved documents, but they represent the authoritative current understanding used to support functional inference from homology for bacterial TrpD-family members (such as Q88QR7). (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8)


2) Target-gene verification and P. putida KT2440-specific functional evidence

2.1 Correct gene/protein mapping: PP_0421 = trpD = anthranilate phosphoribosyltransferase

A functional genetics study of aromatic amino-acid biosynthesis in P. putida KT2440 identifies PP_0421 as trpD, encoding anthranilate phosphoribosyltransferase, and reports a tryptophan-auxotrophic transposon mutant (Aux-3) with a mini-Tn5 insertion at the 184th codon of PP_0421/trpD, linking disruption of PP_0421 to tryptophan auxotrophy. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2)

This directly matches the UniProt identity provided by the user (Q88QR7; PP_0421; “Anthranilate phosphoribosyltransferase”). (molinahenares2009functionalanalysisof pages 2-4)

2.2 Operon/transcriptional organization in KT2440

In KT2440, trpD is embedded in a trpG–trpD–trpC transcriptional unit (trpGDC operon). RT-PCR evidence supports co-transcription across the trpG–trpD–trpC boundaries; the genomic spacing/overlap also supports operon logic (trpG–trpD separated by 9 nt; trpD and trpC overlapping by 6 nt). The same study reports that trpE is transcribed as a monocistronic unit distinct from trpGDC. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8)

Visual evidence for this gene organization and transcriptional-unit mapping is provided in the extracted Table/Figure crops. (molinahenares2009functionalanalysisof media 8eee5d0d, molinahenares2009functionalanalysisof media f71ebd0e)

2.3 Phenotype and pathway placement in KT2440

KT2440 trpD mutants are unable to grow on minimal medium and can be rescued by supplementation. In pathway-intermediate feeding experiments, the trpD mutant grows with tryptophan and indole, consistent with TrpD operating upstream of indole and tryptophan formation (i.e., loss blocks endogenous synthesis but can be bypassed by supplying downstream metabolites). (molinahenares2009functionalanalysisof pages 4-6)

In an independent genome-wide mutant library study focused on growth on minimal medium, a trpD mutant is auxotrophic on M9 and growth is restored by L-tryptophan (while D-tryptophan does not rescue), reinforcing that trpD is required for endogenous tryptophan biosynthesis under minimal conditions. (molina‐henares2010identificationofconditionally pages 6-7)

2.4 Regulation and cellular localization (what is known vs. not found here)

The KT2440 gene organization places tryptophan biosynthesis genes in multiple genomic regions. A regulatory gene trpI is described as divergently transcribed relative to trpAB and annotated as encoding a repressor, indicating local transcriptional regulation at the trpAB locus that is physically separated from trpGDC. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6)

No direct, KT2440-specific experimental evidence for subcellular localization of TrpD (e.g., fractionation, microscopy) was identified in the retrieved texts. Given its biosynthetic role and the general bacterial paradigm for amino-acid biosynthesis enzymes, the most parsimonious functional localization is intracellular (cytosolic), but this report flags that as inference rather than a demonstrated localization in the cited KT2440 sources. (molinahenares2009functionalanalysisof pages 2-4, hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52)


3) Pathways and biological processes

3.1 Tryptophan biosynthesis context

In bacteria, tryptophan is biosynthesized from chorismate through anthranilate and downstream intermediates; TrpD catalyzes the PRPP-dependent conversion of anthranilate to PRA, positioning it as a key node linking shikimate/chorismate-derived aromatic metabolism to the indole/tryptophan branch. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8)

For P. putida KT2440, genetic mapping and rescue phenotypes place trpD squarely in this canonical route, consistent with the conserved trp gene clusters described for pseudomonads. (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8)

3.2 Crosstalk with PRPP metabolism (expert-level interpretation)

Because TrpD consumes PRPP, its effective capacity depends on PRPP supply, which is also demanded by many other phosphoribosyltransferases in nucleotide and amino-acid biosynthesis. A comprehensive PRPP review emphasizes PRPP as a widely shared substrate and discusses structural/mechanistic diversity among PRPP-utilizing enzymes, framing TrpD as one member of a broader PRPP-dependent metabolic network. From a systems viewpoint, TrpD can therefore become a flux-control point not only through its own kinetics and regulation but also through competition for PRPP and impacts on PPi handling. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52)


4) Recent developments (prioritizing 2023–2024) and expert analysis

Direct 2023–2024 studies specifically characterizing KT2440 PP_0421/Q88QR7 were not retrieved here; however, 2023–2024 literature demonstrates that TrpD is a highly actionable control node in microbial engineering (anthranilate and tryptophan-derived product pipelines). These applied findings provide strong, contemporary support for TrpD’s functional centrality and substrate role.

4.1 2023: using trpD disruption to accumulate anthranilate (platform chemical)

Kim et al. engineered E. coli for anthranilate overproduction and explicitly disrupted trpD (described as transferring the phosphoribosyl group to anthranilate) to prevent anthranilate consumption and promote accumulation. In a 7-L fed-batch fermentation, they report ~4 g/L anthranilate. This is an application-level validation that blocking the TrpD step increases anthranilate accumulation in vivo. (kim2023engineeredescherichiacoli pages 1-2)

In their contextualization of the field, the authors also cite prior high performance titers including up to 14 g/L anthranilate in an engineered E. coli strain and 1.5 g/L in a Pseudomonas putida ΔtrpDC strain (the latter supports that trpD-region perturbations are used directly in P. putida backgrounds, though not specifically the KT2440 PP_0421 allele). (kim2023engineeredescherichiacoli pages 2-4)

4.2 2024: tuning trpD translation to optimize anthranilate production

Mutz et al. engineered Corynebacterium glutamicum for anthranilate production and report that adjusting translation efficiency of trpD (anthranilate phosphoribosyltransferase) along with aroK improved production; their final strain accumulated up to 5.9 g/L (43 mM) anthranilate in bioreactors. This demonstrates modern pathway-balancing strategies that treat TrpD as a tunable valve controlling anthranilate drainage to downstream tryptophan. (mutz2024metabolicengineeringof pages 1-3)

4.3 2024: engineering TrpD variants for improved tryptophan availability and specialty products

Putri et al. report a TrpD A162D variant in C. glutamicum described as feedback-resistant to L-tryptophan and with increased substrate affinity versus wild-type. Incorporating this engineering, they report 3.1 g/L L-tryptophan in flask culture (with multiple copies/expressions of the variant), enabling downstream biosynthesis of a halogenated natural-product derivative (APRN) with ~28.1–29.5 mg/L titers. This illustrates recent, protein-level engineering of TrpD beyond knockouts or expression tuning. (putri2024fermentativeaminopyrrolnitrinproduction pages 1-3)

4.4 2024 expert synthesis: TrpD as a co-overexpression target in industrial tryptophan production

A 2024 review on fermentation strategies for tryptophan production in E. coli highlights that the trp operon enzymes are common genetic targets and reports an example where co-overexpression of trpE and trpD produced tryptophan titers up to 45.6 g/L, underscoring that TrpD is not only a drain on anthranilate (when knocked out) but also a needed capacity step when the objective is maximal tryptophan. (ramosvaldovinos2024optimizingfermentationstrategies pages 7-8)


5) Current applications and real-world implementations

5.1 Biomanufacturing (anthranilate as a platform chemical)

Anthranilate is described as a platform chemical relevant to food ingredients, dyes, perfumes, agrochemicals, pharmaceuticals, and plastics; engineering strategies frequently include disrupting trpD to prevent conversion of anthranilate to PRA, thereby increasing anthranilate accumulation, as validated by 7-L fed-batch production at ~4 g/L in E. coli. (kim2023engineeredescherichiacoli pages 1-2)

5.2 Biomanufacturing (tryptophan and tryptophan-derived compounds)

Recent work in C. glutamicum demonstrates that engineering TrpD itself (A162D) can raise tryptophan availability (3.1 g/L) and serve as a foundation for production of tryptophan-derived halogenated specialty metabolites (APRN at ~28–29.5 mg/L). (putri2024fermentativeaminopyrrolnitrinproduction pages 1-3)

5.3 Drug discovery relevance (contextual)

Structural and mechanistic work on bacterial TrpD—especially in pathogens such as Mycobacterium tuberculosis—has informed inhibitor design through characterization of substrate capture and active-site conformational changes. While not directly a KT2440 application, this is an expert-level rationale for why TrpD remains a target of interest for antimicrobial strategies. (parthasarathy2018athreeringcircus pages 6-8, hovejensen2017phosphoribosyldiphosphate(prpp) pages 78-78)


6) Statistics and data highlights (from cited studies)

  • KT2440 genetic mapping: Aux-3 mini-Tn5 insertion at trpD (PP_0421) 184th codon causing tryptophan auxotrophy. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof media 8eee5d0d)
  • KT2440 operon structure: trpGDC operon; trpG–trpD separated by 9 nt; trpD–trpC overlap by 6 nt; RT-PCR supports co-transcription. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6)
  • KT2440 phenotype: trpD mutant growth rescued by L-tryptophan and indole. (molinahenares2009functionalanalysisof pages 4-6, molina‐henares2010identificationofconditionally pages 6-7)
  • Anthranilate titers (2023–2024): ~4 g/L (E. coli, 7-L fed-batch; trpD disrupted). (kim2023engineeredescherichiacoli pages 1-2)
  • Anthranilate titers (2024): up to 5.9 g/L (43 mM) (C. glutamicum, bioreactor; includes trpD translation tuning). (mutz2024metabolicengineeringof pages 1-3)
  • Tryptophan and derivative titers (2024): 3.1 g/L L-tryptophan (C. glutamicum; TrpD A162D engineering) and ~28.1–29.5 mg/L APRN. (putri2024fermentativeaminopyrrolnitrinproduction pages 1-3)
  • Industrial-context tryptophan titer (reviewed 2024): up to 45.6 g/L with trpE/trpD co-overexpression in an E. coli example compiled in review. (ramosvaldovinos2024optimizingfermentationstrategies pages 7-8)

7) Evidence summary table

The table below consolidates organism-specific evidence for KT2440 PP_0421/Q88QR7 versus general mechanistic knowledge and recent engineering applications.

Claim Key evidence and quantitative data Organism context Source with year and URL
Correct gene identity and function PP_0421 in Pseudomonas putida KT2440 is identified as trpD, encoding anthranilate phosphoribosyltransferase; Aux-3 carries a mini-Tn5 insertion at the 184th codon of trpD/PP0421, linking disruption to tryptophan auxotrophy. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2) P. putida KT2440 (target gene Q88QR7 / PP_0421) Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Biochemical reaction Anthranilate phosphoribosyltransferase TrpD (EC 2.4.2.18) transfers the phosphoribosyl group from PRPP to anthranilate to form N-(5-phospho-β-D-ribosyl)-anthranilate (PRA) plus PPi; this is an early committed step in tryptophan biosynthesis. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8) General bacterial/archaeal TrpD framework used to interpret the P. putida ortholog Hove-Jensen et al., 2017, https://doi.org/10.1128/mmbr.00040-16; Parthasarathy et al., 2018, https://doi.org/10.3389/fmolb.2018.00029
Pathway role In P. putida, trpD acts downstream of anthranilate formation and upstream of later indole/tryptophan steps; mutant rescue by tryptophan and indole is consistent with this position in the pathway. (molinahenares2009functionalanalysisof pages 4-6) P. putida KT2440 Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Operon context trpGDC forms an operon in P. putida KT2440; trpE and trpF are monocistronic. Physical organization supports co-transcription: trpG–trpD separated by 9 nt and trpD/trpC overlap by 6 nt; RT-PCR supports a contiguous trpG-trpD-trpC transcript. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8, molinahenares2009functionalanalysisof media 8eee5d0d) P. putida KT2440; conserved organization also noted across Pseudomonas spp. Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Phenotype / essentiality for minimal growth A trpD mutant fails to grow on M9 minimal medium and is rescued by L-tryptophan; in another assay, trpD mutants grew with tryptophan and indole. A genome-wide mutant screen likewise identified trpD as conditionally essential for minimal-medium growth. (molinahenares2009functionalanalysisof pages 4-6, molina‐henares2010identificationofconditionally pages 6-7) P. putida KT2440 Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x; Molina-Henares et al., 2010, https://doi.org/10.1111/j.1462-2920.2010.02166.x
Regulatory / transcriptional context The trpAB locus is separate and transcribed divergently from trpI, a repressor-like regulator; trpGDC is in a distinct transcriptional unit from trpE, indicating split pathway organization and local regulatory separation. Direct specific regulation of PP_0421 itself was not detailed beyond operon structure. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 9-10) P. putida KT2440 with comparative pointers to fluorescent pseudomonads Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Structural/mechanistic features supporting annotation TrpD enzymes are typically homodimeric, with a two-domain fold, Mg2+/Mn2+-assisted PRPP binding, and a conserved glycine-rich PRPP-binding motif; structural studies show anthranilate-binding sites/tunnel features relevant to catalysis. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, hovejensen2017phosphoribosyldiphosphate(prpp) pages 78-78) General bacterial TrpD knowledge; not demonstrated directly for P. putida PP_0421 in the gathered snippets Hove-Jensen et al., 2017, https://doi.org/10.1128/mmbr.00040-16
Localization No direct subcellular localization experiment for PP_0421/TrpD in P. putida was reported in the gathered evidence; the evidence supports a typical intracellular biosynthetic enzyme role rather than extracellular function. (molinahenares2009functionalanalysisof pages 2-4) P. putida KT2440 Molina-Henares et al., 2009, https://doi.org/10.1111/j.1751-7915.2008.00062.x
Application: anthranilate accumulation by blocking TrpD step In engineered E. coli, trpD disruption was used to accumulate anthranilate, reaching about 4 g/L anthranilate in 7-L fed-batch fermentation; the study also cites prior values of 14 g/L in another engineered E. coli strain and 1.5 g/L in P. putida ΔtrpDC as context. (kim2023engineeredescherichiacoli pages 1-2, kim2023engineeredescherichiacoli pages 2-4) Other bacteria (E. coli primary; P. putida cited as comparison, not PP_0421-specific functional proof) Kim et al., 2023, https://doi.org/10.3389/fmicb.2023.1081221
Application: anthranilate production via trpD translation tuning In engineered Corynebacterium glutamicum, translation-efficiency modulation of trpD together with aroK improved anthranilate production; the final strain reached 5.9 g/L (43 mM) anthranilate in bioreactor cultivation. (mutz2024metabolicengineeringof pages 1-3) Other bacteria; demonstrates practical flux-control value of the TrpD step Mutz et al., 2024, https://doi.org/10.1111/1751-7915.14388
Application: engineered TrpD variant In C. glutamicum, TrpD A162D was reported as feedback-resistant to L-tryptophan with increased substrate affinity; strains carrying this variant produced 3.1 g/L L-tryptophan in flask culture, and downstream pathway engineering yielded ~28.1–29.5 mg/L APRN. (putri2024fermentativeaminopyrrolnitrinproduction pages 1-3) Other bacteria; shows TrpD can be engineered to improve tryptophan-derived product formation Putri et al., 2024, https://doi.org/10.1186/s12934-024-02424-y
Application: high-tryptophan production context A 2024 review reports that co-overexpression of trpE and trpD in an E. coli production background yielded 45.6 g/L tryptophan, underscoring TrpD as a relevant engineering node in industrial pathway balancing. (ramosvaldovinos2024optimizingfermentationstrategies pages 7-8) Other bacteria; review-level production context rather than direct P. putida evidence Ramos-Valdovinos and Martínez-Antonio, 2024, https://doi.org/10.3390/pr12112422

Table: This table summarizes direct and indirect evidence for the annotation of Pseudomonas putida KT2440 trpD (UniProt Q88QR7 / PP_0421), separating strain-specific genetic evidence from broader mechanistic and application-oriented TrpD literature. It is useful for distinguishing what is experimentally shown in the target organism versus what is inferred from authoritative studies in other bacteria.


8) Limitations of this evidence set (important for annotation confidence)

  • KT2440-specific enzymology: No direct enzyme kinetic parameters (Km, kcat, Ki) or structural data for P. putida KT2440 TrpD (Q88QR7) were retrieved in the accessible documents; mechanistic/structural claims are therefore grounded in authoritative general TrpD literature. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52, parthasarathy2018athreeringcircus pages 6-8)
  • Localization: No experimental localization data (e.g., cytosolic fractionation, localization tags) were identified for KT2440 TrpD in the retrieved texts. (molinahenares2009functionalanalysisof pages 2-4)
  • 2023–2024 KT2440-focused studies: Recent work was found primarily in other organisms (E. coli; C. glutamicum) or as contextual citations about Pseudomonas strains; this supports functional centrality and engineering relevance but does not replace organism-specific biochemical characterization. (kim2023engineeredescherichiacoli pages 1-2, mutz2024metabolicengineeringof pages 1-3, kim2023engineeredescherichiacoli pages 2-4)

9) Key references (with publication dates and URLs)

  • Molina-Henares et al. 2009-12. Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440. Microbial Biotechnology. https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6)
  • Molina-Henares et al. 2010-06. Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium... Environmental Microbiology. https://doi.org/10.1111/j.1462-2920.2010.02166.x (molina‐henares2010identificationofconditionally pages 6-7)
  • Hove-Jensen et al. 2017-03. PRPP: Biosynthesis, Enzymology, Utilization, and Metabolic Significance. Microbiology and Molecular Biology Reviews. https://doi.org/10.1128/mmbr.00040-16 (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52)
  • Parthasarathy et al. 2018-04. A Three-Ring Circus... Frontiers in Molecular Biosciences. https://doi.org/10.3389/fmolb.2018.00029 (parthasarathy2018athreeringcircus pages 6-8)
  • Kim et al. 2023-03. Engineered Escherichia coli cell factory for anthranilate over-production. Frontiers in Microbiology. https://doi.org/10.3389/fmicb.2023.1081221 (kim2023engineeredescherichiacoli pages 1-2)
  • Mutz et al. 2024-01. Metabolic engineering of Corynebacterium glutamicum for the production of anthranilate... Microbial Biotechnology. https://doi.org/10.1111/1751-7915.14388 (mutz2024metabolicengineeringof pages 1-3)
  • Putri et al. 2024-05. Fermentative aminopyrrolnitrin production... Microbial Cell Factories. https://doi.org/10.1186/s12934-024-02424-y (putri2024fermentativeaminopyrrolnitrinproduction pages 1-3)
  • Ramos-Valdovinos & Martínez-Antonio 2024-11. Optimizing Fermentation Strategies for Enhanced Tryptophan Production... Processes. https://doi.org/10.3390/pr12112422 (ramosvaldovinos2024optimizingfermentationstrategies pages 7-8)

References

  1. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  2. (molinahenares2009functionalanalysisof pages 4-6): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  3. (molina‐henares2010identificationofconditionally pages 6-7): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

  4. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 50-52): Bjarne Hove-Jensen, Kasper R. Andersen, Mogens Kilstrup, Jan Martinussen, Robert L. Switzer, and Martin Willemoës. Phosphoribosyl diphosphate (prpp): biosynthesis, enzymology, utilization, and metabolic significance. Microbiology and Molecular Biology Reviews, Mar 2017. URL: https://doi.org/10.1128/mmbr.00040-16, doi:10.1128/mmbr.00040-16. This article has 283 citations and is from a domain leading peer-reviewed journal.

  5. (parthasarathy2018athreeringcircus pages 6-8): Anutthaman Parthasarathy, Penelope J. Cross, Renwick C. J. Dobson, Lily E. Adams, Michael A. Savka, and André O. Hudson. A three-ring circus: metabolism of the three proteogenic aromatic amino acids and their role in the health of plants and animals. Frontiers in Molecular Biosciences, Apr 2018. URL: https://doi.org/10.3389/fmolb.2018.00029, doi:10.3389/fmolb.2018.00029. This article has 423 citations.

  6. (hovejensen2017phosphoribosyldiphosphate(prpp) pages 78-78): Bjarne Hove-Jensen, Kasper R. Andersen, Mogens Kilstrup, Jan Martinussen, Robert L. Switzer, and Martin Willemoës. Phosphoribosyl diphosphate (prpp): biosynthesis, enzymology, utilization, and metabolic significance. Microbiology and Molecular Biology Reviews, Mar 2017. URL: https://doi.org/10.1128/mmbr.00040-16, doi:10.1128/mmbr.00040-16. This article has 283 citations and is from a domain leading peer-reviewed journal.

  7. (molinahenares2009functionalanalysisof pages 1-2): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  8. (molinahenares2009functionalanalysisof pages 7-8): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  9. (molinahenares2009functionalanalysisof media 8eee5d0d): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  10. (molinahenares2009functionalanalysisof media f71ebd0e): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  11. (kim2023engineeredescherichiacoli pages 1-2): Hye-Jin Kim, Seung-Yeul Seo, Heung-Soon Park, Ji-Young Ko, Si-Sun Choi, Sang Joung Lee, and Eung-Soo Kim. Engineered escherichia coli cell factory for anthranilate over-production. Frontiers in Microbiology, Mar 2023. URL: https://doi.org/10.3389/fmicb.2023.1081221, doi:10.3389/fmicb.2023.1081221. This article has 9 citations and is from a peer-reviewed journal.

  12. (kim2023engineeredescherichiacoli pages 2-4): Hye-Jin Kim, Seung-Yeul Seo, Heung-Soon Park, Ji-Young Ko, Si-Sun Choi, Sang Joung Lee, and Eung-Soo Kim. Engineered escherichia coli cell factory for anthranilate over-production. Frontiers in Microbiology, Mar 2023. URL: https://doi.org/10.3389/fmicb.2023.1081221, doi:10.3389/fmicb.2023.1081221. This article has 9 citations and is from a peer-reviewed journal.

  13. (mutz2024metabolicengineeringof pages 1-3): Mario Mutz, Vincent Brüning, Christian Brüsseler, Moritz‐Fabian Müller, Stephan Noack, and Jan Marienhagen. Metabolic engineering of corynebacterium glutamicum for the production of anthranilate from glucose and xylose. Microbial Biotechnology, Jan 2024. URL: https://doi.org/10.1111/1751-7915.14388, doi:10.1111/1751-7915.14388. This article has 15 citations and is from a peer-reviewed journal.

  14. (putri2024fermentativeaminopyrrolnitrinproduction pages 1-3): Virginia Ryandini Melati Putri, Min-Hee Jung, Ji-Young Lee, Mi-Hyang Kwak, Theavita Chatarina Mariyes, Anastasia Kerbs, Volker F. Wendisch, Hee Jeong Kong, Young-Ok Kim, and Jin-Ho Lee. Fermentative aminopyrrolnitrin production by metabolically engineered corynebacterium glutamicum. Microbial Cell Factories, May 2024. URL: https://doi.org/10.1186/s12934-024-02424-y, doi:10.1186/s12934-024-02424-y. This article has 6 citations and is from a peer-reviewed journal.

  15. (ramosvaldovinos2024optimizingfermentationstrategies pages 7-8): Miguel Angel Ramos-Valdovinos and Agustino Martínez-Antonio. Optimizing fermentation strategies for enhanced tryptophan production in escherichia coli: integrating genetic and environmental controls for industrial applications. Processes, Nov 2024. URL: https://doi.org/10.3390/pr12112422, doi:10.3390/pr12112422. This article has 10 citations.

  16. (molinahenares2009functionalanalysisof pages 9-10): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

Artifacts

Citations

  1. molinahenares2009functionalanalysisof pages 2-4
  2. molinahenares2009functionalanalysisof pages 4-6
  3. kim2023engineeredescherichiacoli pages 1-2
  4. kim2023engineeredescherichiacoli pages 2-4
  5. mutz2024metabolicengineeringof pages 1-3
  6. putri2024fermentativeaminopyrrolnitrinproduction pages 1-3
  7. ramosvaldovinos2024optimizingfermentationstrategies pages 7-8
  8. parthasarathy2018athreeringcircus pages 6-8
  9. molinahenares2009functionalanalysisof pages 1-2
  10. molinahenares2009functionalanalysisof pages 7-8
  11. molinahenares2009functionalanalysisof pages 9-10
  12. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  13. https://doi.org/10.1128/mmbr.00040-16;
  14. https://doi.org/10.3389/fmolb.2018.00029
  15. https://doi.org/10.1111/j.1751-7915.2008.00062.x;
  16. https://doi.org/10.1111/j.1462-2920.2010.02166.x
  17. https://doi.org/10.1128/mmbr.00040-16
  18. https://doi.org/10.3389/fmicb.2023.1081221
  19. https://doi.org/10.1111/1751-7915.14388
  20. https://doi.org/10.1186/s12934-024-02424-y
  21. https://doi.org/10.3390/pr12112422
  22. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  23. https://doi.org/10.1111/j.1462-2920.2010.02166.x,
  24. https://doi.org/10.1128/mmbr.00040-16,
  25. https://doi.org/10.3389/fmolb.2018.00029,
  26. https://doi.org/10.3389/fmicb.2023.1081221,
  27. https://doi.org/10.1111/1751-7915.14388,
  28. https://doi.org/10.1186/s12934-024-02424-y,
  29. https://doi.org/10.3390/pr12112422,

📄 View Raw YAML

id: Q88QR7
gene_symbol: trpD
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: Anthranilate phosphoribosyltransferase (TrpD, EC 2.4.2.18), a cytosolic Mg2+-dependent enzyme that catalyzes the second committed step of L-tryptophan biosynthesis. It transfers the phosphoribosyl group of 5-phospho-alpha-D-ribose 1-diphosphate (PRPP) to anthranilate, producing N-(5-phospho-beta-D-ribosyl)-anthranilate (PRA) and diphosphate. In Pseudomonas putida KT2440 the enzyme is a monofunctional anthranilate phosphoribosyltransferase (the glutamine amidotransferase component of anthranilate synthase is encoded by a separate gene, trpG), belongs to the anthranilate phosphoribosyltransferase family (HAMAP MF_00211), and acts as a homodimer binding two magnesium ions per monomer. Loss of trpD function causes tryptophan auxotrophy that is rescued by L-tryptophan or indole, consistent with its position upstream of the indole/tryptophan branch in the chorismate-derived aromatic amino-acid biosynthetic route.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000104
  title: Electronic Gene Ontology annotations created by transferring manual GO annotations between related proteins based on shared sequence features
  findings: []
- id: GO_REF:0000117
  title: Electronic Gene Ontology annotations created by ARBA machine learning models
  findings: []
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:21261884
  title: Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440.
  findings:
  - statement: PP_0421 is trpD encoding anthranilate phosphoribosyltransferase; a transposon insertion at codon 184 of trpD causes tryptophan auxotrophy, and trpD lies in a trpGDC operon. The trpD mutant is rescued by L-tryptophan and indole.
    reference_section_type: RESULTS
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: 'Corrected identifier. The previous PMID:19825068 could not be resolved on PubMed. PMID:21261884 recovered from DOI 10.1111/j.1751-7915.2008.00062.x (Molina-Henares et al., Microbial Biotechnology 2:91-100, 2009) and PubMed-verified; cached title matches. Establishes PP_0421 = trpD, its role in tryptophan auxotrophy, and the trpGDC operon structure in KT2440.'
- id: PMID:20158506
  title: Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium through the screening of a genome-wide mutant library.
  findings:
  - statement: A trpD mutant is auxotrophic on M9 minimal medium with growth restored by L-tryptophan, confirming trpD is required for endogenous tryptophan biosynthesis.
    reference_section_type: RESULTS
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: 'Corrected identifier. The previous PMID:20236172 was a WRONG identifier (resolved to "Pregabalin in the treatment of post-traumatic peripheral neuropathic pain", an unrelated Eur J Neurol clinical trial). PMID:20158506 recovered from DOI 10.1111/j.1462-2920.2010.02166.x (Molina-Henares et al., Environmental Microbiology 2010) and PubMed-verified; cached title matches.'
existing_annotations:
- term:
    id: GO:0000162
    label: L-tryptophan biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: TrpD catalyzes the second step of tryptophan biosynthesis (anthranilate to PRA). This is the core biological process for this gene, supported by the conserved pathway role and by experimental tryptophan auxotrophy of P. putida KT2440 trpD mutants.
    action: ACCEPT
    reason: Correct core biological process. Consistent with the UniProt pathway annotation (L-tryptophan from chorismate, step 2/5) and with experimental auxotrophy/rescue evidence in KT2440.
- term:
    id: GO:0000287
    label: magnesium ion binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000104
  qualifier: enables
  review:
    summary: The enzyme binds two Mg2+ ions per monomer, which assist PRPP binding. UniProt documents specific Mg2+ binding residues (positions 94, 227, 228).
    action: ACCEPT
    reason: Magnesium is a required cofactor for this PRPP-dependent phosphoribosyltransferase; supported by HAMAP rule and modeled metal-binding sites.
- term:
    id: GO:0004048
    label: anthranilate phosphoribosyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: This is the specific, defining molecular function of TrpD (EC 2.4.2.18), transferring the phosphoribosyl group of PRPP to anthranilate to form PRA.
    action: ACCEPT
    reason: Core molecular function. Maps directly to EC 2.4.2.18 / RHEA:11768 and the anthranilate phosphoribosyltransferase family assignment (HAMAP MF_00211).
- term:
    id: GO:0005829
    label: cytosol
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: As a soluble biosynthetic enzyme in the tryptophan pathway, TrpD acts in the cytosol. No experimental localization in KT2440, but this is the parsimonious and family-consistent compartment.
    action: ACCEPT
    reason: Appropriate cellular component for a cytoplasmic amino-acid biosynthetic enzyme; consistent with TreeGrafter inference across the family.
- term:
    id: GO:0016757
    label: glycosyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: This is a high-level grouping term derived from the InterPro "Glycosyl transferase family 3" domain. The actual activity is the more specific anthranilate phosphoribosyltransferase (a pentosyltransferase, GO:0016763 branch), already captured by GO:0004048.
    action: MARK_AS_OVER_ANNOTATED
    reason: Uninformative parent term that does not accurately describe the enzyme's function; the specific activity (GO:0004048) and its proper parent pentosyltransferase (GO:0016763) are already annotated. Retaining glycosyltransferase activity is redundant and potentially misleading.
- term:
    id: GO:0016763
    label: pentosyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000117
  qualifier: enables
  review:
    summary: Correct but general parent of the specific anthranilate phosphoribosyltransferase activity. EC 2.4.2.18 is a pentosyltransferase, so this term is accurate.
    action: KEEP_AS_NON_CORE
    reason: True but uninformative relative to the specific GO:0004048 annotation; retain as a correct higher-level grouping rather than the core function.
core_functions:
- description: Catalyzes the Mg2+-dependent transfer of the phosphoribosyl group of PRPP to anthranilate, producing N-(5-phospho-beta-D-ribosyl)-anthranilate (PRA), the second committed step of L-tryptophan biosynthesis.
  supported_by:
  - reference_id: PMID:21261884
  molecular_function:
    id: GO:0004048
    label: anthranilate phosphoribosyltransferase activity
  directly_involved_in:
  - id: GO:0000162
    label: L-tryptophan biosynthetic process