trpA

UniProt ID: Q88RP7
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

trpA encodes the alpha subunit of tryptophan synthase (EC 4.2.1.20), the enzyme catalyzing the final step of L-tryptophan biosynthesis. The alpha subunit carries out the retro-aldol (aldol) cleavage of (1S,2R)-1-C-(indol-3-yl)glycerol 3-phosphate (indole-3-glycerol phosphate) to yield indole and D-glyceraldehyde 3-phosphate. The indole intermediate is channeled through an internal intersubunit tunnel to the beta subunit (TrpB), where it is condensed with L-serine in a pyridoxal 5'-phosphate-dependent reaction to form L-tryptophan. The functional enzyme is a tetramer of two alpha and two beta chains (alpha-beta-beta-alpha), and the alpha and beta subunits mutually allosterically activate one another, with the alpha subunit having very low catalytic activity in isolation. In Pseudomonas putida KT2440, trpA (PP_0082) lies in a trpBA operon and is required for tryptophan prototrophy; disruption produces a tryptophan auxotroph. The protein is a soluble, cytosolic enzyme of the aromatic amino acid biosynthetic pathway, adopting a TIM-barrel (ribulose-phosphate-binding barrel) fold.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000162 L-tryptophan biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: Tryptophan synthase alpha subunit catalyzes the final (step 5/5) of L-tryptophan biosynthesis from chorismate. This biological process annotation is strongly supported by the conserved enzymology and by P. putida KT2440 genetics, where trpA disruption produces a tryptophan auxotroph.
Reason: Core biological process of the gene; supported by experimental auxotrophy data (Molina-Henares et al. 2009) and UniPathway/UniProt pathway assignment.
GO:0004834 tryptophan synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: TrpA enables the alpha reaction of tryptophan synthase, the aldol cleavage of indole-3-glycerol phosphate to indole and glyceraldehyde 3-phosphate (EC 4.2.1.20, RHEA:10532). This is the canonical, conserved molecular function captured by HAMAP rule MF_00131 and InterPro family signatures.
Reason: Core molecular function, well supported by family/domain assignment (TrpA family, Pfam PF00290, TIGR00262) and EC/RHEA mapping. GO:0004834 is the standard term applied to both subunits of tryptophan synthase.
GO:0005829 cytosol
IEA
GO_REF:0000118
ACCEPT
Summary: Tryptophan synthase is a soluble cytosolic enzyme complex; cytosolic localization is the expected compartment for this amino acid biosynthetic enzyme in bacteria and is consistent with the lack of any signal/membrane features in the sequence.
Reason: Consistent with the soluble nature of the tryptophan synthase complex and the cytosolic localization of aromatic amino acid biosynthesis. Phylogeny-based (TreeGrafter) inference is reasonable for this conserved cytosolic enzyme.

Core Functions

Catalyzes the alpha reaction of tryptophan synthase, the aldol cleavage of indole-3-glycerol phosphate to indole and D-glyceraldehyde 3-phosphate, as the final step of L-tryptophan biosynthesis.

Molecular Function:
tryptophan synthase activity
Cellular Locations:
Supporting Evidence:

References

TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440.
  • A mini-Tn5 insertion near the start of PP_0082 (trpA) produces a tryptophan auxotroph, demonstrating trpA is required for L-tryptophan biosynthesis; trpA forms a trpBA operon with trpB.

Suggested Questions for Experts

Q: Is the indole intermediate fully channeled to TrpB in P. putida KT2440, or can free indole accumulate under any physiological conditions?

Suggested Experiments

Experiment: Complementation of the trpA auxotroph with wild-type and active-site mutant alleles to confirm catalytic residues (e.g., the conserved proton-acceptor residues) in the P. putida enzyme.

Deep Research

Asta

(trpA-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-05T20:15:39.197172

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.726)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.717)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[3] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.712)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[4] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.712)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.632)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[5] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.710)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[6] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.696)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[7] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.685)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[8] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.676)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[9] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.673)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[10] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[11] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.658)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[12] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[13] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.652)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[14] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.642)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[15] CodingQuarry: highly accurate hidden Markov model gene prediction in fungal genomes using RNA-seq transcripts

  • Authors: Alison C. Testa, James K. Hane, S. Ellwood, R. Oliver
  • Year: 2015
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/06704917bf44cd62eef0bb5cb944b97484056e21
  • DOI: 10.1186/s12864-015-1344-4
  • PMID: 25887563
  • PMCID: 4363200
  • Citations: 164
  • Influential citations: 20
  • Summary: CodingQuarry is a highly accurate, self-training GHMM fungal gene predictor designed to work with assembled, aligned RNA-seq transcripts, which capitalises on the high quality of fungal transcript assemblies by incorporating predictions made directly from transcript sequences.
  • Evidence snippets:
  • Snippet 1 (score: 0.638)
    > Whole-genome sequencing has enabled investigations into the gene content of living many organisms and forms the foundation for further study of gene expression, proteomics and epigenetics. After assembly of a novel genome, gene annotation is often the first step in analysing the gene content of an organism. Accurate annotation of the exonic structure of genes is crucial to the success of all subsequent functional and comparative analyses.
    > Problems that can potentially be caused by incorrect gene annotation are numerous and can lead to incorrect assessments of the lifestyle and ecology of an organism. In comparative genomics where orthologous genes or conserved functional domains are compared between species/isolates, the estimated numbers of such genes/ domains can be distorted by less than perfect annotations (as described by Hane et al. [1], S Text 1). Prediction of extracellular secretion, which can be determined by a short signal peptide at the N-terminus, can miss secreted proteins if the start codon of a gene has been incorrectly annotated. Mis-annotating the start of protein translation could either cut off the signal peptide or bury it within the annotation. While a seemingly benign annotation error, the consequences for downstream research could be detrimental, particularly as the biotic interactions or industrial applications of microbes are largely determined by their secretomes. Additionally, translated protein sequences of novel species are often submitted to databases such as NCBI [2] and Uniprot [3]. It is commonplace to use these database entries to support the annotation of related species or isolates, meaning errors present in the pioneer annotation may be repeated. When these new annotations based on false assumptions are added to databases, there is not only a propagation of errors, but also a perceived strengthening of homology evidence for incorrect protein sequences.
    > In recent years, correction of in silico predicted gene annotations with RNA-seq derived transcripts and read alignments has enabled vastly improved genome annotations and corrections of annotated gene structures [4][5][6].

[16] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.637)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.
  • Authors: Jaap van der Heijden, Asanda Mazubane, Marko Sallisalmi, E. Vorontsov, J. Tenhunen et al.
  • Year: 2025
  • Venue: Clinical Proteomics
  • URL: https://www.semanticscholar.org/paper/ae917565de913693a2713b46c3eda672d9d01c7f
  • DOI: 10.1186/s12014-025-09556-2
  • PMID: 40885913
  • PMCID: 12398169
  • Summary: The altered proteomic profile of hyaluronan-related proteins as reflected by the GO terms indicates a complex dysregulation not only in hyaluronan metabolism and extracellular matrix, but also in the regulation of several proteolytic enzymes.
  • Evidence snippets:
  • Snippet 1 (score: 0.631)
    > The identification of hyaluronan-associated genes was performed using Python 3.10.12. The UniProt REST Web Application Programming Interface (API) was used to query and retrieve gene annotations that met specific criteria (UniProt Consortium). A keyword-based query, utilizing the terms "hyaluronan, " "hyaluronic acid, " "hyaluronidase, " "hyaluronic acid synthase, " "hyaluronate binding protein, " "hyaluronan oligosaccharides, " "hyaluronan synthase, " and "hyaluronan receptor, " was executed across our dataset of 663 genes. Gene symbols were returned as "hits" when the query keywords were found in the annotations, including functional comments, Gene Ontology (GO) terms, and cross-referenced databases, formatted in JSON.

[18] Characterization of holins, the membrane proteins of coliphage ASEC2201: a genomewide in silico approach

  • Authors: Humaira Saeed, Sudhaker Padmesh, Aditi Singh, S. Singh, Mohammed Haris Siddiqui et al.
  • Year: 2025
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/a39392e12bf3bda67bdfe600053e8403deb3b887
  • DOI: 10.3389/fmicb.2025.1550594
  • PMID: 40703241
  • PMCID: 12283622
  • Citations: 3
  • Summary: In silico identification of cell-penetrating peptide motifs within the holin sequences suggests potential for enhanced intracellular delivery in CPP-fusion therapeutic constructs and demonstrates the potential of integrative in silico approaches in developing a comprehensive foundation for future experimental validation for proteins with no prior functional annotation.
  • Evidence snippets:
  • Snippet 1 (score: 0.630)
    > Protein-coding gene annotation is typically a two-step process. Initially, Prodigal is employed to identify open reading frames (ORFs) by locating gene coordinates, but it does not infer gene function. To assign putative functions, Prokka performs hierarchical annotation by comparing candidate genes to curated protein databases. It begins with a user-supplied, high-confidence protein set, using BLAST+ for sequence similarity searches. If no match is found, it progresses to UniProt's verified bacterial proteins, covering \~ 16,000 sequences, and then optionally to RefSeq proteins specific to the organism's genuscapturing nomenclature consistency. When sequence-based annotation fails, Prokka applies profile-based searches using HMMER's hmmscan to query against Pfam and TIGRFAMs databases. An e-value threshold of 10 −6 is consistently applied to ensure significance. If no reliable match is found across all levels, the gene is designated as a "hypothetical protein. " This layered strategy maximizes annotation accuracy and functional insight across diverse bacterial genomes (Seemann, 2014).
    > The genome of coliphage ASEC2201 has been analyzed and three holin protein coding genes were selected. The sequences of all three holin proteins were retrieved from NCBI using accession no. SRX17770782 in the FASTA format. The sequence similarity search was performed via BLAST against the non-redundant database (Altschul et al., 1990).

[19] Integrated genomic-transcriptomic analysis of clavulanic acid production in differentially productive Streptomyces clavuligerus strains

  • Authors: J. Gong, Jeong Sang Yi, Seungchan An, Hang Su Cho, Chang Hun Shin et al.
  • Year: 2025
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/b4903d3729bba93d1d47e38f3353a26f3530a8dd
  • DOI: 10.1038/s41598-025-29509-x
  • PMID: 41310174
  • PMCID: 12749726
  • Citations: 1
  • Summary: Findings include large plasmid deletions, an enrichment of mutations in secondary metabolite biosynthesis and regulatory genes, and metabolic shifts redirecting amino acid and carbon flux toward CA biosynthetic pathways.
  • Evidence snippets:
  • Snippet 1 (score: 0.627)
    > Gene annotation was primarily derived from the S. clavuligerus reference genome in the NCBI database and was annotated using the NCBI Prokaryotic Genome Annotation Pipeline 59 . However, several CA biosynthetic genes were manually corrected based on published literature 9 . For instance, two loci were originally annotated as clavaminate synthase 1 (cas1), but one of these loci is located near the cephamycin C biosynthetic cluster, indicating it was actually cas2. Following this correction, the RefSeq accession numbers of all genes in the reference genome were cross-referenced with the UniProt database to obtain additional annotations 60 . For the mutated genes identified through ICA, protein existence levels were manually assigned based on the UniProt data, including protein existence status, annotation score, similar proteins, and relevant publications.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  3. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  4. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  5. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  6. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  7. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  8. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  9. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  10. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  11. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  12. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  13. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  14. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  15. Alison C. Testa, James K. Hane, S. Ellwood, R. Oliver (2015). CodingQuarry: highly accurate hidden Markov model gene prediction in fungal genomes using RNA-seq transcripts. BMC Genomics. https://www.semanticscholar.org/paper/06704917bf44cd62eef0bb5cb944b97484056e21
  16. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  17. Jaap van der Heijden, Asanda Mazubane, Marko Sallisalmi, E. Vorontsov, J. Tenhunen et al. (2025). Plasma proteomics in septic shock and alcohol-related pancreatitis: a hyaluronan-centered approach. Clinical Proteomics. https://www.semanticscholar.org/paper/ae917565de913693a2713b46c3eda672d9d01c7f
  18. Humaira Saeed, Sudhaker Padmesh, Aditi Singh, S. Singh, Mohammed Haris Siddiqui et al. (2025). Characterization of holins, the membrane proteins of coliphage ASEC2201: a genomewide in silico approach. Frontiers in Microbiology. https://www.semanticscholar.org/paper/a39392e12bf3bda67bdfe600053e8403deb3b887
  19. J. Gong, Jeong Sang Yi, Seungchan An, Hang Su Cho, Chang Hun Shin et al. (2025). Integrated genomic-transcriptomic analysis of clavulanic acid production in differentially productive Streptomyces clavuligerus strains. Scientific Reports. https://www.semanticscholar.org/paper/b4903d3729bba93d1d47e38f3353a26f3530a8dd

Falcon

(trpA-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 35 citations 2 artifacts 2026-06-11T21:13:03.317837

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Comprehensive research report: trpA (UniProt Q88RP7; PP_0082) — Tryptophan synthase α chain in Pseudomonas putida KT2440

0) Scope and identity verification (critical disambiguation)

The target protein UniProt Q88RP7 is annotated as tryptophan synthase α chain (TrpA; EC 4.2.1.20) from Pseudomonas putida strain KT2440, with ordered locus name PP_0082 (gene trpA). The KT2440 gene PP0082 was experimentally linked to the tryptophan synthase α-subunit by (i) sequence similarity (“high similarity to trpA” by BLAST) and (ii) functional genetics, where a mini‑Tn5 insertion near the start of PP0082 produced a tryptophan auxotroph, consistent with loss of TrpA function. (molinahenares2009functionalanalysisof pages 2-4)

1) Key concepts and current understanding

1.1 What TrpA is (definition)

TrpA is the α-subunit of tryptophan synthase, a canonical bacterial enzyme complex (often αββα) catalyzing the final steps of L‑tryptophan biosynthesis. TrpA performs the indole-generating step, producing indole that is subsequently used by the β‑subunit (TrpB) to make L‑tryptophan. (duran2024alteringactivesiteloop pages 1-2, lambert2026sequencebasedgenerativeai pages 3-5)

1.2 Primary biochemical function (reaction, substrates/products, EC)

TrpA catalyzes the retro‑aldol cleavage (lyase reaction; EC 4.2.1.20) of indole‑3‑glycerol phosphate (IGP; also written I3GP) to yield indole and D‑glyceraldehyde‑3‑phosphate (G3P). (duran2024alteringactivesiteloop pages 1-2)

Mechanistically, recent synthesis of experimental/computational work describes a “push–pull” general acid–base mechanism involving Asp61 and Glu50 (numbering as discussed in that work’s TrpA models) and emphasizes that TrpA catalysis depends on access to a closed, catalytically activated conformational state. (duran2024alteringactivesiteloop pages 1-2, duran2024alteringactivesiteloop pages 3-4)

1.3 Functional coupling to TrpB: substrate channeling and allostery

A defining feature of tryptophan synthase is substrate channeling: indole produced at TrpA is transported through an internal intersubunit tunnel to the TrpB active site. One recent synthesis/engineering paper describes a ~20–25 Å substrate tunnel that channels indole from TrpA to TrpB. (lambert2026sequencebasedgenerativeai pages 1-3, lambert2026sequencebasedgenerativeai pages 3-5)

TrpA and TrpB are also mutually allosterically activating: binding and catalytic events in one subunit influence the conformational ensemble and catalytic competence of the other. Contemporary mechanistic discussions emphasize that open (low-activity) vs closed (high-activity) state transitions regulate ligand binding, intermediate stabilization, and product release across the complex. (duran2024alteringactivesiteloop pages 1-2, lambert2026sequencebasedgenerativeai pages 3-5)

2) Gene context in P. putida KT2440: operon structure, regulation, and genetics

2.1 Operon organization

In P. putida KT2440, trpA (PP0082) is clustered with trpB (PP0083 in the cited KT2440 organization) and the two genes overlap by one nucleotide, forming a trpBA operon confirmed by RT‑PCR. The adjacent regulator trpI (PP0084) is divergently transcribed and is monocistronic. (molinahenares2009functionalanalysisof pages 2-4)

This arrangement—trpBA divergently transcribed from trpI—is consistent with a regulatory module in which TrpI controls expression of the tryptophan synthase subunits. (molinahenares2009functionalanalysisof pages 2-4)

2.2 Evidence that trpA is required for tryptophan prototrophy

Disrupting PP0082/trpA yields tryptophan auxotrophy. A mini‑Tn5 insertion at the 7th codon of PP0082 produced a Trp‑requiring mutant (“Aux‑1”) in KT2440. (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2)

A separate phenotypic characterization reported that a trpA mutant grew on minimal medium only when L‑tryptophan (0.6 mM) was supplied, consistent with TrpA being essential for endogenous tryptophan biosynthesis under those conditions. (molinahenares2009functionalanalysisof pages 4-6)

2.3 TrpI-regulated indole responsiveness in KT2440 (quantitative)

A KT2440-derived transcriptional system PpTrpI/PPP_RS00425 (from the trpIAB locus, with locus tags reported as trpI PP_RS00430, trpA PP_RS00420, trpB PP_RS00425) was characterized as indole-inducible and used as a portable whole-cell biosensor module in heterologous hosts. (matulis2022developmentandcharacterization pages 2-4)

Quantitatively, this TrpI-based system achieved up to 639.6‑fold induction and showed a linear response over approximately 0.4–5 mM indole (with fitted inducer Km values in the ~0.9–1.8 mM range depending on host/medium). (matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6)

3) Cellular localization of TrpA (what is known vs inferred)

No direct subcellular localization experiments (e.g., fluorescent tagging, fractionation) were present in the retrieved full-text excerpts. However, the function of TrpA as a core enzyme in amino-acid biosynthesis and its role as a soluble subunit in a cytosolic enzyme complex strongly supports cytosolic localization in bacteria. This inference is consistent with its described participation in the soluble tryptophan synthase complex and with the cytosolic nature of indole channeling between TrpA and TrpB. (duran2024alteringactivesiteloop pages 1-2, lambert2026sequencebasedgenerativeai pages 3-5)

4) Recent developments (prioritizing 2023–2024): dynamics, engineering, and new functional perspectives

4.1 2024: loop dynamics as a design handle for TrpA standalone activity

A 2024 ACS Catalysis study dissected how two active-site loops (loop 2 and loop 6) coordinate transitions between open and closed conformations to enable catalysis and product release. It identifies formation of a catalytically activated enzyme–substrate state as rate-limiting and reports structural/distance metrics (e.g., positioning of a catalytic Asp relative to indole N1) that mark catalytically competent states. (duran2024alteringactivesiteloop pages 3-4)

A central quantitative result is the dramatic dependence of TrpA activity on TrpB: ZmTrpA has extremely low standalone activity (kcat 0.005 s−1; KM 1530 μM; kcat/KM 3.3 M−1 s−1) but in complex with TrpB reaches kcat 2.9 s−1; KM 195 μM; kcat/KM 15,006 M−1 s−1, corresponding to ~4515‑fold activation. (duran2024alteringactivesiteloop pages 3-4)

The authors further report a computationally designed TrpA variant with 163‑fold improved catalytic efficiency for IGP cleavage, illustrating how conformational ensemble engineering can partially decouple TrpA from obligatory TrpB activation. (duran2024alteringactivesiteloop pages 1-2)

Evidence from the paper’s Table 1 and conformational analysis figure supports these quantitative comparisons. (duran2024alteringactivesiteloop media 8730a16f, duran2024alteringactivesiteloop media 001701b4)

4.2 2024: expanding the industrial and synthetic-biology relevance of TrpA/TrpB chemistry

A 2024 review of indole biotechnology highlights that some bacterial TrpA homologs can function as bona fide indole‑3‑glycerol phosphate lyases capable of producing indole without the canonical TrpB partner under certain engineering contexts, and that engineered pathways plus in situ product removal can overcome indole toxicity to reach multi‑g/L titers. (ferrer2024indolesandthe pages 7-9)

Complementing this, 2024 work in biocatalysis shows how the tryptophan synthase system—especially TrpB—has become a platform for creating noncanonical amino acids and even reprogramming chemistry toward tyrosine analog synthesis, underscoring the modern view that the TrpA/TrpB scaffold is a tunable biocatalyst chassis rather than only a “housekeeping” biosynthetic enzyme. (almhjell2024theβsubunitof pages 1-2)

5) Current applications and real-world implementations

5.1 KT2440 aromatic pathway engineering connected to the tryptophan branch

Metabolic engineering in P. putida KT2440 has leveraged the tryptophan pathway branchpoint to produce aromatic chemicals such as anthranilate (o‑aminobenzoate; oAB), a tryptophan precursor. A 2015 study engineered KT2440 by deleting trpDC (blocking conversion of anthranilate onward toward tryptophan) and overexpressing a feedback‑insensitive DAHP synthase and an engineered anthranilate synthase, reaching 1.54 ± 0.3 g/L (11.23 mM) anthranilate from glucose in tryptophan-limited fed‑batch fermentation (with reported yields ~3.5–3.6% g/g under tested feed regimes). (kuepper2015metabolicengineeringof pages 1-2, kuepper2015metabolicengineeringof pages 6-7)

Although this is not a direct “TrpA product” application, it is a real KT2440 implementation demonstrating that the trp network (including downstream steps such as TrpA/TrpB) is an actionable control point in industrial strain design. (kuepper2015metabolicengineeringof pages 1-2)

5.2 Indole and indigoid biomanufacturing (2024 synthesis)

Recent synthesis of industrial biotechnology reports multiple routes to indole and indole-derived products:
- De novo indole production in engineered microbes reached ~0.7 g/L, and with extraction/sequestration strategies could reach 1.4 g/L or 5.7 g/L. (ferrer2024indolesandthe pages 7-9)
- Indigo production has been demonstrated at scale (e.g., 911 mg/L in a 3000 L fermenter from L‑Trp feed in one process) and high-titer fed-batch processes (e.g., 18 g/L indigo reported in the review’s survey). (ferrer2024indolesandthe pages 7-9, ferrer2024indolesandthe pages 9-10)
- Related indole-derived products include indican (2.9 g/L) and indirubin (up to 233 mg/L in one cited context; and 56 mg/L in a de novo pathway with coproduced indigo). (ferrer2024indolesandthe pages 7-9, ferrer2024indolesandthe pages 9-10)

These applications connect directly to TrpA biology because IGP→indole chemistry (whether via TrpA or dedicated IGP lyases) controls flux into indole-derived value chains. (ferrer2024indolesandthe pages 7-9)

5.3 Directed evolution and screening technology (2024)

A 2024 ACS Catalysis study introduced a DNA aptamer-based L‑tryptophan sensor compatible with droplet microfluidics, enabling ultrahigh-throughput directed evolution of tryptophan synthase activity. Reported throughput reaches up to 10^7 experiments/day, and a proof-of-principle screen of ~100,000 variants recovered enzymes with ~5-fold improved catalytic efficiency. (scheele2024ultrahighthroughputevolution pages 1-2, scheele2024ultrahighthroughputevolution pages 2-4)

6) Expert analysis: what the recent literature implies for functional annotation of P. putida TrpA (Q88RP7)

6.1 Most defensible primary function annotation

Given the KT2440 genetics (trpA disruption → Trp auxotrophy) and the conserved, well-defined enzymology of bacterial TrpA, the most defensible functional annotation for Q88RP7 is:
- Enzyme: tryptophan synthase α subunit (EC 4.2.1.20)
- Reaction: IGP → indole + G3P
- Pathway role: provides indole for the TrpB PLP-dependent condensation with L‑serine to form L‑tryptophan
- Complex behavior: functions as part of tryptophan synthase complex with strong allosteric coupling and indole channeling

This is supported by mechanistic synthesis describing TrpA’s reaction and TrpA–TrpB tunnel/allostery, and by P. putida KT2440 genetics/operon architecture establishing trpA as an essential tryptophan biosynthetic gene. (duran2024alteringactivesiteloop pages 1-2, lambert2026sequencebasedgenerativeai pages 3-5, molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6)

6.2 Substrate specificity considerations

The evidence base here primarily supports the canonical TrpA substrate IGP (I3GP) and products indole + G3P. The reviewed biotechnology literature indicates that homologous TSAs can sometimes act as IGP lyases in engineered contexts and that indole flux can be rerouted to diverse derivatives; however, these are generally homolog- and context-dependent properties and should not be assigned to KT2440 TrpA without direct experimental demonstration in P. putida KT2440. (ferrer2024indolesandthe pages 7-9)

6.3 Regulation in Pseudomonas context

KT2440’s TrpI-associated regulatory locus provides a plausible link between indole/IGP availability and trp gene expression. The strong, quantifiable indole inducibility of a KT2440-derived TrpI/promoter module suggests this regulator–promoter pair is a potent sensor of indole-related metabolites, enabling both natural regulation and synthetic-biology reuse. (matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6)

7) Key statistics and quantitative data (selected)

The table below consolidates the most actionable quantitative findings relevant to TrpA function, regulation, and applications.

Topic System/Organism Measurement (with units) Value(s) Notes/Context Source (citation id)
TrpA reaction Tryptophan synthase α-subunit (TrpA) Catalyzed reaction Indole-3-glycerol phosphate (IGP) → indole + D-glyceraldehyde-3-phosphate (G3P) Retro-aldol cleavage step of tryptophan synthase; indole is transferred to TrpB through the intersubunit tunnel (duran2024alteringactivesiteloop pages 1-2)
TrpA kinetics, standalone ZmTrpA alone kcat (s^-1); KM (μM); kcat/KM (M^-1 s^-1) 0.005 ± 0.001; 1530 ± 327; 3.3 ± 0.8 Very low standalone activity of α-subunit without TrpB partner (duran2024alteringactivesiteloop pages 3-4, duran2024alteringactivesiteloop media 8730a16f)
TrpA kinetics, activated complex ZmTrpA in complex with ZmTrpB kcat (s^-1); KM (μM); kcat/KM (M^-1 s^-1) 2.9 ± 0.10; 195 ± 17.9; 15,006 ± 1,430 TrpB strongly activates TrpA catalysis allosterically (duran2024alteringactivesiteloop pages 3-4, duran2024alteringactivesiteloop media 8730a16f)
TrpA activation by TrpB ZmTrpA + ZmTrpB Fold activation 4515-fold Reported increase in catalytic efficiency/activity upon complex formation (duran2024alteringactivesiteloop pages 3-4, duran2024alteringactivesiteloop media 8730a16f)
Engineered standalone TrpA improvement Designed ZmTrpA variant (ZmTrpASPM4-L6BX1) Catalytic-efficiency improvement (fold) 163-fold Loop-dynamics engineering enhanced standalone IGP cleavage (duran2024alteringactivesiteloop pages 1-2)
Indole-responsive regulation PpTrpI/PPP_RS00425 from Pseudomonas putida KT2440 Maximum induction (fold) Up to 639.6-fold Indole-inducible transcriptional system used to build whole-cell biosensors (matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6, matulis2022developmentandcharacterization pages 8-9)
Indole-responsive regulation PpTrpI/PPP_RS00425 from Pseudomonas putida KT2440 Linear response range (mM indole) ~0.4–5 mM Reported linear dose-response window for biosensor output (matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6)
Indole-responsive regulation E. coli host carrying PpTrpI/PPP_RS00425 Dynamic range (fold); Km (mM) 373.5-fold in LB, Km 1.207; 639.6-fold in minimal medium, Km 1.347 Host-dependent biosensor performance (matulis2022developmentandcharacterization pages 4-6)
Indole-responsive regulation Cupriavidus necator host carrying PpTrpI/PPP_RS00425 Dynamic range (fold); Km (mM) 101.4-fold in LB, Km 1.819; 11.9-fold in minimal medium, Km 0.9055 Growth inhibition observed at >=0.125 mM indole in C. necator (matulis2022developmentandcharacterization pages 4-6)
Indole production limitation Prior microbial indole production Titer (mM) ~5 mM Mentioned as an upper level in prior work, likely limited by indole toxicity (matulis2022developmentandcharacterization pages 1-2)
Anthranilate biomanufacturing Pseudomonas putida KT2440 engineered strain Maximum anthranilate titer 1.54 ± 0.3 g/L (11.23 mM) Best strain: ΔtrpDC with aroGD146N + trpES40FG overexpression under tryptophan-limited fed-batch conditions (kuepper2015metabolicengineeringof pages 1-2, kuepper2015metabolicengineeringof pages 5-6, kuepper2015metabolicengineeringof pages 6-7)
Anthranilate biomanufacturing Pseudomonas putida KT2440 engineered strain Shake-flask anthranilate titer 0.25 ± 0.004 g/L (1.83 mM) Initial production level before fed-batch optimization (kuepper2015metabolicengineeringof pages 5-6)
Anthranilate biomanufacturing Pseudomonas putida KT2440 engineered strain Alternative fed-batch anthranilate titer 1.0 ± 0.07 g/L Achieved with different glucose:tryptophan feed regime (kuepper2015metabolicengineeringof pages 5-6, kuepper2015metabolicengineeringof pages 6-7)
Anthranilate biomanufacturing Pseudomonas putida KT2440 engineered strain Product/substrate yield (g/g) 3.6 ± 0.5% and 3.5 ± 0.5% Reported for two fed conditions; yields were relatively similar (kuepper2015metabolicengineeringof pages 6-7)
Droplet evolution platform Directed evolution of TrpB in droplets Throughput (experiments/day) Up to 10^7/day Ultrahigh-throughput droplet microfluidic screening with aptamer readout (scheele2024ultrahighthroughputevolution pages 1-2)
Droplet evolution platform Directed evolution of TrpB in droplets Screened variants; improvement (fold) ~100,000 variants screened; ~5-fold improved variants recovered Demonstrated practical uHT enzyme evolution for tryptophan synthase (scheele2024ultrahighthroughputevolution pages 1-2)
Droplet evolution platform Directed evolution of TrpB in droplets Variants/day; sensor signal-to-noise ≈100,000 variants/day; ≈6-fold S/N at 5 mM Trp CS-10 aptamer sensor performed best and was compatible with droplet incubation (scheele2024ultrahighthroughputevolution pages 2-4)
De novo indole production Engineered Corynebacterium glutamicum with trpA/IGL route Indole titer (g/L) ~0.7 g/L Achieved with shikimate-producing background, trpB deletion, and in situ product removal (ferrer2024indolesandthe pages 7-9)
De novo indole production Engineered production with tributyrin extraction Indole titer (g/L) 1.4 g/L In situ removal improved de novo indole accumulation (ferrer2024indolesandthe pages 7-9)
De novo indole production Engineered production with dibutyl sebacate sequestration Indole titer (g/L) 5.7 g/L Higher final indole titer by mitigating toxicity/product loss (ferrer2024indolesandthe pages 7-9)
Industrial indigo production Biotransformation from 2 g/L L-Trp in 3000 L fermenter Indigo titer (mg/L) 911 mg/L Demonstrates scale-up of indigo bioproduction (ferrer2024indolesandthe pages 7-9)
Indigo production Engineered system with fused FMO–tryptophanase Indigo titer (g/L) 1.7 g/L Biotransformation route from L-tryptophan (ferrer2024indolesandthe pages 7-9)
Indigo production Engineered fed-batch process Indigo titer (g/L) 18 g/L High-titer indigoid production reported in review (ferrer2024indolesandthe pages 9-10)
Indican production Engineered production system Indican titer (g/L) 2.9 g/L Industrially relevant indole-derivative titer (ferrer2024indolesandthe pages 7-9)
Indirubin production Engineered production system Indirubin titer (mg/L) Up to 233 mg/L Reported under cysteine supplementation (ferrer2024indolesandthe pages 7-9)
De novo indirubin production Engineered pathway with coproduced indigo Indirubin titer (mg/L); Indigo coproduction (mg/L) 56 mg/L indirubin; 640 mg/L indigo Combined pathway engineering for indigoid products (ferrer2024indolesandthe pages 9-10)
Halogenated indole production Engineered Corynebacterium glutamicum / tryptophanase routes Final titer (mg/L) 16 mg/L 7-Cl-indole; 23 mg/L 7-Br-indole Demonstrates extension of Trp/indole biomanufacturing to halogenated derivatives (ferrer2024indolesandthe pages 9-10)

Table: This table compiles the main quantitative findings relevant to TrpA/trpA function, regulation, and tryptophan/indole biomanufacturing from the gathered evidence. It is useful as a quick reference for kinetics, regulatory response ranges, engineered production titers, and throughput metrics from recent literature.

Additionally, Duran et al. (2024) provide a direct tabular/figure comparison of TrpA kinetic constants and conformational effects of TrpB binding, useful as mechanistic evidence for strong α–β allosteric activation. (duran2024alteringactivesiteloop media 8730a16f, duran2024alteringactivesiteloop media 001701b4)

8) Limitations and evidence gaps

  • Direct subcellular localization evidence for P. putida KT2440 TrpA (e.g., microscopy/fractionation) was not identified in the retrieved sources; localization is inferred as cytosolic based on enzyme role and complex behavior. (duran2024alteringactivesiteloop pages 1-2, lambert2026sequencebasedgenerativeai pages 3-5)
  • The most detailed structure/dynamics data in the retrieved 2024 TrpA paper are from plant TrpA/BX1 homologs; nonetheless, the principles (loop‑gated catalysis; strong TrpB allosteric activation; open/closed ensemble shift) are widely treated as general features of tryptophan synthase systems and are appropriate for mechanistic context, but organism-specific kinetic constants for KT2440 TrpA were not found in the current corpus. (duran2024alteringactivesiteloop pages 3-4, duran2024alteringactivesiteloop pages 1-2)

9) References (with URLs and publication dates)

  • Molina-Henares MA et al. Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440. Microbial Biotechnology. Dec 2009. https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4)
  • Molina-Henares MA et al. Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium… Environmental Microbiology. Jun 2010. https://doi.org/10.1111/j.1462-2920.2010.02166.x (molina‐henares2010identificationofconditionally pages 6-7)
  • Kuepper J et al. Metabolic Engineering of Pseudomonas putida KT2440 to Produce Anthranilate from Glucose. Frontiers in Microbiology. Nov 2015. https://doi.org/10.3389/fmicb.2015.01310 (kuepper2015metabolicengineeringof pages 1-2)
  • Matulis P et al. Development and Characterization of Indole-Responsive Whole-Cell Biosensor… from P. putida KT2440. Int. J. Mol. Sci. Apr 2022. https://doi.org/10.3390/ijms23094649 (matulis2022developmentandcharacterization pages 1-2)
  • Scheele RA et al. Ultrahigh Throughput Evolution of Tryptophan Synthase in Droplets via an Aptamer Sensor. ACS Catalysis. Apr 2024. https://doi.org/10.1021/acscatal.4c00230 (scheele2024ultrahighthroughputevolution pages 1-2)
  • Almhjell PJ et al. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature Chemical Biology. May 2024. https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 1-2)
  • Ferrer L et al. Indoles and the advances in their biotechnological production for industrial applications. Systems Microbiology and Biomanufacturing. Dec 2024. https://doi.org/10.1007/s43393-023-00223-x (ferrer2024indolesandthe pages 7-9)
  • Duran C et al. Altering Active-Site Loop Dynamics Enhances Standalone Activity of the Tryptophan Synthase Alpha Subunit. ACS Catalysis. Nov 2024. https://doi.org/10.1021/acscatal.4c04587 (duran2024alteringactivesiteloop pages 1-2)

References

  1. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  2. (duran2024alteringactivesiteloop pages 1-2): Cristina Duran, Thomas Kinateder, Caroline Hiefinger, Reinhard Sterner, and Sílvia Osuna. Altering active-site loop dynamics enhances standalone activity of the tryptophan synthase alpha subunit. ACS Catalysis, 14:16986-16995, Nov 2024. URL: https://doi.org/10.1021/acscatal.4c04587, doi:10.1021/acscatal.4c04587. This article has 18 citations and is from a highest quality peer-reviewed journal.

  3. (lambert2026sequencebasedgenerativeai pages 3-5): Théophile Lambert, Amin Tavakoli, Gautham Dharuman, Jason Yang, Vignesh C. Bhethanabotla, Sukhvinder Kaur, Matthew Hill, Arvind Ramanathan, Anima Anandkumar, and Frances H. Arnold. Sequence-based generative ai design of versatile tryptophan synthases. Nature Communications, Jan 2026. URL: https://doi.org/10.1038/s41467-026-68384-6, doi:10.1038/s41467-026-68384-6. This article has 5 citations and is from a highest quality peer-reviewed journal.

  4. (duran2024alteringactivesiteloop pages 3-4): Cristina Duran, Thomas Kinateder, Caroline Hiefinger, Reinhard Sterner, and Sílvia Osuna. Altering active-site loop dynamics enhances standalone activity of the tryptophan synthase alpha subunit. ACS Catalysis, 14:16986-16995, Nov 2024. URL: https://doi.org/10.1021/acscatal.4c04587, doi:10.1021/acscatal.4c04587. This article has 18 citations and is from a highest quality peer-reviewed journal.

  5. (lambert2026sequencebasedgenerativeai pages 1-3): Théophile Lambert, Amin Tavakoli, Gautham Dharuman, Jason Yang, Vignesh C. Bhethanabotla, Sukhvinder Kaur, Matthew Hill, Arvind Ramanathan, Anima Anandkumar, and Frances H. Arnold. Sequence-based generative ai design of versatile tryptophan synthases. Nature Communications, Jan 2026. URL: https://doi.org/10.1038/s41467-026-68384-6, doi:10.1038/s41467-026-68384-6. This article has 5 citations and is from a highest quality peer-reviewed journal.

  6. (molinahenares2009functionalanalysisof pages 1-2): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  7. (molinahenares2009functionalanalysisof pages 4-6): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  8. (matulis2022developmentandcharacterization pages 2-4): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  9. (matulis2022developmentandcharacterization pages 1-2): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  10. (matulis2022developmentandcharacterization pages 4-6): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  11. (duran2024alteringactivesiteloop media 8730a16f): Cristina Duran, Thomas Kinateder, Caroline Hiefinger, Reinhard Sterner, and Sílvia Osuna. Altering active-site loop dynamics enhances standalone activity of the tryptophan synthase alpha subunit. ACS Catalysis, 14:16986-16995, Nov 2024. URL: https://doi.org/10.1021/acscatal.4c04587, doi:10.1021/acscatal.4c04587. This article has 18 citations and is from a highest quality peer-reviewed journal.

  12. (duran2024alteringactivesiteloop media 001701b4): Cristina Duran, Thomas Kinateder, Caroline Hiefinger, Reinhard Sterner, and Sílvia Osuna. Altering active-site loop dynamics enhances standalone activity of the tryptophan synthase alpha subunit. ACS Catalysis, 14:16986-16995, Nov 2024. URL: https://doi.org/10.1021/acscatal.4c04587, doi:10.1021/acscatal.4c04587. This article has 18 citations and is from a highest quality peer-reviewed journal.

  13. (ferrer2024indolesandthe pages 7-9): Lenny Ferrer, Melanie Mindt, Volker F. Wendisch, and Katarina Cankar. Indoles and the advances in their biotechnological production for industrial applications. Systems Microbiology and Biomanufacturing, 4:511-527, Dec 2024. URL: https://doi.org/10.1007/s43393-023-00223-x, doi:10.1007/s43393-023-00223-x. This article has 34 citations.

  14. (almhjell2024theβsubunitof pages 1-2): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  15. (kuepper2015metabolicengineeringof pages 1-2): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

  16. (kuepper2015metabolicengineeringof pages 6-7): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

  17. (ferrer2024indolesandthe pages 9-10): Lenny Ferrer, Melanie Mindt, Volker F. Wendisch, and Katarina Cankar. Indoles and the advances in their biotechnological production for industrial applications. Systems Microbiology and Biomanufacturing, 4:511-527, Dec 2024. URL: https://doi.org/10.1007/s43393-023-00223-x, doi:10.1007/s43393-023-00223-x. This article has 34 citations.

  18. (scheele2024ultrahighthroughputevolution pages 1-2): Remkes A. Scheele, Yanik Weber, Friederike E. H. Nintzel, Michael Herger, Tomasz S. Kaminski, and Florian Hollfelder. Ultrahigh throughput evolution of tryptophan synthase in droplets via an aptamer sensor. ACS Catalysis, 14:6259-6271, Apr 2024. URL: https://doi.org/10.1021/acscatal.4c00230, doi:10.1021/acscatal.4c00230. This article has 17 citations and is from a highest quality peer-reviewed journal.

  19. (scheele2024ultrahighthroughputevolution pages 2-4): Remkes A. Scheele, Yanik Weber, Friederike E. H. Nintzel, Michael Herger, Tomasz S. Kaminski, and Florian Hollfelder. Ultrahigh throughput evolution of tryptophan synthase in droplets via an aptamer sensor. ACS Catalysis, 14:6259-6271, Apr 2024. URL: https://doi.org/10.1021/acscatal.4c00230, doi:10.1021/acscatal.4c00230. This article has 17 citations and is from a highest quality peer-reviewed journal.

  20. (matulis2022developmentandcharacterization pages 8-9): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  21. (kuepper2015metabolicengineeringof pages 5-6): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

  22. (molina‐henares2010identificationofconditionally pages 6-7): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

Artifacts

Citations

  1. molinahenares2009functionalanalysisof pages 2-4
  2. duran2024alteringactivesiteloop pages 1-2
  3. molinahenares2009functionalanalysisof pages 4-6
  4. matulis2022developmentandcharacterization pages 2-4
  5. duran2024alteringactivesiteloop pages 3-4
  6. ferrer2024indolesandthe pages 7-9
  7. kuepper2015metabolicengineeringof pages 1-2
  8. matulis2022developmentandcharacterization pages 4-6
  9. matulis2022developmentandcharacterization pages 1-2
  10. kuepper2015metabolicengineeringof pages 5-6
  11. kuepper2015metabolicengineeringof pages 6-7
  12. scheele2024ultrahighthroughputevolution pages 1-2
  13. scheele2024ultrahighthroughputevolution pages 2-4
  14. ferrer2024indolesandthe pages 9-10
  15. lambert2026sequencebasedgenerativeai pages 3-5
  16. lambert2026sequencebasedgenerativeai pages 1-3
  17. molinahenares2009functionalanalysisof pages 1-2
  18. matulis2022developmentandcharacterization pages 8-9
  19. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  20. https://doi.org/10.1111/j.1462-2920.2010.02166.x
  21. https://doi.org/10.3389/fmicb.2015.01310
  22. https://doi.org/10.3390/ijms23094649
  23. https://doi.org/10.1021/acscatal.4c00230
  24. https://doi.org/10.1038/s41589-024-01619-z
  25. https://doi.org/10.1007/s43393-023-00223-x
  26. https://doi.org/10.1021/acscatal.4c04587
  27. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  28. https://doi.org/10.1021/acscatal.4c04587,
  29. https://doi.org/10.1038/s41467-026-68384-6,
  30. https://doi.org/10.3390/ijms23094649,
  31. https://doi.org/10.1007/s43393-023-00223-x,
  32. https://doi.org/10.1038/s41589-024-01619-z,
  33. https://doi.org/10.3389/fmicb.2015.01310,
  34. https://doi.org/10.1021/acscatal.4c00230,
  35. https://doi.org/10.1111/j.1462-2920.2010.02166.x,

📄 View Raw YAML

id: Q88RP7
gene_symbol: trpA
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: trpA encodes the alpha subunit of tryptophan synthase (EC 4.2.1.20), the enzyme catalyzing the final step of L-tryptophan biosynthesis. The alpha subunit carries out the retro-aldol (aldol) cleavage of (1S,2R)-1-C-(indol-3-yl)glycerol 3-phosphate (indole-3-glycerol phosphate) to yield indole and D-glyceraldehyde 3-phosphate. The indole intermediate is channeled through an internal intersubunit tunnel to the beta subunit (TrpB), where it is condensed with L-serine in a pyridoxal 5'-phosphate-dependent reaction to form L-tryptophan. The functional enzyme is a tetramer of two alpha and two beta chains (alpha-beta-beta-alpha), and the alpha and beta subunits mutually allosterically activate one another, with the alpha subunit having very low catalytic activity in isolation. In Pseudomonas putida KT2440, trpA (PP_0082) lies in a trpBA operon and is required for tryptophan prototrophy; disruption produces a tryptophan auxotroph. The protein is a soluble, cytosolic enzyme of the aromatic amino acid biosynthetic pathway, adopting a TIM-barrel (ribulose-phosphate-binding barrel) fold.
existing_annotations:
- term:
    id: GO:0000162
    label: L-tryptophan biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: Tryptophan synthase alpha subunit catalyzes the final (step 5/5) of L-tryptophan biosynthesis from chorismate. This biological process annotation is strongly supported by the conserved enzymology and by P. putida KT2440 genetics, where trpA disruption produces a tryptophan auxotroph.
    action: ACCEPT
    reason: Core biological process of the gene; supported by experimental auxotrophy data (Molina-Henares et al. 2009) and UniPathway/UniProt pathway assignment.
- term:
    id: GO:0004834
    label: tryptophan synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: TrpA enables the alpha reaction of tryptophan synthase, the aldol cleavage of indole-3-glycerol phosphate to indole and glyceraldehyde 3-phosphate (EC 4.2.1.20, RHEA:10532). This is the canonical, conserved molecular function captured by HAMAP rule MF_00131 and InterPro family signatures.
    action: ACCEPT
    reason: Core molecular function, well supported by family/domain assignment (TrpA family, Pfam PF00290, TIGR00262) and EC/RHEA mapping. GO:0004834 is the standard term applied to both subunits of tryptophan synthase.
- term:
    id: GO:0005829
    label: cytosol
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: Tryptophan synthase is a soluble cytosolic enzyme complex; cytosolic localization is the expected compartment for this amino acid biosynthetic enzyme in bacteria and is consistent with the lack of any signal/membrane features in the sequence.
    action: ACCEPT
    reason: Consistent with the soluble nature of the tryptophan synthase complex and the cytosolic localization of aromatic amino acid biosynthesis. Phylogeny-based (TreeGrafter) inference is reasonable for this conserved cytosolic enzyme.
core_functions:
- description: Catalyzes the alpha reaction of tryptophan synthase, the aldol cleavage of indole-3-glycerol phosphate to indole and D-glyceraldehyde 3-phosphate, as the final step of L-tryptophan biosynthesis.
  supported_by:
  - reference_id: PMID:21261884
  molecular_function:
    id: GO:0004834
    label: tryptophan synthase activity
  directly_involved_in:
  - id: GO:0000162
    label: L-tryptophan biosynthetic process
  locations:
  - id: GO:0005829
    label: cytosol
references:
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:21261884
  title: Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440.
  findings:
  - statement: A mini-Tn5 insertion near the start of PP_0082 (trpA) produces a tryptophan auxotroph, demonstrating trpA is required for L-tryptophan biosynthesis; trpA forms a trpBA operon with trpB.
    reference_section_type: RESULTS
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: PMID recovered from DOI 10.1111/j.1751-7915.2008.00062.x (Molina-Henares et al., Microbial Biotechnology 2009) and PubMed-verified; title matches the cached publication. The previously cited PMID:19302569 was an invalid/wrong identifier (could not be resolved on PubMed) and has been corrected. Supports trpA essentiality for tryptophan prototrophy and trpBA operon structure in KT2440.
suggested_questions:
- question: Is the indole intermediate fully channeled to TrpB in P. putida KT2440, or can free indole accumulate under any physiological conditions?
suggested_experiments:
- description: Complementation of the trpA auxotroph with wild-type and active-site mutant alleles to confirm catalytic residues (e.g., the conserved proton-acceptor residues) in the P. putida enzyme.