trpB

UniProt ID: Q88RP6
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

Tryptophan synthase beta chain (TrpB, EC 4.2.1.20), a pyridoxal 5'-phosphate (PLP)-dependent enzyme that catalyzes the final (beta) reaction of L-tryptophan biosynthesis, condensing indole with L-serine to yield L-tryptophan and water. PLP is bound as an internal aldimine to an active-site lysine (residue 95 in this protein). TrpB is a member of the fold-type II PLP enzyme family (TrpB family) and normally assembles with the alpha subunit (TrpA) into the alpha2-beta2 tryptophan synthase complex, in which the indole produced by TrpA from indole-3-glycerol phosphate is channeled directly to the TrpB active site through an intramolecular tunnel. TrpB carries out the terminal, fifth step of the conversion of chorismate to L-tryptophan and is a soluble cytoplasmic enzyme. In P. putida KT2440 the gene (PP_0083) is adjacent to and co-transcribed with trpA (PP_0082) as a trpBA operon, and its expression is strongly induced by indole.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000162 L-tryptophan biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: TrpB catalyzes the terminal step of L-tryptophan biosynthesis; this annotation correctly captures the core biological process.
Reason: The protein is a UniProt-reviewed tryptophan synthase beta chain (EC 4.2.1.20) belonging to the TrpB family, with UniPathway UPA00035 (L-tryptophan biosynthesis, step 5/5). In P. putida KT2440 trpA disruption causes tryptophan auxotrophy, confirming the trpBA cluster is required for de novo tryptophan synthesis (PMID:21261884; see also file:PSEPK/trpB/trpB-deep-research-falcon.md). The IEA assignment is well-supported and represents a core function.
GO:0004834 tryptophan synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: Correct molecular function. TrpB is the tryptophan synthase beta subunit catalyzing the PLP-dependent beta-reaction (indole + L-serine -> L-tryptophan + H2O).
Reason: Supported by EC 4.2.1.20, RHEA:10532, the conserved PLP-binding lysine (residue 95), HAMAP-Rule MF_00133, and TrpB-family InterPro/PANTHER signatures. This is the core enzymatic activity of the gene product.
GO:0005737 cytoplasm
IEA
GO_REF:0000118
ACCEPT
Summary: Bacterial tryptophan synthase is a soluble cytoplasmic enzyme; cytoplasmic localization is correct.
Reason: TrpB has no signal peptide or transmembrane region and functions in cytoplasmic amino-acid biosynthesis as part of the soluble alpha2-beta2 tryptophan synthase complex. The TreeGrafter IEA assignment is consistent with the well-established localization of this enzyme family. The term is somewhat generic but accurate for a bacterial cytosolic enzyme.

Core Functions

Catalyzes the PLP-dependent beta-replacement reaction forming L-tryptophan from indole (channeled from TrpA) and L-serine, completing the terminal step of L-tryptophan biosynthesis

Molecular Function:
tryptophan synthase activity
Cellular Locations:
Supporting Evidence:
  • GO_REF:0000120
    EC=4.2.1.20; tryptophan synthase activity inferred from InterPro, RHEA:10532, UniRule and PANTHER (TrpB family).
  • file:PSEPK/trpB/trpB-uniprot.txt
    FUNCTION: The beta subunit is responsible for the synthesis of L-tryptophan from indole and L-serine. CATALYTIC ACTIVITY: indol-3-yl glycerol 3-phosphate + L-serine = D-glyceraldehyde 3-phosphate + L-tryptophan + H2O; PATHWAY: L-tryptophan from chorismate, step 5/5; COFACTOR: pyridoxal 5'-phosphate.

References

TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
file:PSEPK/trpB/trpB-uniprot.txt
UniProt entry TRPB_PSEPK (Q88RP6)
  • TrpB synthesizes L-tryptophan from indole and L-serine; PLP cofactor; pathway L-tryptophan from chorismate step 5/5; functions as a tetramer of two alpha and two beta chains.
    "FUNCTION: The beta subunit is responsible for the synthesis of L-tryptophan from indole and L-serine. SUBUNIT: Tetramer of two alpha and two beta chains."
Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440
  • In P. putida KT2440 there is a single pathway from chorismate to tryptophan; the trp genes are in unlinked regions with trpBA organized as an operon, and auxotroph screening shows the pathway is required for de novo tryptophan biosynthesis.
    "Genes for tryptophan biosynthesis are grouped in unlinked regions with the trpBA and trpGDE genes organized as operons... There is a single pathway from chorismate leading to the biosynthesis of tryptophan."

Deep Research

Asta

(trpB-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-05T20:16:06.686172

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.716)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.623)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[3] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.700)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[4] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.698)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[5] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.685)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[6] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.683)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[7] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.677)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[8] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.675)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[9] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[10] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.663)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[11] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.645)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[12] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.640)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[13] Proteome-wide Subcellular Topologies of E. coli Polypeptides Database (STEPdb)*

  • Authors: Georgia Orfanoudaki, A. Economou
  • Year: 2014
  • Venue: Molecular & Cellular Proteomics
  • URL: https://www.semanticscholar.org/paper/8e452aca415b30d995ba9f977086924c8fed8f19
  • DOI: 10.1074/mcp.O114.041137
  • PMID: 25210196
  • Citations: 72
  • Influential citations: 3
  • Summary: The STEP database (STEPdb) is a comprehensive characterization of subcellular localization and topology of the complete proteome of Escherichia coli and is the first database that contains an extensive set of peripheral IM proteins (PIM proteins) and includes their graphical visualization into complexes, cellular functions, and interactions.
  • Evidence snippets:
  • Snippet 1 (score: 0.639)
    > The E. coli K-12 Reference Proteome and Data Sources-Two databases Uniprot (29) and EcoLOCATION (32) and the proposed IM proteome (33) were the main initial starting points for the complete subcellular categorization of K-12 described here. The E. coli K-12/ MG1655 strain is one of the microbial proteomes whose comprehensive annotation is of the highest priority in Uniprot (29). This is the "reference proteome" for E. coli, contains 4303 proteins, and has been annotated here. Our annotation has been formulated in such a way that it can be easily incorporated in Uniprot.
    > EchoLOCATION has an easily accessible table that maps gene names to subcellular locations. However, mapping the gene names given by EchoLOCATION to the respective protein identifiers in Uniprot was not straightforward. Unfortunately, gene names cannot serve as unique identifiers of a protein sequence. In more than 100 cases the gene name of a predicted protein in EchoLOCATION when searched against Uniprot gave as a result more than one K-12 protein hits. That is because there are proteins that have common synonymous gene names with the primary gene name of others.
    > To retrieve updated Uniprot accession identifiers and to map Uniprot accessions identifiers to EchoLOCATION identifiers (termed: EchoBASE IDs) we used the "ID mapping" function of Uniprot. In cases where the only provided identifiers were the gene names, we used mySQL queries to compare with the primary and alternative gene names in Uniprot. In cases where multiple matches existed for the same gene name, we manually resolved the differences based on other information (e.g. protein description, mass etc.).
    > The annotation of pseudogenes, mobile elements, transposons, and insertion elements relied on EcoGene (34), Uniprot (29), and Ochman et al. (35). The list of E. coli K-12 complexes was retrieved from EcoCyc (31) and literature searches.

[14] Undergraduate Bioinformatics Conceptualizing Form and Function on a Molecular Scale

  • Authors: P. Kramer, Jack Treml
  • Year: 2022
  • Venue: Midwestern Journal of Undergraduate Sciences
  • URL: https://www.semanticscholar.org/paper/58a6af68fc2fd768735d429724ffdd52e16dfb27
  • DOI: 10.17161/mjusc.v1i1.18565
  • Summary: The following is a walkthrough of a project designed to overcome the lack of sense for proteins as real objects.
  • Evidence snippets:
  • Snippet 1 (score: 0.634)
    > i. Click "See more" to view a bar chart containing data on where in the body's tissues the gene is expressed (as determined by RNA sequencing). Save and include this bar chart as the deliverable for this step.
    > II. Universal Protein Research Knowledgebase (UniProtKB) 8 6. UniProt Entry Number
    > i. Follow the UniProt link in the Resources then search for the protein using the NCBI Gene ID ii. Carefully select the result that best matches the gene and organism of interest by clicking on the blue entry number. iii. This page will be used later to gather further details about the protein.
    > III. RCSB Protein Data Bank (PDB) 9 7. RCSB PDB Solved Structure Identifi er i. Follow the RCSB PDB link in the Resources and search for the protein by either the common name or the NCBI Gene ID, making sure to select the organism of interest on the left. ii. You must ensure that your chosen protein has an existing solved structure in this data bank in order to do a mutational analysis in later parts of this exercise.
    > IV. NCBI GenBank 10 8. AA Protein Sequence i. From the NCBI Gene page, go to the "Genomic regions, transcripts, and products" section and then click "GenBank" on the right. Scroll down to the fi rst Coding Sequence "CDS" section and look directly after "/translation=" for the full protein sequence. ii. Sequence needs to be in FASTA Format consisting of '>' followed by a simple name, a return, and then the sequence in one continuous line of text. See "FASTA Formatting" link in Resources.

[15] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.632)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[16] GRIMM: Genetic stRatification for Inference in Molecular Modeling

  • Authors: Ashley Babjac, Adrienne Hoarfrost
  • Year: 2026
  • Venue: Unknown venue
  • URL: https://www.semanticscholar.org/paper/a923e3c2800f1ce4cc937015594f74395af624a9
  • Citations: 1
  • Summary: GRIMM (Genetic stRatification for Inference in Molecular Modeling), a benchmark for enzyme function prediction that employs genetic stratification, enables more realistic evaluation of functional prediction models on both familiar and unseen classes and establishes a benchmark that more faithfully assesses model performance and generalizability.
  • Evidence snippets:
  • Snippet 1 (score: 0.631)
    > We used amino acid sequence data from the Universal Protein Resource (UniProt) UniProt Consortium [2025] and associated gene DNA coding sequences in the European Nucleotide Archive (ENA) Leinonen et al. [2010]. The UniProt database is divided into two sections: (i) UniProt/TrEMBL (which includes more sequence diversity but more annotation errors owing to homology-based annotations and error propagation Schnoes et al. [2009]) and (ii) UniProt/SwissProt (which is carefully curated and provides high confidence accurate functional annotations). For this study, we limited the data retrieved to prokaryotic organisms in SwissProt (May 2025) UniProt Consortium [2025].
    > We mapped UniProt/SwissProt accession numbers to corresponding identifiers in UniRef (Universal Protein Resource Reference Clusters) UniRef50, UniRef90, and UniRef100, which define protein clusters based on amino acid sequence identity at 50, 90, and 100 percent amino acid identity respectively; and to their corresponding EMBL CDS IDs for gene coding sequences in the ENA database Leinonen et al. [2010] using ID mapping files uni [2021], Wang et al. [2021] from UniProtKB. Each EMBL CDS ID's corresponding DNA sequence was then obtained from ENA. EC numbers that were either incomplete or missing were removed from the dataset. Individual entries were created for UniProt records with more than one EMBL CDS ID in the nucleotide version of the dataset.

[17] Protocol for gene annotation, prediction, and validation of genomic gene expansion

  • Authors: Quanwei Zhang, Zhengdong D. Zhang
  • Year: 2022
  • Venue: STAR Protocols
  • URL: https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3
  • DOI: 10.1016/j.xpro.2022.101692
  • PMID: 36125934
  • PMCID: 9494284
  • Citations: 1
  • Summary: A detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation is described and steps to discover functionality of each copy of replicated genes are detailed.
  • Evidence snippets:
  • Snippet 1 (score: 0.629)
    > 3. Gene annotation and functional annotation. a. Gene structure annotation.
    > In addition to gene prediction models, evidence from orthologous protein sequences and transcriptome assembly could be used to improve annotation quality. Protein sequences of orthologous genes can be obtained from UniProt (The UniProt, 2017). Ones from Swiss-Port have been reviewed and thus are of higher quality. Transcriptome assembly may be available from previous studies or can be assembled de novo from RNA-seq reads by Trinity (Haas et al., 2013). High quality transcriptome assembly can be selected as described in (Zhang et al., 2021). Note: Details about gene structure annotation (Holt and Yandell, 2011) can be found at http:// gmod.org/wiki/MAKER_Tutorial, https://darencard.net/blog/2017-05-16-maker-genomeannotation/, and the protocol (Campbell et al., 2014).
    > b. Quality measurement and functional annotation.
    > For each predicted gene, Maker2 provides the annotation edit distance (AED) score, which measures the goodness of fit between its predicted gene structure and its evidence support. The lower the score, the more accurate the prediction. If more than 90% genes with AED scores lower than 0.5, the genome can be considered well annotated. In addition to the AED score, a high proportion of recognizable domains contained in predicted protein -e.g., higher than 50% -also indicates a good annotation. Recognizable protein domains can by scanned by InterProScan (Jones et al., 2014), assigning potential function to predicted genes.
    > Note: Besides the aforementioned quality measurement, we strongly recommend measuring the completeness of the genome assembly and annotation by checking the existence of a set of Benchmarking Universal Single-Copy Orthologs (BUSCO) (Simao et al., 2015). A high-level completeness of genome assembly and annotation is imperative for a better identification of gene expansion. Based on the result of this analysis, researchers can decide whether they need to further improve the genome assembly before predicting gene expansion. A detailed protocol of BUSCO is available at

[18] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.626)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[19] GOnet: a tool for interactive Gene Ontology analysis

  • Authors: M. Pomaznoy, Brendan Ha, Bjoern Peters
  • Year: 2018
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  • DOI: 10.1186/s12859-018-2533-3
  • PMID: 30526489
  • PMCID: 6286514
  • Citations: 247
  • Influential citations: 17
  • Summary: The open-source GOnet web-application is created, which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis and provides insight into the functional interconnection of the submitted entries.
  • Evidence snippets:
  • Snippet 1 (score: 0.625)
    > In a basic workflow, the GOnet application receives a list of gene symbols, protein symbols, or protein IDs (UniProt IDs) as an input, and outputs a graph (an example given in Fig. 1). There are various input parameters which will affect the actual structure of the graph visualized and its appearance. The first main user choice is which GO terms the genes are annotated against:
    > 1. GO terms statistically significantly over-represented in the gene list submitted. 2. A predefined subset (also known as 'GO slim'), or a user-supplied list of terms.
    > In the first case the analysis will be referred to as an 'enrichment' analysis, in the second as an 'annotation' analysis.
    > Input parameters 1) Gene list. A mandatory input parameter containing the genes/proteins of interest. Currently human and mouse data is supported. An example of a human gene list might look like this:
    > Fig. 1 Sample network output generated by GOnet application. Gene differentially expressed in CD4 Bulk Memory T cells in Latent TB patients compared to healthy controls were used as an example [22] The gene list can also be accompanied with a contrast value. For example, This contrast value can be any decimal number, such as the log-fold change of gene expression between two conditions. This is merely a visualization enhancement. If the value is supplied it can be used later to differentially color specific genes in the graph (note different colors of gene nodes in Fig. 1), and visually indicate up-or down-regulation of specific genes and gene clusters.
    > The application can process common gene symbols (like in the example above), UniProt IDs, and MGI Accession IDs (mouse only). The former type of ID (gene symbols), although is the most human friendly, can unfortunately be ambiguous. For example, AIM1 can mean 'absent in melanoma' (also called CRYBG1) or 'Aurora and Ipl1-like midbody-associated protein' (also known as AURKB). Due to this ambiguity UniProt IDs or MGI accession IDs (for mouse) are preferred.
    > 2) GO namespace. Can be any of 'biological process', 'molecular function' or 'cellular component'.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  3. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  4. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  5. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  6. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  7. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  8. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  9. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  10. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  11. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  12. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  13. Georgia Orfanoudaki, A. Economou (2014). Proteome-wide Subcellular Topologies of E. coli Polypeptides Database (STEPdb)*. Molecular & Cellular Proteomics. https://www.semanticscholar.org/paper/8e452aca415b30d995ba9f977086924c8fed8f19
  14. P. Kramer, Jack Treml (2022). Undergraduate Bioinformatics Conceptualizing Form and Function on a Molecular Scale. Midwestern Journal of Undergraduate Sciences. https://www.semanticscholar.org/paper/58a6af68fc2fd768735d429724ffdd52e16dfb27
  15. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  16. Ashley Babjac, Adrienne Hoarfrost (2026). GRIMM: Genetic stRatification for Inference in Molecular Modeling. https://www.semanticscholar.org/paper/a923e3c2800f1ce4cc937015594f74395af624a9
  17. Quanwei Zhang, Zhengdong D. Zhang (2022). Protocol for gene annotation, prediction, and validation of genomic gene expansion. STAR Protocols. https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3
  18. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  19. M. Pomaznoy, Brendan Ha, Bjoern Peters (2018). GOnet: a tool for interactive Gene Ontology analysis. BMC Bioinformatics. https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9

Falcon

(trpB-deep-research-falcon.md)
Comprehensive Research Report: trpB Gene (UniProt Q88RP6) in *Pseudomonas putida* KT2440 Falcon Edison Scientific Literature 34 citations 2 artifacts 2026-06-11T21:37:30.759039

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Comprehensive Research Report: trpB Gene (UniProt Q88RP6) in Pseudomonas putida KT2440

Executive Summary

The gene trpB (locus tag PP_0083; UniProt accession Q88RP6) in Pseudomonas putida strain KT2440 encodes the tryptophan synthase beta chain (EC 4.2.1.20), a pyridoxal 5′-phosphate (PLP)-dependent enzyme catalyzing the terminal step of L-tryptophan biosynthesis (kim2013indoletoxicityinvolves pages 3-5). This report provides a detailed functional annotation based on organism-specific experimental evidence and conserved mechanistic knowledge, with emphasis on recent (2023–2024) advances in TrpB enzymology and applications.


1. Gene Identity Verification and Organism-Specific Context

1.1 Confirmed Identity

The gene symbol trpB corresponds unambiguously to PP_0083 in P. putida KT2440, as verified by microarray and functional genomics studies (kim2013indoletoxicityinvolves pages 3-5). This locus encodes the tryptophan synthase beta subunit, matching the UniProt annotation for Q88RP6 (RecName: Tryptophan synthase beta chain; EC 4.2.1.20) (kim2013indoletoxicityinvolves pages 3-5).

Source: Kim et al., 2013, FEMS Microbiology Letters 343(1):89-99. doi:10.1111/1574-6968.12135. URL: https://doi.org/10.1111/1574-6968.12135 (kim2013indoletoxicityinvolves pages 3-5)

1.2 Operon Organization and Genomic Context

In P. putida KT2440, the tryptophan biosynthesis genes are organized into unlinked genomic regions, contrasting with the single-operon architecture in E. coli (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2). Specifically:

  • trpA (PP_0082) and trpB (PP_0083) are adjacent and overlap by one nucleotide, indicating tight co-transcription as a trpBA operon (molinahenares2009functionalanalysisof pages 2-4).
  • RT-PCR experiments using primers spanning the trpA/trpB junction confirmed in vivo co-transcription of this operon (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8).
  • Other tryptophan pathway genes occur as separate operons or monocistronic units: trpGDC forms an operon, while trpE and trpF are monocistronic (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2).
  • The trpI gene (encoding a LysR-family regulator) is divergently oriented from the trpBA cluster (molinahenares2009functionalanalysisof pages 2-4, matulis2022developmentandcharacterization pages 2-4).

Sources:
- Molina-Henares et al., 2009, Microbial Biotechnology 2:91-100. doi:10.1111/j.1751-7915.2008.00062.x. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8, molinahenares2009functionalanalysisof pages 1-2)

1.3 Pathway Role and Genetic Evidence

Genetic disruption experiments demonstrate the essentiality of trpBA for de novo tryptophan biosynthesis in P. putida KT2440:

  • A mini-Tn5 insertion at the 7th codon of trpA (Aux-1 mutant) produces a tryptophan auxotroph, which grows only when supplemented with exogenous tryptophan (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6).
  • This confirms that the trpA/trpB cluster is required for L-tryptophan synthesis from chorismate (molinahenares2009functionalanalysisof pages 1-2).

Source: Molina-Henares et al., 2009 (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2)

1.4 Regulation: Indole-Responsive Expression

Transcriptomic and biosensor studies reveal indole-responsive regulation of trpB in P. putida KT2440:

  • Microarray analysis identified trpB (PP_0083) as the most highly upregulated gene (3.52-fold) in response to indole treatment, out of 47 genes showing >1.5-fold change (kim2013indoletoxicityinvolves pages 3-5, kim2013indoletoxicityinvolves pages 8-9).
  • The authors propose that incorporation of indole into tryptophan biosynthesis may represent a detoxification or stress-mitigation mechanism (kim2013indoletoxicityinvolves pages 8-9).
  • A TrpI-dependent, indole-inducible promoter system (PpTrpI/PPP_RS00425) derived from KT2440 was characterized: it shows up to 639.6-fold induction by indole (linear response ~0.4–5 mM) and requires the trpI regulator for activation (matulis2022developmentandcharacterization pages 2-4).
  • This system functions as a whole-cell indole biosensor in heterologous hosts (E. coli, Cupriavidus necator), demonstrating the conserved regulatory logic (matulis2022developmentandcharacterization pages 2-4).

Sources:
- Kim et al., 2013, FEMS Microbiology Letters (kim2013indoletoxicityinvolves pages 3-5, kim2013indoletoxicityinvolves pages 8-9)
- Matulis et al., 2022, International Journal of Molecular Sciences 23:4649. doi:10.3390/ijms23094649. URL: https://doi.org/10.3390/ijms23094649 (matulis2022developmentandcharacterization pages 2-4)


2. Core Function: The Tryptophan Synthase Beta Reaction

2.1 Enzymatic Reaction and Substrate Specificity

TrpB (EC 4.2.1.20) catalyzes the β-reaction of tryptophan biosynthesis, the final step in de novo L-tryptophan synthesis (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2):

Reaction:
Indole + L-serine → L-tryptophan + H₂O

This is a pyridoxal 5′-phosphate (PLP)-dependent β-replacement reaction in which the hydroxyl group of L-serine is replaced by indole via a nucleophilic substitution mechanism (ghosh2022allostericregulationof pages 1-2, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 1-3, almhjell2024theβsubunitof media fd21782e).

Sources:
- Ghosh et al., 2022, Frontiers in Molecular Biosciences 9:923042. doi:10.3389/fmolb.2022.923042. URL: https://doi.org/10.3389/fmolb.2022.923042 (ghosh2022allostericregulationof pages 1-2)
- Michalska et al., 2019, IUCrJ 6:649-664. doi:10.1107/s2052252519005955. URL: https://doi.org/10.1107/s2052252519005955 (michalska2019conservationofthe pages 1-2)

2.2 Catalytic Mechanism and Key Intermediates

TrpB employs a sophisticated PLP-dependent mechanism involving multiple covalent intermediates (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2, almhjell2024theβsubunitof media fd21782e):

  1. Internal aldimine: PLP is covalently linked to a catalytic lysine residue (Lys87 in Salmonella typhimurium numbering) via a Schiff base (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2).
  2. External aldimine formation: L-serine displaces the lysine, forming an external aldimine.
  3. Amino-acrylate intermediate (E(A-A)): Dehydration of the serine external aldimine generates a kinetically stable α,β-unsaturated Schiff base (amino-acrylate), a key catalytic species that is resistant to β-elimination (unlike in tyrosine phenol-lyase or tryptophanase) (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 1-2, almhjell2024theβsubunitof media fd21782e).
  4. C–C bond formation: Indole performs a Friedel–Crafts-type nucleophilic attack at the Cβ of the amino-acrylate, forming L-tryptophan (almhjell2024theβsubunitof pages 1-2, almhjell2024theβsubunitof media fd21782e).

Visual evidence of the mechanism is presented in Figure 1c of Almhjell et al. (2024), showing the stable amino-acrylate intermediate bound to PLP and the subsequent C3-alkylation of indole (almhjell2024theβsubunitof media fd21782e).

Sources:
- Almhjell et al., 2024, Nature Chemical Biology 20:1086-1093. doi:10.1038/s41589-024-01619-z. URL: https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 1-2, almhjell2024theβsubunitof media fd21782e)

2.3 Quaternary Structure and Substrate Channeling

TrpB normally functions within a heterotetrameric (α₂β₂) tryptophan synthase complex composed of two TrpA (α) and two TrpB (β) subunits (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2, khan2025multienzymesynergyand pages 12-13):

  • TrpA catalyzes the cleavage of indole-3-glycerol phosphate (IGP) to indole and glyceraldehyde-3-phosphate (G3P) (ghosh2022allostericregulationof pages 1-2, khan2025multienzymesynergyand pages 12-13).
  • Indole channeling: The indole intermediate is transferred from the TrpA active site to the TrpB active site through a ~25 Å hydrophobic tunnel, preventing indole loss to the medium and enhancing catalytic efficiency by ~100-fold (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2, khan2025multienzymesynergyand pages 12-13).
  • Allosteric regulation: The two subunits mutually activate each other through conformational changes mediated by the TrpB COMM (communication) domain and specific loops (e.g., αL6 in TrpA) (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2, michalska2021catalyticallyimpairedtrpa pages 4-7).
  • Open (T) and closed (R) states: Ligand binding and the formation of catalytic intermediates (especially the amino-acrylate) trigger transitions between catalytically inactive (open) and active (closed) conformations, synchronizing the α- and β-reactions (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2).

Sources:
- Ghosh et al., 2022 (ghosh2022allostericregulationof pages 1-2)
- Michalska et al., 2019 (michalska2019conservationofthe pages 1-2)
- Khan & Boehr, 2025, Catalysts 15:718. doi:10.3390/catal15080718. URL: https://doi.org/10.3390/catal15080718 (khan2025multienzymesynergyand pages 12-13)
- Michalska et al., 2021, Protein Science 30:1904-1918. doi:10.1002/pro.4143. URL: https://doi.org/10.1002/pro.4143 (michalska2021catalyticallyimpairedtrpa pages 4-7)

2.4 Structural Features and Domains

TrpB is a fold-type II PLP enzyme with conserved structural elements (michalska2019conservationofthe pages 1-2):

  • N-terminal domain: Contains the COMM (communication) domain, which participates in allosteric signaling and conformational gating (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2).
  • C-terminal domain: Houses the PLP-binding active site and catalytic residues.
  • Key catalytic residues (based on conserved bacterial TrpB sequences): catalytic Lys (e.g., Lys87), a near-universally conserved Glu residue (E105 in many TrpB homologs) that coordinates indole for C–C bond formation, and Asp/Ser residues that stabilize intermediates (khan2025multienzymesynergyand pages 12-13, almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 4-6).

Source: Michalska et al., 2019 (michalska2019conservationofthe pages 1-2); Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 12-13); Almhjell et al., 2024 (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 4-6)


3. Cellular Localization

As a cytoplasmic enzyme involved in amino acid biosynthesis, TrpB is expected to localize to the bacterial cytoplasm where it participates in the tryptophan biosynthetic pathway. No experimental evidence for alternative localization (e.g., membrane association or secretion) has been reported for P. putida KT2440 trpB. Bacterial tryptophan synthases are typically soluble, cytoplasmic enzymes (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2).


4. Biological Pathways and Metabolic Context

4.1 The Tryptophan Biosynthesis Pathway

TrpB functions in the terminal step of the canonical L-tryptophan biosynthesis pathway, which proceeds from chorismate via the shikimate pathway (ghosh2022allostericregulationof pages 1-2, khan2025multienzymesynergyand pages 12-13):

  1. Chorismate (common precursor for aromatic amino acids)
  2. Anthranilate (catalyzed by TrpE/TrpD)
  3. Phosphoribosyl-anthranilate (TrpD)
  4. Indole-3-glycerol phosphate (IGP) (via TrpC, TrpF, TrpG)
  5. Indole + G3P (TrpA, α-reaction)
  6. L-Tryptophan (TrpB, β-reaction)

In P. putida KT2440, the pathway is encoded by genes distributed across the genome (trpBA, trpGDC, trpE, trpF, trpI) rather than a single operon (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2).

Source: Molina-Henares et al., 2009 (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2)

4.2 Regulation and Feedback Control

The trp operon/pathway is subject to transcriptional and allosteric regulation in bacteria:

  • TrpI-mediated activation: In P. putida KT2440 and related pseudomonads, TrpI (a LysR-family regulator) acts as an indole-responsive transcriptional activator of trpA and trpB, in contrast to the TrpR repressor system in E. coli (matulis2022developmentandcharacterization pages 2-4).
  • Indole stress response: Indole accumulation strongly induces trpB expression, potentially enabling cells to re-incorporate indole into tryptophan biosynthesis as a detoxification strategy (kim2013indoletoxicityinvolves pages 3-5, kim2013indoletoxicityinvolves pages 8-9).
  • Feedback inhibition: Canonical regulation involves feedback inhibition of early pathway enzymes by L-tryptophan; TrpB itself is regulated indirectly through pathway flux and TrpA allosteric coupling (ghosh2022allostericregulationof pages 1-2).

Sources: Kim et al., 2013 (kim2013indoletoxicityinvolves pages 3-5, kim2013indoletoxicityinvolves pages 8-9); Matulis et al., 2022 (matulis2022developmentandcharacterization pages 2-4); Ghosh et al., 2022 (ghosh2022allostericregulationof pages 1-2)


5. Recent Developments (2023–2024): Enzyme Engineering and New Catalytic Functions

5.1 2024 Breakthrough: TrpB as a Latent Tyrosine Synthase

A landmark study by Almhjell et al. (2024) demonstrated that TrpB can be engineered into a highly selective tyrosine synthase (TyrS), expanding its catalytic repertoire beyond indole substrates (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 1-2, almhjell2024theβsubunitof pages 4-6):

  • Key discovery: A single substitution of the near-universally conserved catalytic glutamate (E105G) unlocked activity on phenol substrates, giving exclusive para-selective C–C bond formation to produce L-tyrosine and analogs (almhjell2024theβsubunitof pages 2-4).
  • Activity enhancement: E105G alone provided an 18-fold rate increase for 1-naphthol and >100-fold for phenol (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 4-6).
  • Directed evolution: Subsequent rounds of random mutagenesis, site-saturation, and recombination produced TmTyrS6 (22 substitutions from the starting Tm9D8 variant), achieving a net >30,000-fold activity improvement* (almhjell2024theβsubunitof pages 4-6).
  • Selectivity: Engineered TyrS variants show ≥99.5% enantioselectivity and regioselectivity, with no detectable D-Tyr or ortho-alkylation products even at >1,000-fold over the limit of quantification (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 4-6).
  • Catalytic performance: TmTyrS6 exhibits an apparent turnover frequency (TOF) of 0.23 min⁻¹ at 37 °C with 50 mM substrates (almhjell2024theβsubunitof pages 2-4).
  • Mechanistic insight: Structural and computational studies reveal that the E105G mutation allows an active-site water molecule to occupy the position of the removed glutamate, coordinating phenol for concerted C–C bond formation and deprotonation (Figure 4 of Almhjell et al., 2024) (almhjell2024theβsubunitof media 0e8ae960, almhjell2024theβsubunitof pages 30-35, almhjell2024theβsubunitof pages 6-7).
  • Preparative utility: Engineered TyrS enables gram-scale synthesis of L-tyrosine analogs, including β-(1-naphthol-4-yl)-L-alanine (NaphAla) (almhjell2024theβsubunitof pages 30-35, almhjell2024theβsubunitof pages 1-2).

Publication: Almhjell et al., 2024, Nature Chemical Biology 20:1086-1093. doi:10.1038/s41589-024-01619-z. URL: https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 4-6, almhjell2024theβsubunitof pages 30-35, almhjell2024theβsubunitof pages 1-2)

5.2 Stand-Alone TrpB Biocatalysts

Engineering efforts prior to 2024 had already established stand-alone TrpB variants freed from dependence on TrpA (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6):

  • PfTrpB0B2: Created via only three rounds of directed evolution (six mutations), achieving an 83-fold increase in catalytic efficiency over wild-type TrpB and activity exceeding the native αββα complex (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6).
  • Thermostable variants: Engineered TrpB from Pyrococcus furiosus and Thermotoga maritima enable high-temperature reactions (55–75 °C), improving indole solubility and substrate loading (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6).
  • Mutations recapitulate TrpA activation: Beneficial mutations stabilize the catalytically active closed conformation of TrpB and enhance persistence of the amino-acrylate intermediate, mimicking the allosteric effects normally provided by TrpA (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6, khan2025multienzymesynergyand pages 13-15).

Sources:
- Watkins-Dulaney et al., 2021, ChemBioChem 22:5-16. doi:10.1002/cbic.202000379. URL: https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
- Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 13-15)


6. Applications and Real-World Implementations

6.1 Noncanonical Amino Acid (ncAA) Synthesis

Engineered TrpB variants have become workhorse biocatalysts for stereoselective ncAA synthesis (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11):

Substrate scope:
- Halogenated indoles: 4-F, 5-F, 5-Cl, etc. (isolated yields typically 70–99%) (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
- Functionalized indoles: Nitro, cyano, carboxamide, boronate, CF₃, azido, amino-substituted (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
- β-Branched amino acids: L-threonine accepted to produce β-methyltryptophans (>6,000-fold activity boost in PfTrpB2B9 vs. wild-type) (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7)
- Quaternary stereocenter formation: PfTrpBquat shows >99% C3 chemoselectivity on oxindoles, yielding 122 mg product from 1 mmol substrate using 100 mL E. coli culture (52% yield) (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9)

Preparative-scale examples:
- 800 mg 4-cyanotryptophan from 1 L E. coli culture (49% yield) (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
- 965 mg azulene-derived amino acid (AzAla) with 57% isolated yield (TmTrpBAzul; turnover improved from 4.6 to 14.0 min⁻¹) (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11)

Sources: Watkins-Dulaney et al., 2021 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11)

6.2 Metabolic Engineering for L-Tryptophan Production

Engineered TrpB and optimized tryptophan pathways have been deployed in industrial microbial hosts (khan2025multienzymesynergyand pages 15-16):

  • Corynebacterium glutamicum: Tryptophan titers reaching 50.5 g/L (khan2025multienzymesynergyand pages 15-16)
  • L-5-Hydroxytryptophan: 86.7% yield under optimized conditions (khan2025multienzymesynergyand pages 15-16)
  • Indigo biosynthesis: Integration of TrpS with flavin monooxygenase yielded 1,288 mg/L indigo (khan2025multienzymesynergyand pages 15-16)
  • One-pot D-tryptophan synthesis: Multi-enzyme cascades (TrpS + L-amino acid deaminase + D-amino acid transaminase) achieved >99% ee (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11, khan2025multienzymesynergyand pages 15-16)

Sources: Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 15-16); Watkins-Dulaney et al., 2021 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11)

6.3 High-Throughput Screening and Continuous Evolution

Recent advances enable ultra-high-throughput engineering of TrpB (khan2025multienzymesynergyand pages 15-16, khan2025multienzymesynergyand pages 22-23):

  • Droplet microfluidics: Screening of >100,000 TrpB mutants per day, identifying variants with 100-fold activity improvements (khan2025multienzymesynergyand pages 15-16)
  • Continuous flow: Thermostable TrpB variants patented for continuous-flow bioreactors and low-temperature manufacturing (khan2025multienzymesynergyand pages 16-18)

Source: Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 15-16, khan2025multienzymesynergyand pages 22-23, khan2025multienzymesynergyand pages 16-18)


7. Expert Analysis and Current Understanding

7.1 Mechanistic Plasticity

TrpB's kinetically stable amino-acrylate intermediate is the key to its versatility as a biocatalyst (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 1-2, almhjell2024theβsubunitof media fd21782e). Unlike TPL or tryptophanase, which favor β-elimination, TrpB's active site stabilizes the amino-acrylate, enabling kinetic control favoring synthesis over degradation (almhjell2024theβsubunitof pages 1-2). This property has been exploited to:

  • Accept diverse nucleophiles beyond indole (phenols, azulene, nitroalkanes, thiols) (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7, almhjell2024theβsubunitof pages 2-4)
  • Form C–C, C–N, C–S, and C–Se bonds (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7)
  • Catalyze previously unknown reactions (e.g., photoredox-pyridoxal radical biocatalysis for β-methyl amino acids) (khan2025multienzymesynergyand pages 15-16)

Sources: Almhjell et al., 2024 (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 1-2, almhjell2024theβsubunitof media fd21782e); Watkins-Dulaney et al., 2021 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7); Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 15-16)

7.2 Allosteric Engineering

Rational and computational approaches target allosteric networks to improve TrpB activity (khan2025multienzymesynergyand pages 13-15):

  • Mutations distal to the active site can increase the population of catalytically competent conformations, recapitulating TrpA activation in the absence of the α-subunit (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6, khan2025multienzymesynergyand pages 13-15).
  • Transferability: Activating mutations can be transferred across TrpB homologs with ≥57% sequence identity (khan2025multienzymesynergyand pages 13-15).
  • Computational tools: Shortest Path Map (SPM) analysis identifies key residues controlling conformational transitions (khan2025multienzymesynergyand pages 15-16, khan2025multienzymesynergyand pages 13-15).

Sources: Watkins-Dulaney et al., 2021 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6); Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 15-16, khan2025multienzymesynergyand pages 13-15)

7.3 Conservation and Drug Target Potential

Bacterial TrpB proteins show high structural and mechanistic conservation (michalska2019conservationofthe pages 1-2, michalska2021catalyticallyimpairedtrpa pages 4-7):

  • The catalytic glutamate (E105) is conserved in ~98.3% of ~18,051 TrpB-like sequences (almhjell2024theβsubunitof pages 4-6, almhjell2024theβsubunitof pages 30-35).
  • TrpAB is essential for bacterial survival in tryptophan-limited environments (e.g., Mycobacterium tuberculosis in macrophages) and is a validated antimicrobial drug target (michalska2019conservationofthe pages 1-2).
  • Allosteric and active-site inhibitors have been developed, exploiting species-specific structural differences (khan2025multienzymesynergyand pages 16-18).

Sources: Michalska et al., 2019 (michalska2019conservationofthe pages 1-2); Michalska et al., 2021 (michalska2021catalyticallyimpairedtrpa pages 4-7); Almhjell et al., 2024 (almhjell2024theβsubunitof pages 4-6, almhjell2024theβsubunitof pages 30-35); Khan & Boehr, 2025 (khan2025multienzymesynergyand pages 16-18)


8. Summary Table of Key Evidence

A comprehensive evidence table summarizing organism-specific findings for P. putida KT2440 trpB (PP_0083; Q88RP6) and general TrpB knowledge is provided below:

Claim/Topic Key finding Organism/system Quantitative details Source (authors, year, journal) URL Evidence citation id
Identity verification of target gene trpB is explicitly annotated as PP_0083, tryptophan synthase beta subunit, in Pseudomonas putida KT2440, matching UniProt Q88RP6. P. putida KT2440 PP_0083 locus tag Kim et al., 2013, FEMS Microbiology Letters https://doi.org/10.1111/1574-6968.12135 (kim2013indoletoxicityinvolves pages 3-5)
Operon organization trpA and trpB are adjacent, overlap by 1 nucleotide, and are consistent with co-transcription as a trpBA operon. P. putida KT2440 1-nt overlap between trpA and trpB Molina-Henares et al., 2009, Microbial Biotechnology https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4)
Experimental operon confirmation RT-PCR across the trpA/trpB region confirmed in vivo co-transcription of trpA and trpB. P. putida KT2440 cDNA band detected with A-1/B-1 primers; negative control lacked RT Molina-Henares et al., 2009, Microbial Biotechnology https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8)
Broader pathway organization Tryptophan biosynthesis genes are split across unlinked regions: trpBA cluster, trpGDC operon, and monocistronic trpE/trpF; trpI is divergently transcribed from trpBA. P. putida KT2440 Genomic organization into 2 clusters + 2 monocistronic units Molina-Henares et al., 2009, Microbial Biotechnology https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2)
Genetic evidence for pathway function Disruption of trpA causes tryptophan auxotrophy, supporting that the trpBA unit is required for de novo tryptophan synthesis. P. putida KT2440 Aux-1 insertion at 7th codon of trpA; TrpA mutant grows only with tryptophan supplementation Molina-Henares et al., 2009, Microbial Biotechnology https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 1-2)
trpI-associated regulation A trpIAB arrangement was identified; trpI is required for indole-responsive activation of the PP_RS00425 promoter system derived from KT2440. P. putida KT2440 system tested in E. coli and Cupriavidus necator Induction observed with 1 mM indole only when trpI was present Matulis et al., 2022, International Journal of Molecular Sciences https://doi.org/10.3390/ijms23094649 (matulis2022developmentandcharacterization pages 2-4)
Indole-responsive biosensor utility The PpTrpI/PPP_RS00425 gene expression system was developed as an indole biosensor and showed strong, specific induction by indole. KT2440-derived regulatory parts in heterologous hosts Up to 639.6-fold induction; linear response ~0.4-5 mM indole Matulis et al., 2022, International Journal of Molecular Sciences https://doi.org/10.3390/ijms23094649 (matulis2022developmentandcharacterization pages 2-4)
Indole stress response trpB (PP_0083) was the most highly induced gene in KT2440 after indole treatment, linking it to indole-responsive physiology and tryptophan biosynthesis. P. putida KT2440 3.52-fold upregulation; 47 genes changed >1.5-fold up or <0.67-fold down Kim et al., 2013, FEMS Microbiology Letters https://doi.org/10.1111/1574-6968.12135 (kim2013indoletoxicityinvolves pages 3-5, kim2013indoletoxicityinvolves pages 8-9)
Functional interpretation of indole response Authors proposed that degradation/incorporation of indole into tryptophan biosynthesis may help mitigate indole stress. P. putida KT2440 Qualitative interpretation from microarray data Kim et al., 2013, FEMS Microbiology Letters https://doi.org/10.1111/1574-6968.12135 (kim2013indoletoxicityinvolves pages 8-9)
Core enzymatic reaction TrpB catalyzes the PLP-dependent β-reaction converting indole + L-serine to L-tryptophan, the terminal step of tryptophan biosynthesis. Bacterial tryptophan synthase Reaction: indole + L-Ser → L-Trp + H2O Ghosh et al., 2022, Frontiers in Molecular Biosciences https://doi.org/10.3389/fmolb.2022.923042 (ghosh2022allostericregulationof pages 1-2)
Catalytic intermediate and cofactor chemistry TrpB forms a PLP-linked aminoacrylate (EAA) Schiff-base intermediate; PLP is present as an internal aldimine with a catalytic Lys. Bacterial tryptophan synthase Catalytic Lys87 noted in Salmonella numbering Ghosh et al., 2022, Frontiers in Molecular Biosciences; Michalska et al., 2019, IUCrJ https://doi.org/10.3389/fmolb.2022.923042 ; https://doi.org/10.1107/S2052252519005955 (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2)
Quaternary structure and channeling TrpB usually functions in an α2β2 tryptophan synthase complex with TrpA; indole is channeled through an intersubunit tunnel. Bacterial tryptophan synthase ~25 Å tunnel connecting α and β active sites Ghosh et al., 2022, Frontiers in Molecular Biosciences; Michalska et al., 2019, IUCrJ https://doi.org/10.3389/fmolb.2022.923042 ; https://doi.org/10.1107/S2052252519005955 (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2)
Allosteric control and structure TrpB contains an N-terminal COMM domain that participates in open/closed transitions and α-β allosteric communication. Bacterial tryptophan synthase Open (T) and closed (R) conformational states regulate catalysis Ghosh et al., 2022, Frontiers in Molecular Biosciences; Michalska et al., 2019, IUCrJ https://doi.org/10.3389/fmolb.2022.923042 ; https://doi.org/10.1107/S2052252519005955 (ghosh2022allostericregulationof pages 1-2, michalska2019conservationofthe pages 1-2)
Structural conservation Bacterial TrpB proteins are broadly conserved in fold and mechanism, supporting annotation transfer across species when identity is verified. Multiple bacterial pathogens / bacterial TrpAB enzymes TrpB described as more sequence- and fold-conserved than TrpA Michalska et al., 2019, IUCrJ; Michalska et al., 2021, Protein Science https://doi.org/10.1107/S2052252519005955 ; https://doi.org/10.1002/pro.4143 (michalska2019conservationofthe pages 1-2, michalska2021catalyticallyimpairedtrpa pages 4-7)
2024 mechanistic advance A single substitution at the near-universally conserved catalytic glutamate can unlock latent tyrosine synthase activity in TrpB. Engineered TrpB (Tm9D8* lineage) E105G gave 18-fold activation on 1-naphthol and >100-fold on phenol in one context Almhjell et al., 2024, Nature Chemical Biology https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 2-4, almhjell2024theβsubunitof pages 4-6)
2024 evolved activity gains Directed evolution converted TrpB into TyrS variants with very large activity gains and exclusive para-selective C-C bond formation on phenols. Engineered TrpB/TyrS >30,000-fold overall rate enhancement; ≥99.5% enantio- and regioselectivity Almhjell et al., 2024, Nature Chemical Biology https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 4-6, almhjell2024theβsubunitof pages 1-2)
2024 catalytic performance Final TyrS variant synthesized tyrosine with measurable catalytic turnover under preparative conditions. Engineered TrpB/TyrS6 Apparent TOF ~0.23 min⁻1 at 37 °C, 50 mM substrate Almhjell et al., 2024, Nature Chemical Biology https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 2-4)
2024 evolutionary/structural insight The catalytic glutamate targeted in TyrS engineering is highly conserved across TrpB-like sequences, underscoring mechanistic importance. TrpB-like sequence space E105 conserved in ~98.3% of ~18,051 sequences Almhjell et al., 2024, Nature Chemical Biology https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 4-6, almhjell2024theβsubunitof pages 30-35)
2024 practical application Engineered TyrS enzymes enabled preparative/gram-scale synthesis of tyrosine analogs. Engineered TrpB/TyrS Preparative scale; multi-gram / gram-scale products reported Almhjell et al., 2024, Nature Chemical Biology https://doi.org/10.1038/s41589-024-01619-z (almhjell2024theβsubunitof pages 30-35, almhjell2024theβsubunitof pages 1-2)
Stand-alone TrpB biocatalyst engineering Directed evolution liberated TrpB from dependence on TrpA, creating stand-alone catalysts suitable for synthesis of noncanonical amino acids. Engineered PfTrpB / TmTrpB variants PfTrpB0B2 obtained with 6 mutations after 3 rounds; 83-fold increase in catalytic efficiency Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
Substrate scope and yields Evolved TrpB variants accept many substituted indoles and often give high isolated yields. Engineered TrpB biocatalysts Typical isolated yields 70-99% across many analogs Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
Preparative Trp analog synthesis A TmTrpB variant was used to synthesize 4-cyanotryptophan on preparative scale. Engineered TmTrpB9D8 in E. coli* cells 800 mg 4-cyanoTrp from 1 L culture; 49% yield Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6)
Expanded bond-forming chemistry Engineered TrpB variants catalyze noncanonical C-C, C-N, C-S, and C-Se bond-forming reactions. Engineered TrpB biocatalysts Up to 2,700 turnovers on nitro-containing substrates Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7)
Quaternary-center synthesis Engineered PfTrpBquat shifted chemoselectivity to C3 alkylation of oxindoles, enabling quaternary stereocenter formation. Engineered PfTrpBquat >99% C3 chemoselectivity; 52% yield; 122 mg from 1 mmol substrate using 100 mL E. coli culture Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9)
Azulene amino acid production Evolution improved azulene alkylation and enabled gram-scale synthesis of AzAla. Engineered TmTrpBAzul Activity improved from 4.6 to 14.0 turnovers/min; 965 mg product; 57% isolated yield Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9, watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11)
β-branched amino acid synthesis TrpB engineering enabled efficient β-methyltryptophan synthesis from threonine-derived chemistry. Engineered PfTrpB variants >6,000-fold boost in β-methylTrp activity vs WT; one variant used 1 equiv Thr and gave 3.5-fold higher yield than prior variant needing 10 equiv Watkins-Dulaney et al., 2021, ChemBioChem https://doi.org/10.1002/cbic.202000379 (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7)

Table: This table compiles organism-specific evidence for Pseudomonas putida KT2440 trpB (PP_0083; UniProt Q88RP6) together with conserved TrpB mechanistic knowledge and recent engineering/application advances. It is useful for separating direct evidence about the target gene from broader, high-confidence functional inference about TrpB enzymes.


9. Conclusions

The trpB gene (PP_0083; UniProt Q88RP6) in Pseudomonas putida KT2440 encodes the tryptophan synthase beta chain, a PLP-dependent enzyme catalyzing the final step of L-tryptophan biosynthesis. Organism-specific evidence confirms:

  1. trpB is co-transcribed with trpA in a trpBA operon and is essential for tryptophan prototrophy.
  2. Expression is strongly induced by indole (3.52-fold), suggesting a role in indole detoxification/re-assimilation.
  3. Regulation is mediated by TrpI, a LysR-family activator responsive to indole or indole-3-glycerol phosphate.

General TrpB enzymology reveals:

  • A sophisticated PLP-dependent mechanism involving a stable amino-acrylate intermediate and Friedel–Crafts alkylation.
  • Quaternary structure and allosteric regulation within the αββα tryptophan synthase complex, with indole channeling through a ~25 Å tunnel.
  • Broad catalytic plasticity enabling engineering for noncanonical amino acid synthesis, with applications ranging from gram-scale ncAA production to metabolic engineering (titers up to 50.5 g/L).

2024 advances include the discovery that TrpB is a latent tyrosine synthase (>30,000-fold activity gain via directed evolution; ≥99.5% selectivity) and the development of ultra-high-throughput screening platforms for continuous TrpB optimization. These findings establish TrpB as a versatile biocatalytic platform for sustainable synthesis of aromatic amino acids and derivatives.


References (Selected)

  • Kim et al., 2013. FEMS Microbiol. Lett. 343(1):89-99. https://doi.org/10.1111/1574-6968.12135
  • Molina-Henares et al., 2009. Microb. Biotechnol. 2:91-100. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  • Matulis et al., 2022. Int. J. Mol. Sci. 23:4649. https://doi.org/10.3390/ijms23094649
  • Ghosh et al., 2022. Front. Mol. Biosci. 9:923042. https://doi.org/10.3389/fmolb.2022.923042
  • Michalska et al., 2019. IUCrJ 6:649-664. https://doi.org/10.1107/s2052252519005955
  • Watkins-Dulaney et al., 2021. ChemBioChem 22:5-16. https://doi.org/10.1002/cbic.202000379
  • Almhjell et al., 2024. Nat. Chem. Biol. 20:1086-1093. https://doi.org/10.1038/s41589-024-01619-z
  • Khan & Boehr, 2025. Catalysts 15:718. https://doi.org/10.3390/catal15080718

Report compiled from peer-reviewed primary literature and authoritative reviews (2009–2025). All major claims are cited with context IDs traceable to source evidence.

References

  1. (kim2013indoletoxicityinvolves pages 3-5): Jisun Kim, Hyerim Hong, Aram Heo, and Woojun Park. Indole toxicity involves the inhibition of adenosine triphosphate production and protein folding in pseudomonas putida. FEMS microbiology letters, 343 1:89-99, Jun 2013. URL: https://doi.org/10.1111/1574-6968.12135, doi:10.1111/1574-6968.12135. This article has 72 citations and is from a peer-reviewed journal.

  2. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  3. (molinahenares2009functionalanalysisof pages 4-6): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  4. (molinahenares2009functionalanalysisof pages 1-2): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  5. (molinahenares2009functionalanalysisof pages 7-8): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  6. (matulis2022developmentandcharacterization pages 2-4): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  7. (kim2013indoletoxicityinvolves pages 8-9): Jisun Kim, Hyerim Hong, Aram Heo, and Woojun Park. Indole toxicity involves the inhibition of adenosine triphosphate production and protein folding in pseudomonas putida. FEMS microbiology letters, 343 1:89-99, Jun 2013. URL: https://doi.org/10.1111/1574-6968.12135, doi:10.1111/1574-6968.12135. This article has 72 citations and is from a peer-reviewed journal.

  8. (ghosh2022allostericregulationof pages 1-2): Rittik K. Ghosh, Eduardo Hilario, Chia-en A. Chang, Leonard J. Mueller, and Michael F. Dunn. Allosteric regulation of substrate channeling: salmonella typhimurium tryptophan synthase. Frontiers in Molecular Biosciences, Sep 2022. URL: https://doi.org/10.3389/fmolb.2022.923042, doi:10.3389/fmolb.2022.923042. This article has 14 citations.

  9. (michalska2019conservationofthe pages 1-2): Karolina Michalska, Jennifer Gale, Grazyna Joachimiak, Changsoo Chang, Catherine Hatzos-Skintges, Boguslaw Nocek, Stephen E. Johnston, Lance Bigelow, Besnik Bajrami, Robert P. Jedrzejczak, Samantha Wellington, Deborah T. Hung, Partha P. Nag, Stewart L. Fisher, Michael Endres, and Andrzej Joachimiak. Conservation of the structure and function of bacterial tryptophan synthases. IUCrJ, 6:649-664, May 2019. URL: https://doi.org/10.1107/s2052252519005955, doi:10.1107/s2052252519005955. This article has 22 citations and is from a peer-reviewed journal.

  10. (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 1-3): Ella Watkins‐Dulaney, Sabine Straathof, and Frances Arnold. Tryptophan synthase: biocatalyst extraordinaire. Sep 2021. URL: https://doi.org/10.1002/cbic.202000379, doi:10.1002/cbic.202000379. This article has 124 citations and is from a peer-reviewed journal.

  11. (almhjell2024theβsubunitof media fd21782e): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  12. (almhjell2024theβsubunitof pages 2-4): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  13. (almhjell2024theβsubunitof pages 1-2): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  14. (khan2025multienzymesynergyand pages 12-13): Sara Khan and David D. Boehr. Multi-enzyme synergy and allosteric regulation in the shikimate pathway: biocatalytic platforms for industrial applications. Catalysts, 15:718, Jul 2025. URL: https://doi.org/10.3390/catal15080718, doi:10.3390/catal15080718. This article has 6 citations.

  15. (michalska2021catalyticallyimpairedtrpa pages 4-7): Karolina Michalska, Samantha Wellington, Natalia Maltseva, Robert Jedrzejczak, Nelly Selem‐Mojica, L. Rodrigo Rosas‐Becerra, Francisco Barona‐Gómez, Deborah T. Hung, and Andrzej Joachimiak. Catalytically impaired trpa subunit of tryptophan synthase from chlamydia trachomatis is an allosteric regulator of trpb. Protein Science, 30:1904-1918, Jun 2021. URL: https://doi.org/10.1002/pro.4143, doi:10.1002/pro.4143. This article has 16 citations and is from a peer-reviewed journal.

  16. (almhjell2024theβsubunitof pages 4-6): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  17. (almhjell2024theβsubunitof media 0e8ae960): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  18. (almhjell2024theβsubunitof pages 30-35): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  19. (almhjell2024theβsubunitof pages 6-7): Patrick J. Almhjell, Kadina E. Johnston, Nicholas J. Porter, Jennifer L. Kennemur, Vignesh C. Bhethanabotla, Julie Ducharme, and Frances H. Arnold. The β-subunit of tryptophan synthase is a latent tyrosine synthase. Nature chemical biology, 20:1086-1093, May 2024. URL: https://doi.org/10.1038/s41589-024-01619-z, doi:10.1038/s41589-024-01619-z. This article has 37 citations and is from a highest quality peer-reviewed journal.

  20. (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 4-6): Ella Watkins‐Dulaney, Sabine Straathof, and Frances Arnold. Tryptophan synthase: biocatalyst extraordinaire. Sep 2021. URL: https://doi.org/10.1002/cbic.202000379, doi:10.1002/cbic.202000379. This article has 124 citations and is from a peer-reviewed journal.

  21. (khan2025multienzymesynergyand pages 13-15): Sara Khan and David D. Boehr. Multi-enzyme synergy and allosteric regulation in the shikimate pathway: biocatalytic platforms for industrial applications. Catalysts, 15:718, Jul 2025. URL: https://doi.org/10.3390/catal15080718, doi:10.3390/catal15080718. This article has 6 citations.

  22. (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 7-9): Ella Watkins‐Dulaney, Sabine Straathof, and Frances Arnold. Tryptophan synthase: biocatalyst extraordinaire. Sep 2021. URL: https://doi.org/10.1002/cbic.202000379, doi:10.1002/cbic.202000379. This article has 124 citations and is from a peer-reviewed journal.

  23. (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 6-7): Ella Watkins‐Dulaney, Sabine Straathof, and Frances Arnold. Tryptophan synthase: biocatalyst extraordinaire. Sep 2021. URL: https://doi.org/10.1002/cbic.202000379, doi:10.1002/cbic.202000379. This article has 124 citations and is from a peer-reviewed journal.

  24. (watkins‐dulaney2021tryptophansynthasebiocatalyst pages 9-11): Ella Watkins‐Dulaney, Sabine Straathof, and Frances Arnold. Tryptophan synthase: biocatalyst extraordinaire. Sep 2021. URL: https://doi.org/10.1002/cbic.202000379, doi:10.1002/cbic.202000379. This article has 124 citations and is from a peer-reviewed journal.

  25. (khan2025multienzymesynergyand pages 15-16): Sara Khan and David D. Boehr. Multi-enzyme synergy and allosteric regulation in the shikimate pathway: biocatalytic platforms for industrial applications. Catalysts, 15:718, Jul 2025. URL: https://doi.org/10.3390/catal15080718, doi:10.3390/catal15080718. This article has 6 citations.

  26. (khan2025multienzymesynergyand pages 22-23): Sara Khan and David D. Boehr. Multi-enzyme synergy and allosteric regulation in the shikimate pathway: biocatalytic platforms for industrial applications. Catalysts, 15:718, Jul 2025. URL: https://doi.org/10.3390/catal15080718, doi:10.3390/catal15080718. This article has 6 citations.

  27. (khan2025multienzymesynergyand pages 16-18): Sara Khan and David D. Boehr. Multi-enzyme synergy and allosteric regulation in the shikimate pathway: biocatalytic platforms for industrial applications. Catalysts, 15:718, Jul 2025. URL: https://doi.org/10.3390/catal15080718, doi:10.3390/catal15080718. This article has 6 citations.

Artifacts

Citations

  1. kim2013indoletoxicityinvolves pages 3-5
  2. molinahenares2009functionalanalysisof pages 2-4
  3. molinahenares2009functionalanalysisof pages 1-2
  4. kim2013indoletoxicityinvolves pages 8-9
  5. matulis2022developmentandcharacterization pages 2-4
  6. ghosh2022allostericregulationof pages 1-2
  7. michalska2019conservationofthe pages 1-2
  8. khan2025multienzymesynergyand pages 12-13
  9. michalska2021catalyticallyimpairedtrpa pages 4-7
  10. khan2025multienzymesynergyand pages 13-15
  11. khan2025multienzymesynergyand pages 15-16
  12. khan2025multienzymesynergyand pages 16-18
  13. molinahenares2009functionalanalysisof pages 4-6
  14. molinahenares2009functionalanalysisof pages 7-8
  15. khan2025multienzymesynergyand pages 22-23
  16. https://doi.org/10.1111/1574-6968.12135
  17. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  18. https://doi.org/10.3390/ijms23094649
  19. https://doi.org/10.3389/fmolb.2022.923042
  20. https://doi.org/10.1107/s2052252519005955
  21. https://doi.org/10.1038/s41589-024-01619-z
  22. https://doi.org/10.3390/catal15080718
  23. https://doi.org/10.1002/pro.4143
  24. https://doi.org/10.1002/cbic.202000379
  25. https://doi.org/10.1107/S2052252519005955
  26. https://doi.org/10.1111/1574-6968.12135,
  27. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  28. https://doi.org/10.3390/ijms23094649,
  29. https://doi.org/10.3389/fmolb.2022.923042,
  30. https://doi.org/10.1107/s2052252519005955,
  31. https://doi.org/10.1002/cbic.202000379,
  32. https://doi.org/10.1038/s41589-024-01619-z,
  33. https://doi.org/10.3390/catal15080718,
  34. https://doi.org/10.1002/pro.4143,

📄 View Raw YAML

id: Q88RP6
gene_symbol: trpB
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: Tryptophan synthase beta chain (TrpB, EC 4.2.1.20), a pyridoxal 5'-phosphate (PLP)-dependent enzyme that catalyzes the final (beta) reaction of L-tryptophan biosynthesis, condensing indole with L-serine to yield L-tryptophan and water. PLP is bound as an internal aldimine to an active-site lysine (residue 95 in this protein). TrpB is a member of the fold-type II PLP enzyme family (TrpB family) and normally assembles with the alpha subunit (TrpA) into the alpha2-beta2 tryptophan synthase complex, in which the indole produced by TrpA from indole-3-glycerol phosphate is channeled directly to the TrpB active site through an intramolecular tunnel. TrpB carries out the terminal, fifth step of the conversion of chorismate to L-tryptophan and is a soluble cytoplasmic enzyme. In P. putida KT2440 the gene (PP_0083) is adjacent to and co-transcribed with trpA (PP_0082) as a trpBA operon, and its expression is strongly induced by indole.
core_functions:
- description: Catalyzes the PLP-dependent beta-replacement reaction forming L-tryptophan from indole (channeled from TrpA) and L-serine, completing the terminal step of L-tryptophan biosynthesis
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: EC=4.2.1.20; tryptophan synthase activity inferred from InterPro, RHEA:10532, UniRule and PANTHER (TrpB family).
  - reference_id: file:PSEPK/trpB/trpB-uniprot.txt
    supporting_text: "FUNCTION: The beta subunit is responsible for the synthesis of L-tryptophan from indole and L-serine. CATALYTIC ACTIVITY: indol-3-yl glycerol 3-phosphate + L-serine = D-glyceraldehyde 3-phosphate + L-tryptophan + H2O; PATHWAY: L-tryptophan from chorismate, step 5/5; COFACTOR: pyridoxal 5'-phosphate."
  molecular_function:
    id: GO:0004834
    label: tryptophan synthase activity
  directly_involved_in:
  - id: GO:0000162
    label: L-tryptophan biosynthetic process
  locations:
  - id: GO:0005737
    label: cytoplasm
existing_annotations:
- term:
    id: GO:0000162
    label: L-tryptophan biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: TrpB catalyzes the terminal step of L-tryptophan biosynthesis; this annotation correctly captures the core biological process.
    reason: The protein is a UniProt-reviewed tryptophan synthase beta chain (EC 4.2.1.20) belonging to the TrpB family, with UniPathway UPA00035 (L-tryptophan biosynthesis, step 5/5). In P. putida KT2440 trpA disruption causes tryptophan auxotrophy, confirming the trpBA cluster is required for de novo tryptophan synthesis (PMID:21261884; see also file:PSEPK/trpB/trpB-deep-research-falcon.md). The IEA assignment is well-supported and represents a core function.
    action: ACCEPT
- term:
    id: GO:0004834
    label: tryptophan synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: Correct molecular function. TrpB is the tryptophan synthase beta subunit catalyzing the PLP-dependent beta-reaction (indole + L-serine -> L-tryptophan + H2O).
    reason: Supported by EC 4.2.1.20, RHEA:10532, the conserved PLP-binding lysine (residue 95), HAMAP-Rule MF_00133, and TrpB-family InterPro/PANTHER signatures. This is the core enzymatic activity of the gene product.
    action: ACCEPT
- term:
    id: GO:0005737
    label: cytoplasm
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: Bacterial tryptophan synthase is a soluble cytoplasmic enzyme; cytoplasmic localization is correct.
    reason: TrpB has no signal peptide or transmembrane region and functions in cytoplasmic amino-acid biosynthesis as part of the soluble alpha2-beta2 tryptophan synthase complex. The TreeGrafter IEA assignment is consistent with the well-established localization of this enzyme family. The term is somewhat generic but accurate for a bacterial cytosolic enzyme.
    action: ACCEPT
references:
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: file:PSEPK/trpB/trpB-uniprot.txt
  title: UniProt entry TRPB_PSEPK (Q88RP6)
  findings:
  - statement: TrpB synthesizes L-tryptophan from indole and L-serine; PLP cofactor; pathway L-tryptophan from chorismate step 5/5; functions as a tetramer of two alpha and two beta chains.
    supporting_text: "FUNCTION: The beta subunit is responsible for the synthesis of L-tryptophan from indole and L-serine. SUBUNIT: Tetramer of two alpha and two beta chains."
- id: PMID:21261884
  title: "Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440"
  findings:
  - statement: In P. putida KT2440 there is a single pathway from chorismate to tryptophan; the trp genes are in unlinked regions with trpBA organized as an operon, and auxotroph screening shows the pathway is required for de novo tryptophan biosynthesis.
    supporting_text: "Genes for tryptophan biosynthesis are grouped in unlinked regions with the trpBA and trpGDE genes organized as operons... There is a single pathway from chorismate leading to the biosynthesis of tryptophan."
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: "PubMed-verified (PMID:21261884, DOI 10.1111/j.1751-7915.2008.00062.x, Molina-Henares et al., Microb Biotechnol 2:91-100). Abstract confirms the trpBA operon organization and tryptophan-auxotroph screening in KT2440. Establishes the P. putida KT2440-specific genomic context and pathway essentiality for trpB."