trpF

UniProt ID: Q88LE0
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

trpF encodes N-(5'-phosphoribosyl)anthranilate isomerase (PRAI, EC 5.3.1.24), a cytoplasmic monomeric TIM-barrel ((beta/alpha)8) enzyme that catalyzes the third step of L-tryptophan biosynthesis from chorismate. It converts N-(5-phospho-beta-D-ribosyl)anthranilate (PRA) into 1-(2-carboxyphenylamino)-1-deoxy-D-ribulose 5-phosphate (CdRP) via an Amadori rearrangement that opens the ribose ring. This intermediate is subsequently used by TrpC (indole-3-glycerol-phosphate synthase) and tryptophan synthase (TrpAB) to complete tryptophan biosynthesis. In Pseudomonas putida KT2440 the gene (locus PP_1995) is unlinked to the other trp clusters and is most likely monocistronic; a targeted chromosomal knockout produces a tryptophan auxotroph that is rescued by tryptophan or indole but not by anthranilate, placing the enzyme downstream of anthranilate and upstream of indole as expected for PRAI.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000162 L-tryptophan biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: Correct and core. PRAI catalyzes step 3 of 5 in tryptophan biosynthesis from chorismate, and the KT2440 trpF knockout is a tryptophan auxotroph, directly confirming a required role in this process.
Reason: The biological process is supported both by family/pathway assignment (UniPathway UPA00035) and by experimental auxotrophy/complementation evidence in KT2440 (PMID:21261884).
GO:0004640 phosphoribosylanthranilate isomerase activity
IEA
GO_REF:0000120
ACCEPT
Summary: Correct core molecular function. This is the precise EC 5.3.1.24 activity (RHEA:21540) defining the TrpF family, consistent with the HAMAP rule, InterPro TrpF family (IPR044643) and PRAI domain (IPR001240), and the conserved TIM-barrel active site.
Reason: Directly matches the curated catalytic activity in UniProt and the protein's family/domain assignment; this is the defining function of the gene product.
GO:0046394 carboxylic acid biosynthetic process
IEA
GO_REF:0000117
MARK AS OVER ANNOTATED
Summary: Not wrong but uninformative and over-general. Tryptophan is a carboxylic acid, so this high-level ARBA-derived term is technically true but adds nothing beyond the more specific GO:0000162 (L-tryptophan biosynthetic process), which is already annotated.
Reason: Generic parent term auto-generated from sequence features; the specific L-tryptophan biosynthetic process annotation already captures the biology precisely.

Core Functions

Catalyzes the isomerization of N-(5-phospho-beta-D-ribosyl)anthranilate to 1-(2-carboxyphenylamino)-1-deoxy-D-ribulose 5-phosphate (CdRP), the third step of L-tryptophan biosynthesis.

Supporting Evidence:
  • PMID:21261884
    Targeted chromosomal knockout of trpF (PP_1995) in KT2440 yields a tryptophan auxotroph rescued by tryptophan or indole but not anthranilate, consistent with loss of PRAI activity acting downstream of anthranilate and upstream of indole.

References

Electronic Gene Ontology annotations created by ARBA machine learning models
Combined Automated Annotation using Multiple IEA Methods
Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440
  • trpF (PP_1995) is unlinked to the other trp gene clusters and most likely monocistronic; its knockout is a tryptophan auxotroph rescued by tryptophan or indole but not by anthranilate, confirming its role as PRAI in tryptophan biosynthesis.

Deep Research

Asta

(trpF-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-05T20:22:12.385847

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.748)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.665)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[2] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.730)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[3] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.723)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[4] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.717)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[5] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.713)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[6] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.702)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[7] AgAnimalGenomes: browsers for viewing and manually annotating farm animal genomes

  • Authors: D. Triant, Amy T. Walsh, Gabrielle Hartley, B. Petry, Morgan R. Stegemiller et al.
  • Year: 2023
  • Venue: Mammalian Genome
  • URL: https://www.semanticscholar.org/paper/38a969fd5641e503106cb215010f84ea0a271f99
  • DOI: 10.1007/s00335-023-10008-1
  • PMID: 37460664
  • PMCID: 10382368
  • Citations: 5
  • Summary: This work presents genome visualization and annotation tools to support seven livestock species, available in a new resource called AgAnimalGenomes, and describes the data and search methods available and how to use the provided tools to edit and create new gene models.
  • Evidence snippets:
  • Snippet 1 (score: 0.698)
    > As previously described (Triant et al. 2020), once a proteincoding gene annotation is complete, each new or modified isoform should be compared to a well-curated protein sequence database to check for congruency with known proteins. The sequence of an annotation is obtained by right clicking it and selecting Get Sequence. The first choice of database to search is the well-curated UniProtKB/Swissprot database using BLAST at either the UniProt (https:// www. unipr ot. org/ blast) or NCBI website (https:// blast. ncbi. nlm. nih. gov/ Blast. cgi) (Sayers et al. 2023a;UniProt Consortium 2023). If there is no match with a significant e-value (< 1e−05) in UniProtKB/Swissprot, the next database to try is the Model Organisms (landmark) database at NCBI. If that fails, select the RefSeq Proteins database and exclude your organism of interest from the search. Although RefSeq includes computationally predicted and hypothetical proteins, an alignment to a homologous protein from another organism provides support for the annotation. An alignment that covers the full length of both the annotated protein and the database protein sequence suggests the annotation is correct. An alignment that encompasses the full length of an annotated protein sequence but only part of a database protein suggests that the annotation is truncated. You may be able to correct the annotation with additional evidence, but if there is not sufficient evidence the issue can be noted in the Annotation Information Panel under the Comment tab. A partial alignment of an annotated protein to a database protein suggests the annotation has a reading frame shift or was extended incorrectly. Aligning the coding sequence (CDS) to the protein database will reveal whether the problem is due to a reading frame shift. Further annotation editing should be performed to correct the reading frame. If an incorrect extension was due to the merging of two genes, you should edit or redo the annotation. Any unresolved issues should be entered in the Comment section of the Annotation Information Panel.

[8] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.692)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[9] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.690)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[10] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.688)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[11] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.680)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[12] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.679)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[13] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[14] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[15] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[16] Characterization of holins, the membrane proteins of coliphage ASEC2201: a genomewide in silico approach

  • Authors: Humaira Saeed, Sudhaker Padmesh, Aditi Singh, S. Singh, Mohammed Haris Siddiqui et al.
  • Year: 2025
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/a39392e12bf3bda67bdfe600053e8403deb3b887
  • DOI: 10.3389/fmicb.2025.1550594
  • PMID: 40703241
  • PMCID: 12283622
  • Citations: 3
  • Summary: In silico identification of cell-penetrating peptide motifs within the holin sequences suggests potential for enhanced intracellular delivery in CPP-fusion therapeutic constructs and demonstrates the potential of integrative in silico approaches in developing a comprehensive foundation for future experimental validation for proteins with no prior functional annotation.
  • Evidence snippets:
  • Snippet 1 (score: 0.664)
    > Protein-coding gene annotation is typically a two-step process. Initially, Prodigal is employed to identify open reading frames (ORFs) by locating gene coordinates, but it does not infer gene function. To assign putative functions, Prokka performs hierarchical annotation by comparing candidate genes to curated protein databases. It begins with a user-supplied, high-confidence protein set, using BLAST+ for sequence similarity searches. If no match is found, it progresses to UniProt's verified bacterial proteins, covering \~ 16,000 sequences, and then optionally to RefSeq proteins specific to the organism's genuscapturing nomenclature consistency. When sequence-based annotation fails, Prokka applies profile-based searches using HMMER's hmmscan to query against Pfam and TIGRFAMs databases. An e-value threshold of 10 −6 is consistently applied to ensure significance. If no reliable match is found across all levels, the gene is designated as a "hypothetical protein. " This layered strategy maximizes annotation accuracy and functional insight across diverse bacterial genomes (Seemann, 2014).
    > The genome of coliphage ASEC2201 has been analyzed and three holin protein coding genes were selected. The sequences of all three holin proteins were retrieved from NCBI using accession no. SRX17770782 in the FASTA format. The sequence similarity search was performed via BLAST against the non-redundant database (Altschul et al., 1990).

[17] Functional annotation of parasitic worm genomes, by assigning protein names and GO terms

  • Authors: Avril Coghlan, M. Berriman
  • Year: 2018
  • Venue: Unknown venue
  • URL: https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  • DOI: 10.1038/protex.2018.055
  • Citations: 1
  • Summary: A computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and Go terms from orthologues in other species.
  • Evidence snippets:
  • Snippet 1 (score: 0.664)
    > Given a set of predicted protein-coding genes for a newly sequenced genome, functional annotation involves assigning putative functions to the predicted genes. Two ways in which this can be done are assigning protein names and Gene Ontology (GO;Gene Ontology Consortium, 2010) terms to the predicted proteins. Here we describe a computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and GO terms from orthologues in other species.
    > When assigning protein names, UniProt protein naming rules (www.uniprot.org/docs/nameprot) are followed where possible. This recommends that a good and stable name for a protein is "as neutral as possible"; that a protein name "should be, as far as possible, unique and attributed to all orthologs"; and a protein name "should not contain a specific characteristic of the protein, and in particular it should not reflect the function or role of the protein, nor its subcellular location, its domain structure, its tissue specificity, its molecular weight or its species of origin".
    > In our protocol, a protein name is assigned to each predicted protein based on curated names in UniProt (Bairoch & Apweiler, 2000) for human, zebrafish, Drosophila melanogaster, Caenorhabditis elegans, and Schistosoma mansoni orthologues identified from a database of gene families (e.g. built using Ensembl Compara; Vilella et al. 2009), or (if no information is found from orthologues) based on InterPro (Hunter et al. 2012) domains. Figure 1 shows an example of using our protein naming pipeline for four Strongyloides ratti genes that belong to the tubulin polyglutamylase family (underlined in pink), where four different protein names were assigned to them (in pink), based on names of their C. elegans or human orthologues.
    > Since each of the S. ratti genes belonged to a different subfamily of the tubulin polyglutamylase family, they were assigned different names.

[18] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[19] Role of histone-lysine N-methyltransferase 2D (KMT2D) in MEK-ERK signaling-mediated epigenetic regulation: a phosphoproteomics perspective

  • Authors: Sreeshma Ravindran Kammarambath, Leona Dcunha, Athira Perunelly Gopalakrishnan, Amal Fahma, N. Krishna et al.
  • Year: 2025
  • Venue: Frontiers in Bioinformatics
  • URL: https://www.semanticscholar.org/paper/0ac0729148aff3d839e6a15984e11532e9e740f9
  • DOI: 10.3389/fbinf.2025.1683469
  • PMID: 41341998
  • PMCID: 12669113
  • Citations: 3
  • Summary: The phosphoregulatory network of Histone-lysine N-methyltransferase 2D is delineated, positioning it as a dynamic epigenetic effector modulated by MEK-ERK signaling, with broader implications for cancer and developmental disorders.
  • Evidence snippets:
  • Snippet 1 (score: 0.661)
    > Each protein was mapped to its corresponding gene symbol based on the HGNC (downloaded on 30.05.2023) and to its corresponding UniProt (13.04.2023) (UniProt, 2023) accessions using our in-built mapping tool to ensure consistent and standardized annotation. We conducted the analysis using the methodologies outlined in (Sanjeev et al., 2024). The overall workflow used in this study is outlined in Figure 1.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  2. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  3. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  4. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  5. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  6. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  7. D. Triant, Amy T. Walsh, Gabrielle Hartley, B. Petry, Morgan R. Stegemiller et al. (2023). AgAnimalGenomes: browsers for viewing and manually annotating farm animal genomes. Mammalian Genome. https://www.semanticscholar.org/paper/38a969fd5641e503106cb215010f84ea0a271f99
  8. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  9. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  10. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  11. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  12. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  13. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  14. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  15. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  16. Humaira Saeed, Sudhaker Padmesh, Aditi Singh, S. Singh, Mohammed Haris Siddiqui et al. (2025). Characterization of holins, the membrane proteins of coliphage ASEC2201: a genomewide in silico approach. Frontiers in Microbiology. https://www.semanticscholar.org/paper/a39392e12bf3bda67bdfe600053e8403deb3b887
  17. Avril Coghlan, M. Berriman (2018). Functional annotation of parasitic worm genomes, by assigning protein names and GO terms. https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  18. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  19. Sreeshma Ravindran Kammarambath, Leona Dcunha, Athira Perunelly Gopalakrishnan, Amal Fahma, N. Krishna et al. (2025). Role of histone-lysine N-methyltransferase 2D (KMT2D) in MEK-ERK signaling-mediated epigenetic regulation: a phosphoproteomics perspective. Frontiers in Bioinformatics. https://www.semanticscholar.org/paper/0ac0729148aff3d839e6a15984e11532e9e740f9

Falcon

(trpF-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 25 citations 2 artifacts 2026-06-11T21:45:57.896548

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: trpF (UniProt Q88LE0; locus PP_1995) in Pseudomonas putida KT2440

0) Identity verification (mandatory disambiguation)

The target protein is unambiguously Pseudomonas putida KT2440 trpF (locus tag PP_1995), functionally assigned as phosphoribosyl anthranilate isomerase (PRAI) in the tryptophan biosynthesis pathway. This is supported by a KT2440-specific targeted chromosomal knockout of trpF (PP_1995) and pathway-consistent auxotrophy/complementation assays, confirming that this trpF is the canonical tryptophan-pathway enzyme rather than an unrelated protein sharing the symbol (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8, molinahenares2009functionalanalysisof media f6402c94).

1) Key concepts and definitions (current understanding)

1.1 Enzyme and reaction definition

TrpF (phosphoribosyl anthranilate isomerase; PRAI) catalyzes an isomerization step in the conversion of anthranilate-derived intermediates toward tryptophan. In pathway terms, anthranilate is first converted to phosphoribosyl anthranilate (PRA) by TrpD; TrpF then opens the ribose ring of PRA to yield 1-carboxyphenylamino-1′-deoxyribulose-5′-phosphate (CdRP), and TrpC then converts CdRP onward toward indole ring formation (guida2024aminoacidbiosynthesis pages 4-6).

A mechanistic picture consistent with this biochemical definition is supported by modern computational/structural analyses of TrpF-family enzymes: TrpF catalysis is described as proceeding through an Amadori-type rearrangement and involves enzyme-catalyzed opening of the ribose ring (romerorivera2022complexloopdynamics pages 2-3, romerorivera2022complexloopdynamics pages 3-4).

1.2 Pathway context in Pseudomonas

In Pseudomonas, tryptophan biosynthesis genes are often not in a single contiguous operon; instead, they can be distributed in separate loci. In KT2440 specifically, trpBA and trpGDE form operons, whereas trpF (and some other trp genes) are organized as single transcriptional units (molinahenares2009functionalanalysisof pages 1-2). Downstream in the pathway, TrpA/TrpB (tryptophan synthase) convert indole + L-serine to L-tryptophan; regulation can differ from E. coli and in some Pseudomonas involves the LysR-family regulator TrpI influencing expression of tryptophan synthase genes (matulis2022developmentandcharacterization pages 1-2).

2) KT2440-specific functional annotation: gene organization, phenotype, and pathway placement

2.1 Gene organization and genomic context

Molina-Henares et al. (2009) show that KT2440 trpF/PP_1995 is unlinked to the main trp gene clusters and is likely monocistronic. They report short intergenic distances suggestive of independent transcriptional units (PP1994 is 61 nt upstream; PP1996 begins 222 nt after the PP_1995 stop codon), and RT-PCR did not support cotranscription, leading them to conclude trpF is “most probably a single cistron” (molinahenares2009functionalanalysisof pages 2-4). They further note that in sequenced Pseudomonas spp. trpF is consistently flanked by truA and accD (molinahenares2009functionalanalysisof pages 4-6).

2.2 Targeted knockout and auxotrophy

Because spontaneous trpF mutants were not recovered, Molina-Henares et al. constructed a site-specific chromosomal trpF (PP_1995) knockout using a pCHESI-based strategy; insertion was confirmed by PCR and Southern blotting (molinahenares2009functionalanalysisof pages 7-8). The resulting mutants were tryptophan auxotrophs, with only very small colonies detectable without tryptophan supplementation (molinahenares2009functionalanalysisof pages 4-6).

2.3 Precursor-feeding evidence maps trpF step relative to anthranilate and indole

In minimal medium supplementation assays, the trpF mutant is rescued by tryptophan and by indole, but shows only weak/variable growth with anthranilate or chorismate supplementation, a pattern consistent with trpF acting downstream of anthranilate formation and upstream of indole (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media f6402c94).

These assays were performed in M9 minimal medium with 16 mM citrate as carbon source and supplemented nutrients at 0.2 mM (with a higher tryptophan condition used for a trpA comparison), providing quantitative experimental conditions for the phenotype (molinahenares2009functionalanalysisof pages 4-6).

3) Structure/function and substrate specificity (TrpF family; relevance to Q88LE0)

3.1 TIM-barrel enzyme family and loop dynamics as determinants of specificity

Recent mechanistic work on (βα)8-barrel enzymes involved in histidine and tryptophan biosynthesis highlights that TrpF is a specialist enzyme with specificity for the smaller substrate PRA, whereas related enzymes (e.g., PriA) can be bifunctional. A key determinant is active-site architecture and loop dynamics:
* TrpF catalysis includes ribose-ring opening as a first mechanistic step and proceeds to products such as CdRP (romerorivera2022complexloopdynamics pages 3-4).
* TrpF active sites can be more compact, supporting selectivity for PRA over bulkier analogs like ProFAR; substrate volume comparisons (ProFAR 829 ų vs PRA 559 ų) were used to rationalize specificity in a TrpF model (romerorivera2022complexloopdynamics pages 3-4).
* Catalytically important residues and loops are emphasized, including a catalytically important Asp in loop 6 (example numbering D126 in TmTrpF) and multiple long-loop rearrangements that gate catalytic competence (romerorivera2022complexloopdynamics pages 2-3, romerorivera2022complexloopdynamics pages 6-8).

While these detailed mechanistic/structural data are not from P. putida KT2440 specifically, they provide authoritative, modern support for how TrpF-family enzymes (the family to which Q88LE0 belongs, per the provided UniProt context) achieve substrate specificity and catalysis (romerorivera2022complexloopdynamics pages 2-3, romerorivera2022complexloopdynamics pages 3-4, romerorivera2022complexloopdynamics pages 6-8).

3.2 What is (and is not) known specifically for KT2440 TrpF biochemistry

In the retrieved KT2440-focused experimental genetics paper, the evidence is phenotype-based (auxotrophy/complementation), not purified-enzyme kinetics. No KT2440 TrpF kinetic parameters (kcat/Km for PRA) were identified in the retrieved corpus, so substrate specificity is supported indirectly by (i) pathway-consistent rescue by indole/tryptophan and (ii) broader TrpF-family mechanistic literature (molinahenares2009functionalanalysisof pages 4-6, romerorivera2022complexloopdynamics pages 3-4).

4) Cellular localization

No retrieved source explicitly states subcellular localization for P. putida KT2440 TrpF. Given the pathway role in amino-acid biosynthesis and lack of secretion/transport context in the experimental genetics paper, the most defensible statement from this evidence set is that TrpF functions as an intracellular enzyme, but the precise compartmental assignment (e.g., “cytosolic”) cannot be directly cited here and should be treated as an inference rather than a proven KT2440-specific fact (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4).

5) Recent developments and latest research (prioritizing 2023–2024)

5.1 2024 perspective: Trp pathway enzymes as antimicrobial targets

A 2024 review on amino-acid biosynthesis inhibitors in tuberculosis drug discovery summarizes the tryptophan pathway and explicitly describes the TrpF step: TrpF opens the ribose ring of PRA to form CdRP, which is then processed by TrpC (guida2024aminoacidbiosynthesis pages 4-6). While focused on Mycobacterium tuberculosis, this review reflects contemporary interest in Trp biosynthesis (including TrpF) as a drug target and provides an up-to-date pathway description (guida2024aminoacidbiosynthesis pages 4-6).

5.2 Current mechanistic consensus: conformational dynamics and evolvability

Modern mechanistic analyses emphasize that TrpF-family catalysis depends on interdependent loop motions and that differences in loop dynamics/active-site volumes can encode specialization versus bifunctionality (romerorivera2022complexloopdynamics pages 2-3, romerorivera2022complexloopdynamics pages 3-4). This informs current thinking on how TrpF-like enzymes might be engineered or inhibited.

6) Current applications and real-world implementations (with quantitative data)

6.1 Metabolic engineering in P. putida KT2440: anthranilate production

Anthranilate (o-aminobenzoate; a tryptophan-pathway intermediate precursor) is an industrially relevant aromatic. In P. putida KT2440, Kuepper et al. (2015) engineered strains to accumulate anthranilate from glucose by deleting trpDC (blocking conversion toward tryptophan) and overexpressing feedback-insensitive variants (including trpES40FG and aroGD146N) (kuepper2015metabolicengineeringof pages 1-2). In tryptophan-limited fed-batch fermentations, the best strain reached a maximum of 1.54 ± 0.3 g/L anthranilate (11.23 mM) (kuepper2015metabolicengineeringof pages 1-2).

Although this study targets upstream steps (anthranilate synthesis and blocking its consumption), it is a concrete example of how the tryptophan biosynthesis branch—including downstream steps like TrpF—interfaces with engineering strategies that modulate flux through the pathway (kuepper2015metabolicengineeringof pages 1-2).

6.2 Biosensing built from P. putida KT2440 regulation: indole-responsive biosensor

Matulis et al. (2022) repurposed an indole-inducible gene expression system from P. putida KT2440 (PpTrpI/PPP_RS00425) to build whole-cell biosensors in E. coli and Cupriavidus necator. Key reported performance metrics include:
* Up to 639.6-fold induction by indole (in E. coli biosensor, minimal medium) (matulis2022developmentandcharacterization pages 1-2).
* A linear response range of approximately 0.4–5 mM indole (matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6).
* Apparent Km values ~0.9–1.8 mM, depending on host and medium (matulis2022developmentandcharacterization pages 4-6).
* Specificity: structurally similar compounds (including L-tryptophan and several indole-acid derivatives) did not induce the system (matulis2022developmentandcharacterization pages 6-8, matulis2022developmentandcharacterization pages 4-6).

This is a real-world implementation derived from KT2440’s tryptophan/indole regulatory biology, enabling monitoring of indole production/accumulation—a pathway output connected to tryptophan metabolism (matulis2022developmentandcharacterization pages 6-8).

7) Expert interpretation and analysis (evidence-based)

  1. Strongest KT2440-specific evidence for function is genetic/phenotypic. The targeted PP_1995 knockout yields tryptophan auxotrophy and is rescued by indole/tryptophan, establishing that trpF is required for de novo tryptophan synthesis and placing it upstream of indole (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media f6402c94).
  2. Modern mechanistic literature suggests TrpF specificity is a structural/dynamic property rather than just a static active-site motif. Comparative MD/EVB work highlights loop rearrangements and active-site volume as determinants of PRA specificity and as levers of evolvability toward bifunctional PriA-like behavior (romerorivera2022complexloopdynamics pages 3-4, romerorivera2022complexloopdynamics pages 6-8).
  3. Applied biotechnology around this pathway in P. putida focuses more on flux control than on TrpF itself. Industrially oriented studies often modulate upstream anthranilate synthesis and block consumption (e.g., trpDC deletion) to accumulate valuable aromatics; TrpF becomes relevant as part of the native pathway capacity that may need to be bypassed, balanced, or monitored (kuepper2015metabolicengineeringof pages 1-2).

8) Key statistics/data points (from cited studies)

  • KT2440 trpF (PP_1995) genomic context: PP1994 is 61 nt upstream; PP1996 starts 222 nt downstream of PP_1995 stop codon; RT-PCR does not support cotranscription → likely monocistronic (molinahenares2009functionalanalysisof pages 2-4).
  • KT2440 trpF mutant feeding assay conditions: M9 minimal medium + 16 mM citrate, supplements at 0.2 mM; TrpF growth pattern indicates rescue by tryptophan and indole (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media f6402c94).
  • Anthranilate bioproduction in KT2440 (2015): maximum 1.54 ± 0.3 g/L (11.23 mM) in tryptophan-limited fed-batch (kuepper2015metabolicengineeringof pages 1-2).
  • Indole biosensor derived from KT2440 regulation (2022): up to 639.6-fold induction; linear range ~0.4–5 mM; Km ~0.9–1.8 mM (matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6).

9) Evidence summary table

Target Enzyme name / function EC number Reaction step in tryptophan biosynthesis Gene organization / genomic context Key experimental evidence in P. putida KT2440 Growth complementation / quantitative conditions Key citation(s)
trpF (UniProt Q88LE0; locus PP_1995) Phosphoribosyl anthranilate isomerase / N-(5'-phosphoribosyl)anthranilate isomerase; enzyme assigned to the tryptophan branch from anthranilate toward indole-3-glycerol phosphate (molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2) EC 5.3.1.24 reported for TrpF/PRAI in general pathway annotations and genome annotations cited in retrieved literature; KT2440 paper functionally assigns PP_1995 as phosphoribosyl anthranilate isomerase (molinahenares2009functionalanalysisof pages 2-4) Catalyzes the PRAI step after anthranilate phosphoribosyltransferase (TrpD) and before indole-3-glycerol phosphate synthase (TrpC); thus it converts the phosphoribosyl-anthranilate intermediate within the anthranilate → indole-3-glycerol phosphate segment of the pathway (molinahenares2009functionalanalysisof pages 2-4, matulis2022developmentandcharacterization pages 1-2) Monocistronic, unlinked to the other main trp clusters; trp genes in KT2440 are distributed in separate regions, with trpBA and trpGDC in operons while trpF is a single transcriptional unit. In sequenced Pseudomonas spp., trpF is consistently flanked by truA and accD. In KT2440, the upstream ORF (PP1994) is 61 nt away and the downstream ORF (PP1996) starts 222 nt after the PP_1995 stop codon; RT-PCR for cotranscription was negative, supporting that trpF is most probably a single cistron (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2) Because spontaneous mutants were not recovered, the authors made a targeted chromosomal knockout using a pCHESI/pCHESIWKm site-specific homologous inactivation strategy. An internal fragment of about 500 bp was amplified with primers TrpF-XbaI and TrpF-2, cloned, introduced by electroporation, and confirmed by colony PCR and Southern blotting. The resulting trpF-deficient clones were tryptophan auxotrophs, with only very small colonies appearing without tryptophan supplementation (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 7-8) In feeding assays on M9 minimal medium with 16 mM citrate as carbon source, supplements were added to 0.2 mM final concentration (except tryptophan at 0.6 mM for the trpA assay comparison). The TrpF mutant showed growth pattern: M9 +/-, + tryptophan: +, + chorismate: +/-, + anthranilate: +/-, + indole: +. This indicates rescue by tryptophan or indole, but not effective rescue by anthranilate, consistent with TrpF acting downstream of anthranilate formation and upstream of indole production (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof media f6402c94) Molina-Henares et al., 2009-12, Microbial Biotechnology, DOI: 10.1111/j.1751-7915.2008.00062.x, URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2, molinahenares2009functionalanalysisof pages 7-8, molinahenares2009functionalanalysisof media f6402c94); pathway/regulatory context from Matulis et al., 2022-04, Int. J. Mol. Sci., DOI: 10.3390/ijms23094649, URL: https://doi.org/10.3390/ijms23094649 (matulis2022developmentandcharacterization pages 1-2)

Table: This table summarizes the validated functional annotation of Pseudomonas putida KT2440 trpF/PP_1995, including its enzymatic role, pathway position, genomic organization, and direct mutant evidence. It is useful as a concise evidence map linking the locus to tryptophan biosynthesis and experimentally observed auxotrophy/complementation.

References (URLs and publication dates)

  • Molina-Henares MA et al. Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440. Microbial Biotechnology. Dec 2009. https://doi.org/10.1111/j.1751-7915.2008.00062.x (molinahenares2009functionalanalysisof pages 4-6, molinahenares2009functionalanalysisof pages 2-4, molinahenares2009functionalanalysisof pages 1-2, molinahenares2009functionalanalysisof pages 7-8, molinahenares2009functionalanalysisof media f6402c94)
  • Guida M et al. Amino Acid Biosynthesis Inhibitors in Tuberculosis Drug Discovery. Pharmaceutics. May 2024. https://doi.org/10.3390/pharmaceutics16060725 (guida2024aminoacidbiosynthesis pages 4-6)
  • Romero-Rivera A et al. Complex Loop Dynamics Underpin Activity, Specificity, and Evolvability in the (βα)8 Barrel Enzymes of Histidine and Tryptophan Biosynthesis. JACS Au. Apr 2022. https://doi.org/10.1021/jacsau.2c00063 (romerorivera2022complexloopdynamics pages 2-3, romerorivera2022complexloopdynamics pages 3-4, romerorivera2022complexloopdynamics pages 6-8)
  • Kuepper J et al. Metabolic Engineering of Pseudomonas putida KT2440 to Produce Anthranilate from Glucose. Frontiers in Microbiology. Nov 2015. https://doi.org/10.3389/fmicb.2015.01310 (kuepper2015metabolicengineeringof pages 1-2, kuepper2015metabolicengineeringof pages 5-6)
  • Matulis P et al. Development and Characterization of Indole-Responsive Whole-Cell Biosensor Based on the Inducible Gene Expression System from Pseudomonas putida KT2440. Int. J. Mol. Sci. Apr 2022. https://doi.org/10.3390/ijms23094649 (matulis2022developmentandcharacterization pages 6-8, matulis2022developmentandcharacterization pages 1-2, matulis2022developmentandcharacterization pages 4-6)

References

  1. (molinahenares2009functionalanalysisof pages 4-6): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  2. (molinahenares2009functionalanalysisof pages 7-8): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  3. (molinahenares2009functionalanalysisof media f6402c94): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  4. (guida2024aminoacidbiosynthesis pages 4-6): Michela Guida, Chiara Tammaro, Miriana Quaranta, Benedetta Salvucci, Mariangela Biava, Giovanna Poce, and Sara Consalvi. Amino acid biosynthesis inhibitors in tuberculosis drug discovery. Pharmaceutics, 16:725, May 2024. URL: https://doi.org/10.3390/pharmaceutics16060725, doi:10.3390/pharmaceutics16060725. This article has 3 citations.

  5. (romerorivera2022complexloopdynamics pages 2-3): Adrian Romero-Rivera, Marina Corbella, Antonietta Parracino, Wayne M. Patrick, and Shina Caroline Lynn Kamerlin. Complex loop dynamics underpin activity, specificity, and evolvability in the (βα)8 barrel enzymes of histidine and tryptophan biosynthesis. JACS Au, 2:943-960, Apr 2022. URL: https://doi.org/10.1021/jacsau.2c00063, doi:10.1021/jacsau.2c00063. This article has 35 citations and is from a peer-reviewed journal.

  6. (romerorivera2022complexloopdynamics pages 3-4): Adrian Romero-Rivera, Marina Corbella, Antonietta Parracino, Wayne M. Patrick, and Shina Caroline Lynn Kamerlin. Complex loop dynamics underpin activity, specificity, and evolvability in the (βα)8 barrel enzymes of histidine and tryptophan biosynthesis. JACS Au, 2:943-960, Apr 2022. URL: https://doi.org/10.1021/jacsau.2c00063, doi:10.1021/jacsau.2c00063. This article has 35 citations and is from a peer-reviewed journal.

  7. (molinahenares2009functionalanalysisof pages 1-2): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  8. (matulis2022developmentandcharacterization pages 1-2): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  9. (molinahenares2009functionalanalysisof pages 2-4): M. A. Molina-Henares, Adela García‐Salamanca, A. Molina-Henares, J. de la Torre, M. C. Herrera, J. Ramos, and E. Duque. Functional analysis of aromatic biosynthetic pathways in pseudomonas putida kt2440. Microbial biotechnology, 2:91-100, Dec 2009. URL: https://doi.org/10.1111/j.1751-7915.2008.00062.x, doi:10.1111/j.1751-7915.2008.00062.x. This article has 32 citations and is from a peer-reviewed journal.

  10. (romerorivera2022complexloopdynamics pages 6-8): Adrian Romero-Rivera, Marina Corbella, Antonietta Parracino, Wayne M. Patrick, and Shina Caroline Lynn Kamerlin. Complex loop dynamics underpin activity, specificity, and evolvability in the (βα)8 barrel enzymes of histidine and tryptophan biosynthesis. JACS Au, 2:943-960, Apr 2022. URL: https://doi.org/10.1021/jacsau.2c00063, doi:10.1021/jacsau.2c00063. This article has 35 citations and is from a peer-reviewed journal.

  11. (kuepper2015metabolicengineeringof pages 1-2): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

  12. (matulis2022developmentandcharacterization pages 4-6): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  13. (matulis2022developmentandcharacterization pages 6-8): Paulius Matulis, Ingrida Kutraite, Ernesta Augustiniene, Egle Valanciene, Ilona Jonuskiene, and Naglis Malys. Development and characterization of indole-responsive whole-cell biosensor based on the inducible gene expression system from pseudomonas putida kt2440. International Journal of Molecular Sciences, 23:4649, Apr 2022. URL: https://doi.org/10.3390/ijms23094649, doi:10.3390/ijms23094649. This article has 8 citations.

  14. (kuepper2015metabolicengineeringof pages 5-6): Jannis Kuepper, Jasmin Dickler, Michael Biggel, Swantje Behnken, Gernot Jäger, Nick Wierckx, and Lars M. Blank. Metabolic engineering of pseudomonas putida kt2440 to produce anthranilate from glucose. Frontiers in Microbiology, Nov 2015. URL: https://doi.org/10.3389/fmicb.2015.01310, doi:10.3389/fmicb.2015.01310. This article has 66 citations and is from a peer-reviewed journal.

Artifacts

Citations

  1. guida2024aminoacidbiosynthesis pages 4-6
  2. molinahenares2009functionalanalysisof pages 1-2
  3. matulis2022developmentandcharacterization pages 1-2
  4. molinahenares2009functionalanalysisof pages 2-4
  5. molinahenares2009functionalanalysisof pages 4-6
  6. molinahenares2009functionalanalysisof pages 7-8
  7. romerorivera2022complexloopdynamics pages 3-4
  8. kuepper2015metabolicengineeringof pages 1-2
  9. matulis2022developmentandcharacterization pages 4-6
  10. matulis2022developmentandcharacterization pages 6-8
  11. romerorivera2022complexloopdynamics pages 2-3
  12. romerorivera2022complexloopdynamics pages 6-8
  13. kuepper2015metabolicengineeringof pages 5-6
  14. https://doi.org/10.1111/j.1751-7915.2008.00062.x**
  15. https://doi.org/10.3390/ijms23094649**
  16. https://doi.org/10.1111/j.1751-7915.2008.00062.x
  17. https://doi.org/10.3390/pharmaceutics16060725
  18. https://doi.org/10.1021/jacsau.2c00063
  19. https://doi.org/10.3389/fmicb.2015.01310
  20. https://doi.org/10.3390/ijms23094649
  21. https://doi.org/10.1111/j.1751-7915.2008.00062.x,
  22. https://doi.org/10.3390/pharmaceutics16060725,
  23. https://doi.org/10.1021/jacsau.2c00063,
  24. https://doi.org/10.3390/ijms23094649,
  25. https://doi.org/10.3389/fmicb.2015.01310,

📄 View Raw YAML

id: Q88LE0
gene_symbol: trpF
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: trpF encodes N-(5'-phosphoribosyl)anthranilate isomerase (PRAI, EC 5.3.1.24), a cytoplasmic monomeric TIM-barrel ((beta/alpha)8) enzyme that catalyzes the third step of L-tryptophan biosynthesis from chorismate. It converts N-(5-phospho-beta-D-ribosyl)anthranilate (PRA) into 1-(2-carboxyphenylamino)-1-deoxy-D-ribulose 5-phosphate (CdRP) via an Amadori rearrangement that opens the ribose ring. This intermediate is subsequently used by TrpC (indole-3-glycerol-phosphate synthase) and tryptophan synthase (TrpAB) to complete tryptophan biosynthesis. In Pseudomonas putida KT2440 the gene (locus PP_1995) is unlinked to the other trp clusters and is most likely monocistronic; a targeted chromosomal knockout produces a tryptophan auxotroph that is rescued by tryptophan or indole but not by anthranilate, placing the enzyme downstream of anthranilate and upstream of indole as expected for PRAI.
existing_annotations:
- term:
    id: GO:0000162
    label: L-tryptophan biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: Correct and core. PRAI catalyzes step 3 of 5 in tryptophan biosynthesis from chorismate, and the KT2440 trpF knockout is a tryptophan auxotroph, directly confirming a required role in this process.
    action: ACCEPT
    reason: The biological process is supported both by family/pathway assignment (UniPathway UPA00035) and by experimental auxotrophy/complementation evidence in KT2440 (PMID:21261884).
- term:
    id: GO:0004640
    label: phosphoribosylanthranilate isomerase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: Correct core molecular function. This is the precise EC 5.3.1.24 activity (RHEA:21540) defining the TrpF family, consistent with the HAMAP rule, InterPro TrpF family (IPR044643) and PRAI domain (IPR001240), and the conserved TIM-barrel active site.
    action: ACCEPT
    reason: Directly matches the curated catalytic activity in UniProt and the protein's family/domain assignment; this is the defining function of the gene product.
- term:
    id: GO:0046394
    label: carboxylic acid biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000117
  qualifier: involved_in
  review:
    summary: Not wrong but uninformative and over-general. Tryptophan is a carboxylic acid, so this high-level ARBA-derived term is technically true but adds nothing beyond the more specific GO:0000162 (L-tryptophan biosynthetic process), which is already annotated.
    action: MARK_AS_OVER_ANNOTATED
    reason: Generic parent term auto-generated from sequence features; the specific L-tryptophan biosynthetic process annotation already captures the biology precisely.
core_functions:
- description: Catalyzes the isomerization of N-(5-phospho-beta-D-ribosyl)anthranilate to 1-(2-carboxyphenylamino)-1-deoxy-D-ribulose 5-phosphate (CdRP), the third step of L-tryptophan biosynthesis.
  molecular_function:
    id: GO:0004640
    label: phosphoribosylanthranilate isomerase activity
  supported_by:
  - reference_id: PMID:21261884
    full_text_unavailable: true
    supporting_text: Targeted chromosomal knockout of trpF (PP_1995) in KT2440 yields a tryptophan auxotroph rescued by tryptophan or indole but not anthranilate, consistent with loss of PRAI activity acting downstream of anthranilate and upstream of indole.
  directly_involved_in:
  - id: GO:0000162
    label: L-tryptophan biosynthetic process
references:
- id: GO_REF:0000117
  title: Electronic Gene Ontology annotations created by ARBA machine learning models
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:21261884
  title: Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440
  findings:
  - statement: trpF (PP_1995) is unlinked to the other trp gene clusters and most likely monocistronic; its knockout is a tryptophan auxotroph rescued by tryptophan or indole but not by anthranilate, confirming its role as PRAI in tryptophan biosynthesis.
    reference_section_type: RESULTS
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: PMID confirmed via PubMed search for the exact title (Molina-Henares et al., Microb Biotechnol 2009). Provides the experimental auxotrophy/complementation evidence underpinning the function annotations.