aroH

UniProt ID: Q88LR3
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

aroH (PP_1866) encodes a class-II (type-II) 3-deoxy-D-arabino-heptulosonate 7-phosphate synthase (DAHP synthase / DAH7PS; EC 2.5.1.54) in Pseudomonas putida KT2440. The enzyme catalyzes the aldol-like condensation of phosphoenolpyruvate (PEP) and D-erythrose 4-phosphate (E4P), with release of phosphate, to form 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP). This is the first committed step of the shikimate pathway, which ultimately produces chorismate, the branch-point precursor for the aromatic amino acids (phenylalanine, tyrosine, tryptophan) and many other aromatic metabolites. Class-II DAH7PS enzymes adopt a (beta/alpha)8 TIM-barrel fold and require a divalent metal cation for catalysis; the UniProt record for this protein indicates activity with manganese, cobalt or cadmium ions, with one cation bound per subunit. In bacterial DAHP-synthase nomenclature, AroH is conventionally the tryptophan-sensitive isoenzyme, although the regulatory properties of this specific P. putida protein have not been characterized experimentally. As a soluble central-metabolism enzyme acting on cytosolic substrates, AroH is expected to function in the cytoplasm.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0003849 3-deoxy-7-phosphoheptulonate synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: Core molecular function. The protein belongs to the class-II DAHP synthase family (RuleBase RU363071; InterPro IPR002480/PF01474; TIGR01358 DAHP_synth_II) and the UniProt catalytic-activity record (Rhea:14717, EC 2.5.1.54) describes condensation of PEP + E4P + H2O to DAHP + phosphate, exactly matching this GO term.
Reason: This term precisely captures the enzymatic activity of a class-II DAHP synthase. The assignment is strongly supported by sequence/family evidence (InterPro, RuleBase, NCBIfam TIGR01358) and conserved PEP- and metal-binding residues annotated in UniProt. This is the gene's core molecular function.
Supporting Evidence:
file:PSEPK/aroH/aroH-deep-research-falcon.md
AroH is a DAH7PS catalyzing PEP + E4P -> DAHP/DAH7P and feeding the shikimate pathway; best-supported assignment is class-II / type-II DAH7PS (TIM-barrel fold, conserved metal-binding site), compatible with the UniProt/InterPro/Pfam context for Q88LR3 (PF01474 / DAHP_synth_2). No direct biochemical characterization of P. putida KT2440 AroH was found, so this is family/sequence-level inference.
GO:0009073 aromatic amino acid family biosynthetic process
IEA
GO_REF:0000002
ACCEPT
Summary: Correct biological process. DAHP synthase catalyzes the first committed step of the shikimate pathway, which feeds chorismate biosynthesis and the downstream aromatic amino acid (Phe/Tyr/Trp) biosynthetic branches.
Reason: This is the canonical pathway role of a DAHP synthase and is consistent with the molecular function annotation. The term is appropriate and represents a core biological process for this gene. (The GOA stub used the label "aromatic amino acid biosynthetic process"; the current GO label for GO:0009073 is "aromatic amino acid family biosynthetic process" and is used here.)

Core Functions

Catalyzes the first committed step of the shikimate pathway, the condensation of phosphoenolpyruvate and D-erythrose 4-phosphate to form 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP), as a class-II DAHP synthase.

Supporting Evidence:
  • GO_REF:0000120
    3-deoxy-7-phosphoheptulonate synthase activity (GO:0003849); EC 2.5.1.54; Rhea:14717 PEP + E4P + H2O = DAHP + phosphate.

References

Gene Ontology annotation through association of InterPro records with GO terms
Combined Automated Annotation using Multiple IEA Methods
Microbial origin of plant-type 2-keto-3-deoxy-D-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from Xanthomonas campestris
  • Biochemical characterization of a class-II (plant-type / AroAII) DAHP synthase shows divalent-metal dependence (activity abolished by EDTA) and feedback inhibition by chorismate and tryptophan, providing the family-level basis for annotating P. putida AroH as a class-II DAHP synthase.
    "Microbial origin of plant-type 2-keto-3-deoxy-D-arabino-heptulosonate 7-phosphate synthases from Xanthomonas campestris; Km(PEP) and Km(E4P) reported; chorismate and L-tryptophan inhibition characterized."

Deep Research

Asta

(aroH-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:21:18.788908

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.802)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.745)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[3] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.733)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[4] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.718)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[5] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.712)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[6] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.710)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[7] Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser

  • Authors: Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała
  • Year: 2025
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  • DOI: 10.1186/s12864-025-12329-8
  • PMID: 41408139
  • PMCID: 12822150
  • Citations: 2
  • Summary: An R script is developed that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database and provides a single list of updated symbols with annotation about their ambiguity.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > Gene nomenclature contains current official symbols and various numbers of synonyms, which pose a challenge to integrating genomic data and increase the probability that different genes share the same symbol. Therefore, we retrieved identifiers assigned to all protein-coding genes in human, mouse, rat and pig genomes that are available in the Ensembl genome browser (release 113) to assess the number of genes, compare species and identify ambiguous symbols. Results: Our analysis revealed that the total number of symbols, both official symbols and synonyms, used to identify protein-coding genes ranges from 16,600 in pigs to 64,580 in mice. Furthermore, the gene nomenclature is not complete because there are also genes without an assigned symbol, which indicates gaps in understanding protein-coding genes, especially in pigs. We also found a large number of gene symbols that map to more than one gene. These symbols might complicate the identification of about 10% of rat and mouse genes and 18% of human protein-coding genes. A simple solution for this problem is the usage of stable gene IDs assigned by scientific institutions and committees (Ensembl, NCBI, RGD, HGNC and VGNC) provided that the genomic information associated with these IDs is retrieved directly from proprietary databases containing the most accurate data. Finally, although gene symbols may pose a problem with unequivocal identification of genes, there are instances when no other identifiers are available in the literature. Therefore, we have developed an R script performing search of the Ensembl database and integrating data to provide a single list of updated symbols with annotation about their ambiguity. Conclusions: Gene symbols are not always reliable and should be reported together with stable IDs to enable unequivocal identification of genes. Therefore, data containing only gene symbols should be used cautiously to avoid misidentification of genes. A solution for this problem is our R script REgeness that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database.

[8] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.700)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[9] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.697)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[10] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.696)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[11] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.696)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[12] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.693)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[13] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.680)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[14] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.679)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[15] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[16] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.661)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[17] A Genome-Wide Association Study Identifying Novel Genetic Markers of Response to Treatment with Interleukin-23 Inhibitors in Psoriasis

  • Authors: Sophia Zachari, K. Liadaki, Angeliki Planaki, E. Zafiriou, Olga Kouvarou et al.
  • Year: 2025
  • Venue: Genes
  • URL: https://www.semanticscholar.org/paper/d5f656311b54e222e7487ea32a061869b30178a1
  • DOI: 10.3390/genes16101195
  • PMID: 41153410
  • PMCID: 12564705
  • Summary: These findings provide promising pharmacogenetic markers which, upon validation in larger, independent cohorts, will enable the translation of a patient’s genotype into a response phenotype, thereby guiding clinical decisions and improving drug effectiveness.
  • Evidence snippets:
  • Snippet 1 (score: 0.661)
    > The UniProt knowledgebase (www.uniprot.org/uniprotkb/), (accessed on 20 June 2025), the central hub for the collection of functional information on proteins, with accurate and rich annotation [33], was used to retrieve the approved human gene and protein names and symbols.

[18] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.659)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[19] MultiLoc2: integrating phylogeny and Gene Ontology terms improves subcellular protein localization prediction

  • Authors: Torsten Blum, S. Briesemeister, O. Kohlbacher
  • Year: 2009
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/c2f00f9a94fe72eeeadc54a37a816731f329bfa4
  • DOI: 10.1186/1471-2105-10-274
  • PMID: 19723330
  • PMCID: 2745392
  • Citations: 294
  • Influential citations: 36
  • Summary: MultiLoc2 is an extensive high-performance subcellular protein localization prediction system that outperforms other prediction systems in two benchmarks studies and yields higher accuracies compared to its previous version.
  • Evidence snippets:
  • Snippet 1 (score: 0.659)
    > The Gene Ontology (GO) is a controlled vocabulary for uniformly describing gene products in terms of biological processes, cellular components and molecular function across all organisms [46]. It has been shown that GO terms can be used to improve the performance of subcellular protein localization prediction methods [47,48]. In the literature to date, there are three possibilities for obtaining GO annotation terms for a query sequence. If the UniProt [49] accession number is known, one can simply extract the GO annotation from the UniProt database [50]. However, this procedure fails for novel proteins without accession number. Another possibility is to search for homologous proteins annotated with GO terms using BLAST [28,29]. This becomes difficult in cases where proteins have no close homolog or proteins have many homologs, because no GO term can be obtained or GO terms might be ambiguous. A further method of inferring GO terms is InterProScan [51] used, for example, by Chou and Cai [52]. Given a protein sequence, the tool scans against various pattern and signature data sources collected by the InterPro project [53]. InterPro also provides a mapping of the detected protein domains and functional sites to GO terms.
    > Our subpredictor GOLoc is based on GO terms calculated using InterProScan. Since the GO terms are derived directly from the query sequence, we avoid the drawbacks of using accession numbers or BLAST. The input of GOLoc is a binary-coded vector which represents all GO terms of the training sequences (see Fig. 2). GO terms present in the query sequence are set to 1 in the vector and to 0 otherwise [see Additional file 1].

[20] The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize

  • Authors: Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al.
  • Year: 2016
  • Venue: BMC Systems Biology
  • URL: https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  • DOI: 10.1186/s12918-016-0369-x
  • PMID: 27899149
  • PMCID: 5129634
  • Citations: 11
  • Influential citations: 1
  • Summary: CornCyc’s computational predictions are more accurate than those in MaizeCyc when compared to experimentally determined function assignments, demonstrating the relative strength of the enzymatic function assignment pipeline used to generate CornCyc.
  • Evidence snippets:
  • Snippet 1 (score: 0.657)
    > A gold standard set of protein functional annotations was generated by extracting data from UniProt [16] and BRENDA [17]. We extracted all protein sequence and annotation data from UniProt (release 2016_05) for the organism Zea mays, keeping the EC annotations only from the manually reviewed component of UniProt, while removing those annotations that had not undergone manual review. We also extracted experimentally verified protein annotations for Zea mays from BRENDA (release 2016.1). The UniProt and BRENDA annotations were then merged by matching proteins based on the database crosslinks provided by BRENDA, resulting in the union of the reviewed annotations from UniProt and the experimentally verified annotations of BRENDA with duplicates removed. The merged protein annotations were then matched to the B73 RefGen_v2 translated gene models using BLASTP based on a sequence identity cutoff of 96% and an e-value cutoff of 1e-20. We selected the top scoring hit for each protein which resulted in matches to 1,815 unique maize proteins. EC annotations for alternate isoforms were consolidated at the gene level, resulting in 1,475 experimentally verified or manually reviewed protein functional annotations across 1,450 maize genes.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  3. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  4. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  5. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  6. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  7. Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała (2025). Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser. BMC Genomics. https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  8. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  9. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  10. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  11. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  12. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  13. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  14. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  15. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  16. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  17. Sophia Zachari, K. Liadaki, Angeliki Planaki, E. Zafiriou, Olga Kouvarou et al. (2025). A Genome-Wide Association Study Identifying Novel Genetic Markers of Response to Treatment with Interleukin-23 Inhibitors in Psoriasis. Genes. https://www.semanticscholar.org/paper/d5f656311b54e222e7487ea32a061869b30178a1
  18. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  19. Torsten Blum, S. Briesemeister, O. Kohlbacher (2009). MultiLoc2: integrating phylogeny and Gene Ontology terms improves subcellular protein localization prediction. BMC Bioinformatics. https://www.semanticscholar.org/paper/c2f00f9a94fe72eeeadc54a37a816731f329bfa4
  20. Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al. (2016). The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize. BMC Systems Biology. https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8

Falcon

(aroH-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 23 citations 2 artifacts 2026-06-11T20:59:50.592002

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: Pseudomonas putida KT2440 aroH (PP_1866; UniProt Q88LR3)

Executive summary

The target protein (UniProt Q88LR3; locus PP_1866; gene symbol aroH) is annotated as a class-II (type-II) 3-deoxy-D-arabino-heptulosonate 7-phosphate synthase (DAH7PS; DAHP synthase; EC 2.5.1.54), catalyzing the first committed step into the shikimate pathway by condensing phosphoenolpyruvate (PEP) and erythrose-4-phosphate (E4P) to produce DAHP/DAH7P, a precursor to chorismate and ultimately aromatic amino acids and many aromatic metabolites. Across bacteria, AroH is typically the tryptophan-inhibited DAHP synthase isoenzyme, and P. putida KT2440 shikimate-pathway engineering work reiterates this isoenzyme logic. However, direct biochemical/structural characterization of P. putida KT2440 AroH (Q88LR3) itself was not identified in the retrieved literature, so several mechanistic points must be treated as family-level inference supported by characterized type-II DAH7PS enzymes in other bacteria. (wang2022uncoveringtherole pages 10-12, sterritt2018structuralandfunctional pages 1-2, dias2023fromdegraderto pages 4-6)

1) Key concepts and definitions (current understanding)

1.1 DAHP synthase / DAH7PS and the shikimate pathway

DAHP synthase (DAH7PS) catalyzes an aldol-like condensation of PEP + E4P → DAHP (DAH7P), which is widely considered the entry/first committed step of the shikimate pathway toward chorismate and downstream aromatic metabolites. (wang2022uncoveringtherole pages 10-12, gosset2001microbialoriginof pages 1-2, sterritt2018structuralandfunctional pages 1-2)

In many bacteria, multiple DAHP synthase isoenzymes exist, and their regulation partitions flux in response to aromatic amino acids; in this framework, AroH is the tryptophan-sensitive isoenzyme. (wang2022uncoveringtherole pages 10-12, dias2023fromdegraderto pages 4-6)

1.2 “Class-II/type-II” DAHP synthase family (relevant to Q88LR3)

DAH7PS enzymes are commonly grouped into type I and type II. Characterized type-II enzymes share a (β/α)8 barrel (TIM-barrel) fold, a conserved divalent metal-binding site, and conserved substrate-binding features. (sterritt2018structuralandfunctional pages 1-2)

Sterritt et al. further emphasize that type-II DAH7PS enzymes may use diverse allosteric architectures; some have extra-barrel structural extensions generating distinct allosteric sites for aromatic amino acids, while others lack these elements and thereby display different regulation and oligomeric assembly. (sterritt2018structuralandfunctional pages 1-2)

2) Target verification and gene/protein identity

2.1 Symbol ambiguity check

The gene symbol aroH can denote different functions in different organisms; within bacterial shikimate-pathway nomenclature, it commonly denotes a DAHP synthase isoenzyme (tryptophan-inhibited). In the retrieved P. putida KT2440 engineering literature, the AroF/AroG/AroH triad is explicitly discussed as DAHP synthases inhibited by Tyr/Phe/Trp, respectively—consistent with the UniProt-supplied identity for Q88LR3. (dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4)

2.2 Organism and locus context

No retrieved paper explicitly mentions PP_1866 or UniProt Q88LR3 by accession in the text examined. Therefore, P. putida KT2440-specific conclusions are based on (i) the user-provided UniProt identity and (ii) P. putida KT2440 pathway-engineering studies that describe the DAHP-synthase isoenzyme logic that includes AroH. (dias2023fromdegraderto pages 4-6, yu2016metabolicengineeringof pages 1-3)

3) Molecular function: reaction, substrate specificity, and mechanism

3.1 Primary enzymatic reaction (AroH/DAH7PS)

Across DAHP synthases, the core reaction is the condensation of PEP and E4P to yield DAHP/DAH7P. This reaction is stated directly in Pseudomonas DAHP synthase context (phzC/DAH7PS) and in broader DAH7PS structural analysis. (wang2022uncoveringtherole pages 10-12, sterritt2018structuralandfunctional pages 1-2)

For P. putida KT2440 AroH (Q88LR3), no experimental substrate-range studies were retrieved; thus, the most defensible annotation is that AroH is a canonical DAH7PS using PEP and E4P in central metabolism. (wang2022uncoveringtherole pages 10-12)

3.2 Metal dependence (family-level evidence)

A characterized class-II DAH7PS (Xanthomonas campestris AroAII) shows strong metal dependence: activity is abolished by EDTA and can be restored by divalent metals (e.g., Mn2+ partially restoring activity). (gosset2001microbialoriginof pages 5-7, gosset2001microbialoriginof pages 7-8)

This supports the inference that class-II DAH7PS enzymes—including AroH family members—typically require a divalent metal for catalysis, though the specific metal preference for P. putida AroH is not established here. (sterritt2018structuralandfunctional pages 1-2, gosset2001microbialoriginof pages 5-7)

3.3 Quantitative kinetics and inhibition constants (comparator class-II enzyme)

Because P. putida AroH-specific kinetics were not found, the best available quantitative reference in the retrieved corpus is the class-II AroAII enzyme from X. campestris (Gosset et al., 2001). Reported apparent Michaelis constants were Km(PEP) = 0.13 mM and Km(E4P) = 0.23 mM. (gosset2001microbialoriginof pages 5-7, gosset2001microbialoriginof media 69aabd91)

Regulatory inhibitor constants in this class-II example include:
- Chorismate competitive inhibition: Ki = 0.31 mM (competitive vs PEP) and Ki = 0.09 mM (competitive vs E4P). (gosset2001microbialoriginof pages 5-7, gosset2001microbialoriginof pages 7-8)
- L-tryptophan noncompetitive inhibition: Ki = 0.35 mM (vs PEP) and Ki = 0.61 mM (vs E4P). (gosset2001microbialoriginof pages 5-7, gosset2001microbialoriginof pages 7-8)

These values provide an order-of-magnitude view of inhibitor potency in a class-II DAH7PS, but they should not be treated as parameters for P. putida KT2440 AroH without direct measurement. (gosset2001microbialoriginof pages 5-7)

4) Biological role and pathways in P. putida KT2440

4.1 Role in aromatic amino acid and chorismate-derived metabolism

DAH7PS catalyzes the first committed step into the shikimate pathway, which produces chorismate as a branchpoint precursor for aromatic amino acids and many aromatic metabolites. (gosset2001microbialoriginof pages 1-2, sterritt2018structuralandfunctional pages 1-2)

In Pseudomonas systems, DAH7PS entry can also supply specialized aromatic secondary metabolites (e.g., phenazine/pyocyanin routes), reinforcing that DAH7PS enzymes can couple primary and secondary metabolism depending on paralog context and regulation. (sterritt2018structuralandfunctional pages 1-2, wang2022uncoveringtherole pages 1-3)

4.2 Regulation and isoenzyme logic (AroF/AroG/AroH)

Multiple sources reiterate the common regulatory logic that AroF, AroG, and AroH are feedback inhibited by tyrosine, phenylalanine, and tryptophan, respectively, situating AroH as the Trp-sensitive control point. (wang2022uncoveringtherole pages 10-12, dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4)

Sterritt et al. emphasize that type-II DAH7PS allostery can be implemented through different structural “componentry”; some type-II enzymes have additional structural extensions that create allosteric sites and enable complex feedback regimes, while other type-II enzymes (including a Pseudomonas phenazine-associated enzyme) lack those structural elements and therefore differ in inhibition behavior and oligomerization. This underscores that the existence and mechanism of AroH allostery must be validated experimentally for each organism/enzyme. (sterritt2018structuralandfunctional pages 1-2)

5) Subcellular localization

For a canonical shikimate-pathway DAH7PS such as AroH, the most plausible localization is cytosolic, because its substrates (PEP, E4P) are soluble intermediates of central metabolism. This is a functional inference rather than a direct localization measurement for P. putida AroH in the retrieved corpus. (wang2022uncoveringtherole pages 10-12)

Importantly, a class-II DAH7PS example (X. campestris AroAII) was predicted by PSORT to have a transmembrane region and a topology consistent with a membrane-anchored N-terminus and cytosolic catalytic domain, indicating that membrane association is possible in some class-II DAH7PS lineages. (gosset2001microbialoriginof pages 5-7)

Thus, while cytosolic localization is the best default annotation for P. putida AroH, class-II family diversity warrants checking N-terminal features/topology in the specific sequence when possible. (gosset2001microbialoriginof pages 10-11, gosset2001microbialoriginof pages 5-7)

6) Recent developments and latest research (prioritizing 2023–2024)

6.1 2023: P. putida KT2440 shikimate flux redirection to gallic acid

Dias et al. (published 2023-11; International Microbiology) engineered P. putida KT2440 from a gallic-acid degrader to a producer by introducing a heterologous operon (including a feedback-resistant DAHP synthase variant aroG4 (Pro150Leu)) and deleting native catabolic operons (ΔpcaHG and ΔgalTAPR) to prevent degradation of the target product and intermediate. (dias2023fromdegraderto pages 1-2, dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4)

The study explicitly re-states DAHP synthase isoenzyme inhibition logic: aromatic amino acids inhibit DAHP synthases, with AroH inhibited by tryptophan. (dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4)

6.2 2024: structural/allosteric themes remain central (but KT2440 AroH-specific updates not found)

No 2023–2024 primary papers directly characterizing P. putida KT2440 AroH (Q88LR3) were retrieved. The closest “authoritative mechanistic framework” in the retrieved set remains the structural and regulatory analysis of type-II DAH7PS enzymes (e.g., Sterritt et al. 2018), which emphasizes diversification of allostery and oligomeric assembly among type-II enzymes. (sterritt2018structuralandfunctional pages 1-2)

7) Current applications and real-world implementations (with quantitative data)

Because DAH7PS entry is often rate-controlling and feedback regulated, biotechnological exploitation frequently uses feedback-resistant DAHP synthase variants (often AroG variants) to push carbon into chorismate-derived products in P. putida KT2440. (yu2016metabolicengineeringof pages 1-3, dias2023fromdegraderto pages 4-6)

7.1 para-Hydroxybenzoic acid (PHBA) production in P. putida KT2440 (2016)

Yu et al. (published 2016-11; Frontiers in Bioengineering and Biotechnology) engineered P. putida KT2440 for PHBA production from glucose via chorismate by overexpressing E. coli ubiC (chorismate lyase) and a feedback-resistant DAHP synthase variant aroG D146N; they also deleted pobA (to prevent product degradation) and pheA/trpE (to reduce chorismate drain to aromatic amino acids), and deleted hexR to increase E4P/NADPH availability. (yu2016metabolicengineeringof pages 1-3, yu2016metabolicengineeringof pages 5-6)

Reported performance: maximum 1.73 g/L PHBA and 18.1% carbon yield (C-mol/C-mol) in a non-optimized fed-batch process—an example of direct industrially relevant output linked to shikimate-pathway entry control. (yu2016metabolicengineeringof pages 1-3)

7.2 Gallic acid production in P. putida KT2440 (2023)

Dias et al. reported 346.7 ± 0.004 mg/L gallic acid after 72 h in shaker culture following deletions that blocked degradation plus expression of a synthetic operon (including a feedback-resistant DAHP synthase variant). (dias2023fromdegraderto pages 1-2)

8) Expert opinions and analysis (authoritative sources)

A key expert-level insight from structural enzymology is that type-II DAH7PS enzymes retain a conserved catalytic scaffold (TIM barrel, metal-binding site) while evolving diverse allosteric “modules” and oligomerization strategies. This implies that annotation of AroH as “Trp-inhibited DAH7PS” is reasonable at the pathway level, but the precise molecular mechanism and strength of inhibition can vary substantially by lineage and should not be assumed without measurement. (sterritt2018structuralandfunctional pages 1-2)

Additionally, the class-II AroAII work emphasizes that in some organisms a class-II enzyme can be the sole DAHP synthase supporting primary aromatic amino acid synthesis, and that sequential feedback inhibition by chorismate and tryptophan is possible in class-II enzymes—highlighting why DAH7PS is a frequent metabolic-engineering target. (gosset2001microbialoriginof pages 1-2, gosset2001microbialoriginof pages 5-7)

9) Evidence gaps and confidence assessment (for Q88LR3 specifically)

High confidence (family/pathway level):
- AroH is a DAH7PS catalyzing PEP + E4P → DAHP/DAH7P and feeding the shikimate pathway. (wang2022uncoveringtherole pages 10-12, sterritt2018structuralandfunctional pages 1-2)
- AroH is commonly described as the Trp-inhibited DAHP synthase isoenzyme in bacterial nomenclature, reiterated in P. putida KT2440 context. (dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4)

Moderate confidence (structural/mechanistic inference):
- Type-II DAH7PS enzymes generally use a TIM-barrel fold and a conserved metal-binding site, but exact allostery differs among subtypes. (sterritt2018structuralandfunctional pages 1-2)

Low confidence / not established for KT2440 AroH in retrieved data:
- AroH-specific Km, kcat, Ki, metal preference, oligomeric state, and inhibition mechanism in P. putida KT2440.
- Experimental subcellular localization of Q88LR3.

Summary evidence table

Annotation aspect Evidence summary for aroH / PP_1866 / UniProt Q88LR3 in Pseudomonas putida KT2440 Notes / citations
Gene/protein identity Q88LR3 is annotated as phospho-2-dehydro-3-deoxyheptonate aldolase (EC 2.5.1.54), i.e. a DAHP synthase / DAH7PS family enzyme; ordered locus PP_1866 and gene name aroH match the requested target. Literature on P. putida KT2440 is limited, so much mechanistic detail is inferred from class-II DAHP synthase literature and DAHP-synthase isoenzyme conventions. Functional inference is consistent with DAHP synthase annotations and class-II family discussion; direct P. putida biochemical characterization was not found in retrieved evidence. (wang2022uncoveringtherole pages 10-12, sterritt2018structuralandfunctional pages 1-2)
Enzyme family / structural class Best-supported assignment is class-II / type-II DAH7PS. Characterized type-II DAH7PS enzymes share a (β/α)8 TIM-barrel fold, conserved metal-binding site, and conserved PEP/E4P-binding residues. This is compatible with the UniProt/InterPro/Pfam context for Q88LR3 (PF01474 / DAHP_synth_2). Type-II DAH7PS structural features were described directly for characterized enzymes; application to Q88LR3 is by family inference. (sterritt2018structuralandfunctional pages 1-2, wang2022uncoveringtherole pages 15-16)
Catalytic reaction / substrate specificity DAHP synthases catalyze the condensation of phosphoenolpyruvate (PEP) and erythrose-4-phosphate (E4P) to form DAHP/DAH7P, the first committed step of the shikimate pathway leading to chorismate and aromatic amino acids. No alternative substrate specificity for P. putida AroH was found. This reaction is directly stated for DAHP synthases and class-II examples. (gosset2001microbialoriginof pages 1-2, wang2022uncoveringtherole pages 10-12, sterritt2018structuralandfunctional pages 1-2)
Pathway role in P. putida AroH is inferred to function in entry into the shikimate pathway, supplying flux toward chorismate and downstream biosynthesis of Trp, Phe, Tyr and other chorismate-derived metabolites. In engineering literature, increasing DAHP synthase flux is treated as a major leverage point for aromatic production in P. putida. Rate control at DAHP synthase is emphasized in shikimate-pathway engineering studies. (wang2022uncoveringtherole pages 1-3, yu2016metabolicengineeringof pages 1-3, yu2016metabolicengineeringof pages 3-4)
Feedback regulation General DAHP-synthase isoenzyme convention assigns AroH as the Trp-sensitive isoenzyme; the 2023 P. putida engineering study explicitly reiterates that AroF, AroG, and AroH are inhibited by Tyr, Phe, and Trp, respectively. For a class-II comparator, Xanthomonas campestris AroAII shows Trp noncompetitive inhibition and chorismate competitive inhibition. P. putida KT2440-specific inhibition constants for AroH were not found. Comparator class-II kinetics: Km PEP 0.13 mM; Km E4P 0.23 mM; Ki chorismate 0.31 mM vs PEP and 0.09 mM vs E4P; Ki Trp 0.35 mM vs PEP and 0.61 mM vs E4P. (dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4, gosset2001microbialoriginof pages 5-7, gosset2001microbialoriginof media 69aabd91)
Localization inference The most likely localization for P. putida AroH is cytosolic, because canonical DAHP synthases act on soluble central-metabolism substrates (PEP, E4P) in the shikimate pathway. However, class-II exceptions exist: some bacterial AroAII enzymes, such as the X. campestris example, have a predicted N-terminal membrane/periplasm-associated segment; smaller AroAII proteins lacking that region are inferred to be soluble. Thus, membrane association is a known class-II exception, but no evidence was found that PP_1866/Q88LR3 has such an extension; soluble/cytosolic annotation remains the best inference. (gosset2001microbialoriginof pages 1-2, gosset2001microbialoriginof pages 10-11, gosset2001microbialoriginof pages 5-7)
Real-world implementation in P. putida (2016) In P. putida KT2440, boosting shikimate-pathway entry with a feedback-resistant DAHP synthase aroG D146N plus ubiC, combined with deletion of pobA, pheA, trpE, hexR, enabled para-hydroxybenzoic acid (PHBA) production. Reported best performance: 1.73 g/L PHBA and 18.1% C-mol/C-mol in non-optimized fed-batch. This shows the practical importance of relieving DAHP synthase control, even though the engineered enzyme was AroG rather than native AroH. (yu2016metabolicengineeringof pages 1-3, yu2016metabolicengineeringof pages 5-6, yu2016metabolicengineeringof pages 3-4)
Real-world implementation in P. putida (2023) A 2023 KT2440 gallic-acid study used a heterologous feedback-resistant aroG4 (Pro150Leu) with downstream pathway genes and deletion of pcaHG and galTAPR to redirect shikimate-derived flux. The study again states AroH is the Trp-inhibited DAHP synthase isoenzyme. Final reported production: 346.7 ± 0.004 mg/L gallic acid after 72 h in shaker culture. This further supports DAHP-synthase deregulation as a real-world strategy in P. putida aromatic bioproduction. (dias2023fromdegraderto pages 4-6, dias2023fromdegraderto pages 2-4, dias2023fromdegraderto pages 1-2)

Table: This table summarizes the strongest available evidence for functional annotation of Pseudomonas putida KT2440 aroH (Q88LR3/PP_1866), combining direct P. putida pathway-engineering evidence with mechanistic data from characterized class-II DAHP synthases. It is useful for separating target-specific facts from family-level inference where direct biochemical characterization is limited.

Key references (with URLs and publication dates when available)

  • Dias FMS et al. From degrader to producer: reversing the gallic acid metabolism of Pseudomonas putida KT2440. International Microbiology (published online 2023; issue 2023-11). https://doi.org/10.1007/s10123-022-00282-5 (dias2023fromdegraderto pages 1-2)
  • Yu S et al. Metabolic Engineering of Pseudomonas putida KT2440 for the Production of para-Hydroxy Benzoic Acid. Frontiers in Bioengineering and Biotechnology (2016-11). https://doi.org/10.3389/fbioe.2016.00090 (yu2016metabolicengineeringof pages 1-3)
  • Sterritt OW et al. Structural and functional characterisation… defines a new DAH7PS subclass. Bioscience Reports (2018-09). https://doi.org/10.1042/bsr20181605 (sterritt2018structuralandfunctional pages 1-2)
  • Gosset G et al. Microbial Origin of Plant-Type… chorismate- and tryptophan-regulated enzyme from Xanthomonas campestris. Journal of Bacteriology (2001-07). https://doi.org/10.1128/jb.183.13.4061-4070.2001 (gosset2001microbialoriginof pages 5-7)

References

  1. (wang2022uncoveringtherole pages 10-12): Songwei Wang, Dongliang Liu, Muhammad Bilal, Wei Wang, and Xuehong Zhang. Uncovering the role of phzc as dahp synthase in shikimate pathway of pseudomonas chlororaphis ht66. Biology, 11:86, Jan 2022. URL: https://doi.org/10.3390/biology11010086, doi:10.3390/biology11010086. This article has 12 citations.

  2. (sterritt2018structuralandfunctional pages 1-2): O. W. Sterritt, Eric J. M. Lang, Sarah A Kessans, T. Ryan, B. Demeler, G. Jameson, and E. Parker. Structural and functional characterisation of the entry point to pyocyanin biosynthesis in pseudomonas aeruginosa defines a new 3-deoxy-d-arabino-heptulosonate 7-phosphate synthase subclass. Bioscience Reports, Sep 2018. URL: https://doi.org/10.1042/bsr20181605, doi:10.1042/bsr20181605. This article has 23 citations and is from a peer-reviewed journal.

  3. (dias2023fromdegraderto pages 4-6): Felipe M. S. Dias, Raoní K. Pantoja, José Gregório C. Gomez, and Luiziana F. Silva. From degrader to producer: reversing the gallic acid metabolism of pseudomonas putida kt2440. International Microbiology, 26:243-255, Nov 2023. URL: https://doi.org/10.1007/s10123-022-00282-5, doi:10.1007/s10123-022-00282-5. This article has 7 citations and is from a peer-reviewed journal.

  4. (gosset2001microbialoriginof pages 1-2): Guillermo Gosset, Carol A. Bonner, and Roy A. Jensen. Microbial origin of plant-type 2-keto-3-deoxy-d-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from xanthomonas campestris. Journal of Bacteriology, 183:4061-4070, Jul 2001. URL: https://doi.org/10.1128/jb.183.13.4061-4070.2001, doi:10.1128/jb.183.13.4061-4070.2001. This article has 84 citations and is from a peer-reviewed journal.

  5. (dias2023fromdegraderto pages 2-4): Felipe M. S. Dias, Raoní K. Pantoja, José Gregório C. Gomez, and Luiziana F. Silva. From degrader to producer: reversing the gallic acid metabolism of pseudomonas putida kt2440. International Microbiology, 26:243-255, Nov 2023. URL: https://doi.org/10.1007/s10123-022-00282-5, doi:10.1007/s10123-022-00282-5. This article has 7 citations and is from a peer-reviewed journal.

  6. (yu2016metabolicengineeringof pages 1-3): Shiqin Yu, Manuel R. Plan, Gal Winter, and Jens O. Krömer. Metabolic engineering of pseudomonas putida kt2440 for the production of para-hydroxy benzoic acid. Frontiers in Bioengineering and Biotechnology, Nov 2016. URL: https://doi.org/10.3389/fbioe.2016.00090, doi:10.3389/fbioe.2016.00090. This article has 76 citations.

  7. (gosset2001microbialoriginof pages 5-7): Guillermo Gosset, Carol A. Bonner, and Roy A. Jensen. Microbial origin of plant-type 2-keto-3-deoxy-d-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from xanthomonas campestris. Journal of Bacteriology, 183:4061-4070, Jul 2001. URL: https://doi.org/10.1128/jb.183.13.4061-4070.2001, doi:10.1128/jb.183.13.4061-4070.2001. This article has 84 citations and is from a peer-reviewed journal.

  8. (gosset2001microbialoriginof pages 7-8): Guillermo Gosset, Carol A. Bonner, and Roy A. Jensen. Microbial origin of plant-type 2-keto-3-deoxy-d-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from xanthomonas campestris. Journal of Bacteriology, 183:4061-4070, Jul 2001. URL: https://doi.org/10.1128/jb.183.13.4061-4070.2001, doi:10.1128/jb.183.13.4061-4070.2001. This article has 84 citations and is from a peer-reviewed journal.

  9. (gosset2001microbialoriginof media 69aabd91): Guillermo Gosset, Carol A. Bonner, and Roy A. Jensen. Microbial origin of plant-type 2-keto-3-deoxy-d-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from xanthomonas campestris. Journal of Bacteriology, 183:4061-4070, Jul 2001. URL: https://doi.org/10.1128/jb.183.13.4061-4070.2001, doi:10.1128/jb.183.13.4061-4070.2001. This article has 84 citations and is from a peer-reviewed journal.

  10. (wang2022uncoveringtherole pages 1-3): Songwei Wang, Dongliang Liu, Muhammad Bilal, Wei Wang, and Xuehong Zhang. Uncovering the role of phzc as dahp synthase in shikimate pathway of pseudomonas chlororaphis ht66. Biology, 11:86, Jan 2022. URL: https://doi.org/10.3390/biology11010086, doi:10.3390/biology11010086. This article has 12 citations.

  11. (gosset2001microbialoriginof pages 10-11): Guillermo Gosset, Carol A. Bonner, and Roy A. Jensen. Microbial origin of plant-type 2-keto-3-deoxy-d-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from xanthomonas campestris. Journal of Bacteriology, 183:4061-4070, Jul 2001. URL: https://doi.org/10.1128/jb.183.13.4061-4070.2001, doi:10.1128/jb.183.13.4061-4070.2001. This article has 84 citations and is from a peer-reviewed journal.

  12. (dias2023fromdegraderto pages 1-2): Felipe M. S. Dias, Raoní K. Pantoja, José Gregório C. Gomez, and Luiziana F. Silva. From degrader to producer: reversing the gallic acid metabolism of pseudomonas putida kt2440. International Microbiology, 26:243-255, Nov 2023. URL: https://doi.org/10.1007/s10123-022-00282-5, doi:10.1007/s10123-022-00282-5. This article has 7 citations and is from a peer-reviewed journal.

  13. (yu2016metabolicengineeringof pages 5-6): Shiqin Yu, Manuel R. Plan, Gal Winter, and Jens O. Krömer. Metabolic engineering of pseudomonas putida kt2440 for the production of para-hydroxy benzoic acid. Frontiers in Bioengineering and Biotechnology, Nov 2016. URL: https://doi.org/10.3389/fbioe.2016.00090, doi:10.3389/fbioe.2016.00090. This article has 76 citations.

  14. (wang2022uncoveringtherole pages 15-16): Songwei Wang, Dongliang Liu, Muhammad Bilal, Wei Wang, and Xuehong Zhang. Uncovering the role of phzc as dahp synthase in shikimate pathway of pseudomonas chlororaphis ht66. Biology, 11:86, Jan 2022. URL: https://doi.org/10.3390/biology11010086, doi:10.3390/biology11010086. This article has 12 citations.

  15. (yu2016metabolicengineeringof pages 3-4): Shiqin Yu, Manuel R. Plan, Gal Winter, and Jens O. Krömer. Metabolic engineering of pseudomonas putida kt2440 for the production of para-hydroxy benzoic acid. Frontiers in Bioengineering and Biotechnology, Nov 2016. URL: https://doi.org/10.3389/fbioe.2016.00090, doi:10.3389/fbioe.2016.00090. This article has 76 citations.

Artifacts

Citations

  1. sterritt2018structuralandfunctional pages 1-2
  2. wang2022uncoveringtherole pages 10-12
  3. gosset2001microbialoriginof pages 5-7
  4. yu2016metabolicengineeringof pages 1-3
  5. dias2023fromdegraderto pages 1-2
  6. dias2023fromdegraderto pages 4-6
  7. gosset2001microbialoriginof pages 1-2
  8. dias2023fromdegraderto pages 2-4
  9. gosset2001microbialoriginof pages 7-8
  10. wang2022uncoveringtherole pages 1-3
  11. gosset2001microbialoriginof pages 10-11
  12. yu2016metabolicengineeringof pages 5-6
  13. wang2022uncoveringtherole pages 15-16
  14. yu2016metabolicengineeringof pages 3-4
  15. https://doi.org/10.1007/s10123-022-00282-5
  16. https://doi.org/10.3389/fbioe.2016.00090
  17. https://doi.org/10.1042/bsr20181605
  18. https://doi.org/10.1128/jb.183.13.4061-4070.2001
  19. https://doi.org/10.3390/biology11010086,
  20. https://doi.org/10.1042/bsr20181605,
  21. https://doi.org/10.1007/s10123-022-00282-5,
  22. https://doi.org/10.1128/jb.183.13.4061-4070.2001,
  23. https://doi.org/10.3389/fbioe.2016.00090,

📄 View Raw YAML

id: Q88LR3
gene_symbol: aroH
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: aroH (PP_1866) encodes a class-II (type-II) 3-deoxy-D-arabino-heptulosonate 7-phosphate synthase (DAHP synthase / DAH7PS; EC 2.5.1.54) in Pseudomonas putida KT2440. The enzyme catalyzes the aldol-like condensation of phosphoenolpyruvate (PEP) and D-erythrose 4-phosphate (E4P), with release of phosphate, to form 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP). This is the first committed step of the shikimate pathway, which ultimately produces chorismate, the branch-point precursor for the aromatic amino acids (phenylalanine, tyrosine, tryptophan) and many other aromatic metabolites. Class-II DAH7PS enzymes adopt a (beta/alpha)8 TIM-barrel fold and require a divalent metal cation for catalysis; the UniProt record for this protein indicates activity with manganese, cobalt or cadmium ions, with one cation bound per subunit. In bacterial DAHP-synthase nomenclature, AroH is conventionally the tryptophan-sensitive isoenzyme, although the regulatory properties of this specific P. putida protein have not been characterized experimentally. As a soluble central-metabolism enzyme acting on cytosolic substrates, AroH is expected to function in the cytoplasm.
existing_annotations:
- term:
    id: GO:0003849
    label: 3-deoxy-7-phosphoheptulonate synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: Core molecular function. The protein belongs to the class-II DAHP synthase family (RuleBase RU363071; InterPro IPR002480/PF01474; TIGR01358 DAHP_synth_II) and the UniProt catalytic-activity record (Rhea:14717, EC 2.5.1.54) describes condensation of PEP + E4P + H2O to DAHP + phosphate, exactly matching this GO term.
    action: ACCEPT
    reason: This term precisely captures the enzymatic activity of a class-II DAHP synthase. The assignment is strongly supported by sequence/family evidence (InterPro, RuleBase, NCBIfam TIGR01358) and conserved PEP- and metal-binding residues annotated in UniProt. This is the gene's core molecular function.
    supported_by:
    - reference_id: file:PSEPK/aroH/aroH-deep-research-falcon.md
      supporting_text: AroH is a DAH7PS catalyzing PEP + E4P -> DAHP/DAH7P and feeding the shikimate pathway; best-supported assignment is class-II / type-II DAH7PS (TIM-barrel fold, conserved metal-binding site), compatible with the UniProt/InterPro/Pfam context for Q88LR3 (PF01474 / DAHP_synth_2). No direct biochemical characterization of P. putida KT2440 AroH was found, so this is family/sequence-level inference.
- term:
    id: GO:0009073
    label: aromatic amino acid family biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: involved_in
  review:
    summary: Correct biological process. DAHP synthase catalyzes the first committed step of the shikimate pathway, which feeds chorismate biosynthesis and the downstream aromatic amino acid (Phe/Tyr/Trp) biosynthetic branches.
    action: ACCEPT
    reason: This is the canonical pathway role of a DAHP synthase and is consistent with the molecular function annotation. The term is appropriate and represents a core biological process for this gene. (The GOA stub used the label "aromatic amino acid biosynthetic process"; the current GO label for GO:0009073 is "aromatic amino acid family biosynthetic process" and is used here.)
core_functions:
- description: Catalyzes the first committed step of the shikimate pathway, the condensation of phosphoenolpyruvate and D-erythrose 4-phosphate to form 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP), as a class-II DAHP synthase.
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: 3-deoxy-7-phosphoheptulonate synthase activity (GO:0003849); EC 2.5.1.54; Rhea:14717 PEP + E4P + H2O = DAHP + phosphate.
  molecular_function:
    id: GO:0003849
    label: 3-deoxy-7-phosphoheptulonate synthase activity
  directly_involved_in:
  - id: GO:0009423
    label: chorismate biosynthetic process
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:11395471
  title: 'Microbial origin of plant-type 2-keto-3-deoxy-D-arabino-heptulosonate 7-phosphate synthases, exemplified by the chorismate- and tryptophan-regulated enzyme from Xanthomonas campestris'
  findings:
  - statement: Biochemical characterization of a class-II (plant-type / AroAII) DAHP synthase shows divalent-metal dependence (activity abolished by EDTA) and feedback inhibition by chorismate and tryptophan, providing the family-level basis for annotating P. putida AroH as a class-II DAHP synthase.
    supporting_text: Microbial origin of plant-type 2-keto-3-deoxy-D-arabino-heptulosonate 7-phosphate synthases from Xanthomonas campestris; Km(PEP) and Km(E4P) reported; chorismate and L-tryptophan inhibition characterized.
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: 'Citation-integrity fix: the original identifier PMID:11442790 was a wrong identifier (resolves to an unrelated veterinary pharmacokinetics article, "Pharmacokinetics of intravenous imipramine hydrochloride in cattle"). Replaced with PMID:11395471, the correct Gosset, Bonner & Jensen, J Bacteriol 2001;183:4061-4070 paper (DOI 10.1128/jb.183.13.4061-4070.2001), recovered via DOI lookup and PubMed-verified to match the title and supporting text (Xanthomonas campestris AroA(II) class-II DAHP synthase, chorismate/tryptophan feedback inhibition). Characterizes a class-II DAHP synthase in a different organism; relevant as family-level mechanistic support, not direct evidence for P. putida AroH.'