aroB

UniProt ID: Q88CV2
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

aroB (PP_5078) encodes 3-dehydroquinate synthase (DHQS; EC 4.2.3.4), a cytoplasmic enzyme that catalyzes the second step of the shikimate pathway. It converts 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP) into 3-dehydroquinate (DHQ), the first carbocyclic intermediate of the pathway, with release of inorganic phosphate. Catalysis requires a tightly bound NAD(+) used catalytically (transiently reduced and reoxidized) and a divalent metal cation (Co2+ or Zn2+) per subunit, driving a multi-step cascade of alcohol oxidation, phosphate elimination, carbonyl reduction, ring opening, and intramolecular aldol cyclization within a single active site. The shikimate pathway proceeds through chorismate, the branch-point precursor for the aromatic amino acids (phenylalanine, tyrosine, tryptophan) and other aromatic metabolites such as folate and ubiquinone. The pathway is present in bacteria, fungi, algae, and plants but absent in animals. The protein belongs to the sugar phosphate cyclases superfamily, dehydroquinate synthase family.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0003856 3-dehydroquinate synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: Core molecular function. aroB/PP_5078 is the 3-dehydroquinate synthase of P. putida KT2440 (EC 4.2.3.4, RHEA:21968), supported by UniProt/HAMAP rule MF_00110, conserved domain architecture (TIGR01357 aroB, Pfam PF01761), and KT2440 literature mapping the locus to this activity.
GO:0005737 cytoplasm
IEA
GO_REF:0000120
ACCEPT
Summary: DHQS acts on soluble cytosolic metabolites (DAHP, NAD+, divalent cation) and is a soluble intracellular enzyme of central metabolism. Consistent with UniProt subcellular location (cytoplasm) and the general architecture of the bacterial shikimate pathway. No experimental KT2440 localization assay exists, but the inference is well supported.
GO:0009073 aromatic amino acid biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: DHQS provides the second step of the shikimate pathway that supplies chorismate, the precursor of phenylalanine, tyrosine, and tryptophan. This biological process annotation is correct and represents a core role of the gene.
GO:0009423 chorismate biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: DHQS catalyzes step 2 of 7 in chorismate biosynthesis from D-erythrose 4-phosphate and phosphoenolpyruvate (UniPathway UPA00053). This is the most precise biological-process term for the gene and is well supported.
GO:0016838 carbon-oxygen lyase activity, acting on phosphates
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: InterPro2GO-derived MF term that is a more general parent of the specific 3-dehydroquinate synthase activity (GO:0003856) already annotated. DHQS is classified under EC 4.2.3.- (carbon-oxygen lyases acting on phosphates), so the term is not wrong, but it is redundant with and less informative than the specific child term.
Reason: The precise activity is already captured by GO:0003856. This broad lyase grouping term adds no information beyond the specific annotation and is a less informative restatement of the same molecular function.

Core Functions

3-dehydroquinate synthase catalyzing the second step of the shikimate pathway, converting DAHP to 3-dehydroquinate with release of phosphate, using catalytic NAD(+) and a divalent metal cofactor

Supporting Evidence:
  • GO_REF:0000120
    aroB/PP_5078 annotated as 3-dehydroquinate synthase activity (GO:0003856) via UniRule UR000001254 / HAMAP MF_00110, EC 4.2.3.4, RHEA:21968.
  • PMID:12534463
    KT2440 genome annotation assigns PP_5078 (aroB) as 3-dehydroquinate synthase; deep research (aroB-deep-research-falcon.md) corroborates the PP_5078 to aroB / DHQS mapping and the DAHP to DHQ reaction.

References

Gene Ontology annotation through association of InterPro records with GO terms
Combined Automated Annotation using Multiple IEA Methods
Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.
  • Genome sequence of P. putida KT2440 in which the aroB/PP_5078 locus (Q88CV2) is annotated as 3-dehydroquinate synthase.

Deep Research

Asta

(aroB-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:27:25.934890

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.705)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[2] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.700)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[3] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[4] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[5] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.685)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[6] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.680)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[7] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.680)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[8] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.679)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[9] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[10] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.669)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[11] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.665)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[12] The y-ome defines the 35% of Escherichia coli genes that lack experimental evidence of function

  • Authors: S. Ghatak, Zachary A. King, Anand V. Sastry, B. Palsson
  • Year: 2019
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/c0336e0a70554304893a9e2d010ee30bd6872b10
  • DOI: 10.1093/nar/gkz030
  • PMID: 30698741
  • PMCID: 6412132
  • Citations: 133
  • Influential citations: 4
  • Summary: The misconception that a gene in E. coli whose primary name starts with ‘y’ is unannotated is resolved, and the value of the y-ome is discussed for systematic improvement ofE.
  • Evidence snippets:
  • Snippet 1 (score: 0.659)
    > There are several knowledge bases that represent the collected knowledge of the E. coli K-12 MG1655 genome: EcoCyc (11), EcoGene (12), UniProt (13) and RefSeq (14). Other useful knowledge bases cater to specific classes of gene products, such as the RegulonDB, which contains manually curated functional information about transcription factors in E. coli (15). Our initial review of these knowledge bases yielded conflicting information on gene function and level of annotation for many E. coli genes. Any attempt to systematically assess the function of unannotated genes must therefore draw from multiple knowledge bases and resolve these conflicts.
    > Many research groups have categorized E. coli genes and proteins by annotation quality as a part of their studies. In 2009, Hu et (16). First, they identified all unannotated proteins in the K-12 W3110 and MG1655 genomes. In order for a protein-encoding gene to be considered functionally uncharacterized in their analysis, it had to meet the following criteria: (i) The gene name begins with 'y', (ii) the gene does not have a known pathway within EcoCyc and (iii) the gene does not have a functional description in Gen-ProtEC (17) (any gene with a description containing the words 'predicted', 'hypothetical', or 'conserved'). Based on these criteria, it was determined that 1431 of 4225 protein coding sequences were not functionally annotated. In 2015, Kim et al. published a database called EcoliNet that curated and predicted cofunctional gene networks for every protein coding gene in the E. coli genome (18). This study also quantified the number of uncharacterized protein coding genes in E. coli. To assess functional annotation, they used the presence of experimentally supported 'biological process' annotations in the Gene Ontology database (19). They concluded that ∼2000 protein coding genes in E. coli were not functionally annotated. The most comprehensive effort to assess the level of annotation in bacterial genomes has been Computational Bridges to Experiments (COM-BREX) (20,21).

[13] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.657)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[14] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.653)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[15] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.653)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[16] Protocol for gene annotation, prediction, and validation of genomic gene expansion

  • Authors: Quanwei Zhang, Zhengdong D. Zhang
  • Year: 2022
  • Venue: STAR Protocols
  • URL: https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3
  • DOI: 10.1016/j.xpro.2022.101692
  • PMID: 36125934
  • PMCID: 9494284
  • Citations: 1
  • Summary: A detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation is described and steps to discover functionality of each copy of replicated genes are detailed.
  • Evidence snippets:
  • Snippet 1 (score: 0.651)
    > 3. Gene annotation and functional annotation. a. Gene structure annotation.
    > In addition to gene prediction models, evidence from orthologous protein sequences and transcriptome assembly could be used to improve annotation quality. Protein sequences of orthologous genes can be obtained from UniProt (The UniProt, 2017). Ones from Swiss-Port have been reviewed and thus are of higher quality. Transcriptome assembly may be available from previous studies or can be assembled de novo from RNA-seq reads by Trinity (Haas et al., 2013). High quality transcriptome assembly can be selected as described in (Zhang et al., 2021). Note: Details about gene structure annotation (Holt and Yandell, 2011) can be found at http:// gmod.org/wiki/MAKER_Tutorial, https://darencard.net/blog/2017-05-16-maker-genomeannotation/, and the protocol (Campbell et al., 2014).
    > b. Quality measurement and functional annotation.
    > For each predicted gene, Maker2 provides the annotation edit distance (AED) score, which measures the goodness of fit between its predicted gene structure and its evidence support. The lower the score, the more accurate the prediction. If more than 90% genes with AED scores lower than 0.5, the genome can be considered well annotated. In addition to the AED score, a high proportion of recognizable domains contained in predicted protein -e.g., higher than 50% -also indicates a good annotation. Recognizable protein domains can by scanned by InterProScan (Jones et al., 2014), assigning potential function to predicted genes.
    > Note: Besides the aforementioned quality measurement, we strongly recommend measuring the completeness of the genome assembly and annotation by checking the existence of a set of Benchmarking Universal Single-Copy Orthologs (BUSCO) (Simao et al., 2015). A high-level completeness of genome assembly and annotation is imperative for a better identification of gene expansion. Based on the result of this analysis, researchers can decide whether they need to further improve the genome assembly before predicting gene expansion. A detailed protocol of BUSCO is available at

[17] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.648)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[18] Telomere-to-telomere genome assembly of Phoxinus lagowskii

  • Authors: Yanfeng Zhou, Chunhai Chen, Di’an Fang, Chenhe Wang, Yajuan Peng et al.
  • Year: 2025
  • Venue: Scientific Data
  • URL: https://www.semanticscholar.org/paper/23f573aff23769df3b733d508b46f721cadb94ac
  • DOI: 10.1038/s41597-025-05367-0
  • PMID: 40533492
  • PMCID: 12177035
  • Citations: 1
  • Summary: A T2T (Telomere-to-telomere) genome for P. lagowskii with chromosome-level is reported, serving as an invaluable resource for studies in evolution, comparative genomics, fish breeding applications, and ecological research.
  • Evidence snippets:
  • Snippet 1 (score: 0.644)
    > The gene set of 24,610 genes were functionally annotated using diamond v0.8.23
    > with an E-value threshold of 1E-5 based on the five databases including NR (NCBI nonredundant protein), TrEMB (http://www.uniprot.org), KOG 42 , KEEG (Kyoto Encyclopedia of Genes and Genomes, http://www. genome.jp/kegg/), Swiss-Prot (http://www.gpmaw.com/html/swiss-prot.html). The protein motifs and domains were identified using the InterProScan with InterPro 93.0 43 . The GO Ontology (GO) was classified from the results of InterProScan 44 . The annotation of 24,599 predicted genes (99.96%) out of the total 24,610 genes can be found by at least one database (Fig. 3b and Table 6). Of these functional proteins, 19,367 genes (~78.7%) were supported by five databases (Fig. 3a).

[19] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.643)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.
  • Authors: Jaap van der Heijden, Asanda Mazubane, Marko Sallisalmi, E. Vorontsov, J. Tenhunen et al.
  • Year: 2025
  • Venue: Clinical Proteomics
  • URL: https://www.semanticscholar.org/paper/ae917565de913693a2713b46c3eda672d9d01c7f
  • DOI: 10.1186/s12014-025-09556-2
  • PMID: 40885913
  • PMCID: 12398169
  • Summary: The altered proteomic profile of hyaluronan-related proteins as reflected by the GO terms indicates a complex dysregulation not only in hyaluronan metabolism and extracellular matrix, but also in the regulation of several proteolytic enzymes.
  • Evidence snippets:
  • Snippet 1 (score: 0.641)
    > The identification of hyaluronan-associated genes was performed using Python 3.10.12. The UniProt REST Web Application Programming Interface (API) was used to query and retrieve gene annotations that met specific criteria (UniProt Consortium). A keyword-based query, utilizing the terms "hyaluronan, " "hyaluronic acid, " "hyaluronidase, " "hyaluronic acid synthase, " "hyaluronate binding protein, " "hyaluronan oligosaccharides, " "hyaluronan synthase, " and "hyaluronan receptor, " was executed across our dataset of 663 genes. Gene symbols were returned as "hits" when the query keywords were found in the annotations, including functional comments, Gene Ontology (GO) terms, and cross-referenced databases, formatted in JSON.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  2. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  3. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  4. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  5. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  6. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  7. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  8. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  9. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  10. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  11. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  12. S. Ghatak, Zachary A. King, Anand V. Sastry, B. Palsson (2019). The y-ome defines the 35% of Escherichia coli genes that lack experimental evidence of function. Nucleic Acids Research. https://www.semanticscholar.org/paper/c0336e0a70554304893a9e2d010ee30bd6872b10
  13. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  14. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  15. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  16. Quanwei Zhang, Zhengdong D. Zhang (2022). Protocol for gene annotation, prediction, and validation of genomic gene expansion. STAR Protocols. https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3
  17. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  18. Yanfeng Zhou, Chunhai Chen, Di’an Fang, Chenhe Wang, Yajuan Peng et al. (2025). Telomere-to-telomere genome assembly of Phoxinus lagowskii. Scientific Data. https://www.semanticscholar.org/paper/23f573aff23769df3b733d508b46f721cadb94ac
  19. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  20. Jaap van der Heijden, Asanda Mazubane, Marko Sallisalmi, E. Vorontsov, J. Tenhunen et al. (2025). Plasma proteomics in septic shock and alcohol-related pancreatitis: a hyaluronan-centered approach. Clinical Proteomics. https://www.semanticscholar.org/paper/ae917565de913693a2713b46c3eda672d9d01c7f

Falcon

(aroB-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 33 citations 1 artifacts 2026-06-11T20:48:04.491466

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research Report: Functional Annotation of aroB (UniProt Q88CV2; locus PP_5078) in Pseudomonas putida KT2440

0) Target verification (mandatory)

The target protein is correctly identified as 3-dehydroquinate synthase (DHQS) encoded by aroB in Pseudomonas putida KT2440. A KT2440 transcriptomics study explicitly maps the locus tag PP_5078 to aroB and describes it as 3-dehydroquinate synthase (“3-dehydroquinate synthase (aroB) (PP_5078)”). (lopezlara2020influenceofrehydration pages 2-4). Independent KT2440 metabolic-engineering literature also refers to aroB as encoding the native 3-dehydroquinate synthase in the shikimate pathway. (ling2022muconicacidproduction pages 1-2). Authoritative pathway reviews define AroB/DHQS as EC 4.2.3.4 and place it within the sugar phosphate cyclase superfamily, consistent with the UniProt-provided family/domain context. (derrer2013theshikimatepathway pages 3-4, derrer2013theshikimatepathway pages 4-5).

1) Key concepts and current understanding

1.1 The shikimate pathway and the role of AroB/DHQS

The shikimate pathway is a conserved biosynthetic route (present in bacteria, fungi, algae, and plants but absent in animals) that converts central-carbon precursors into chorismate, the branch-point precursor for aromatic amino acids and many specialized metabolites. (shende2024theshikimatepathway pages 3-4).

AroB/DHQS (3-dehydroquinate synthase) catalyzes the second step of the pathway: conversion of 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP) into 3-dehydroquinate (DHQ), which is the first carbocyclic product (and a key intermediate) of the shikimate pathway. (derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4, maeda2012theshikimatepathway pages 7-8).

1.2 Enzymatic reaction, substrate specificity, and products

Across authoritative reviews, DHQS is described as catalyzing:
- Substrate: DAHP (a sugar phosphate; often described as the pyranose form)
- Product: 3-dehydroquinate (DHQ) (carbocyclic intermediate)
- Phosphate: phosphate chemistry is integral; phosphate is eliminated during the mechanism and inorganic phosphate is discussed as a product/activator in mechanistic descriptions. (dev2012structureandfunction pages 2-4, derrer2013theshikimatepathway pages 3-4).

The enzyme’s functional specificity is therefore primarily for DAHP within the shikimate pathway (i.e., it is not a broad-spectrum cyclase in this context but a dedicated shikimate-pathway enzyme). (derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4).

1.3 Cofactors and mechanistic steps

DHQS is widely characterized as an enzyme with an unusually complex, multi-step catalytic cascade occurring in a single active site:
- It uses NAD+ catalytically (transient reduction to NADH and reoxidation back to NAD+), and requires a divalent metal ion, typically described as Co2+ or Zn2+. (derrer2013theshikimatepathway pages 3-4, dev2012structureandfunction pages 2-4, maeda2012theshikimatepathway pages 7-8, derrer2013theshikimatepathway pages 4-5).
- Mechanistic steps described in reviews include an ordered sequence of alcohol oxidation, phosphate elimination, carbonyl reduction, ring opening, and intramolecular aldol condensation/cyclization to yield DHQ. (dev2012structureandfunction pages 2-4, derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4).

A high-level mechanistic consensus is that the enzyme initiates a redox-triggered cascade that yields the carbocycle while suppressing side reactions (“shunt” reactions). (shende2024theshikimatepathway pages 3-4).

1.4 Quantitative enzymology (comparative, conserved bacterial DHQS)

While not specific to P. putida KT2440, bacterial DHQS enzymes share conserved chemistry and cofactor usage; thus kinetic data from well-studied bacterial homologs provide useful bounds on expected catalytic behavior. For Mycobacterium tuberculosis DHQS, a review summarizes reported values KM(DAHP) = 6.3 µM, KM(NAD+) = 70 µM, and kcat ≈ 0.63 s−1, and also notes EDTA sensitivity with best restoration by Co2+ (and partial support by other divalent cations, with Zn often present). (nunes2020mycobacteriumtuberculosisshikimate pages 10-12).

2) aroB (PP_5078) functional annotation in P. putida KT2440

2.1 Biological function and pathway placement

Given the direct PP_5078→aroB mapping and the conserved reaction, aroB (PP_5078) in KT2440 is best annotated as a cytosolic metabolic enzyme of the shikimate pathway catalyzing DAHP → DHQ. (lopezlara2020influenceofrehydration pages 2-4, derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4).

In bacteria, the shikimate pathway supplies chorismate, which supports aromatic amino acid biosynthesis and numerous downstream aromatic metabolites. (shende2024theshikimatepathway pages 3-4). In P. putida, this central role underpins why aroB is frequently targeted in metabolic engineering to raise flux to aromatic products. (ling2022muconicacidproduction pages 1-2, camposmagana2024combinatorialengineeringreveals pages 1-4).

2.2 Cellular localization and site of action

No retrieved source provided a direct experimental subcellular localization assay for KT2440 AroB (e.g., fractionation, microscopy tagging). However, the enzyme acts on soluble cytosolic metabolites (DAHP, NAD+, divalent cations), and DHQS is treated in the literature as a soluble, intracellular (cytosolic) enzyme of central metabolism. This inference is consistent with the enzyme’s chemistry and its use in intracellular pathway engineering. (derrer2013theshikimatepathway pages 3-4, dev2012structureandfunction pages 2-4, ling2022muconicacidproduction pages 1-2).

2.3 Expression and physiological context in KT2440

In a KT2440 desiccation/rehydration transcriptome study, aroB (PP_5078) is reported among genes in an overrepresented amino-acid biosynthesis functional group and is described as upregulated during resuscitation after desiccation. (lopezlara2020influenceofrehydration pages 2-4). The same excerpt links this group to aromatic amino acid biosynthesis (phenylalanine/tyrosine), consistent with increased demand for shikimate-pathway flux during recovery. (lopezlara2020influenceofrehydration pages 2-4).

3) Recent developments and latest research (prioritizing 2023–2024)

3.1 2024 mechanistic/structural synthesis of DHQS and pathway context

A 2024 Natural Product Reports review (“The shikimate pathway: gateway to metabolic diversity”) explicitly highlights that mechanistic proposals for DHQS are informed by high-resolution crystal structures of DHQS in complex with metals, nicotinamide cofactors, and substrate-based inhibitors, and summarizes DHQS as a multi-step enzyme stabilizing intermediates and minimizing off-pathway reactions. (shende2024theshikimatepathway pages 3-4). This consolidates the current understanding that DHQS is not merely a simple cyclase but a tightly choreographed redox-enabled cyclization catalyst.

3.2 2024–2025 systems metabolic engineering in P. putida: aroB emerges as a bottleneck

A 2024 bioRxiv preprint applying combinatorial Design-of-Experiments (DoE) to P. putida shikimate and pABA pathways identifies aroB expression as a significant limiting factor for para-aminobenzoic acid (pABA) production. (camposmagana2024combinatorialengineeringreveals pages 1-4). Quantitatively, from 14 representative strains spanning a theoretical 512-combination design space, the authors report 2–186.2 mg/L pABA in the initial screen and 232.1 mg/L after a second engineering round guided by modeling. (camposmagana2024combinatorialengineeringreveals pages 1-4). A peer-reviewed continuation (2025) reports the same workflow and titers and again identifies aroB as a critical bottleneck (included here for completeness beyond the user’s 2024 window). (camposmagana2025combinatorialengineeringpinpoints pages 1-2).

This line of work supports an expert interpretation that, in P. putida, DHQS capacity can constrain carbon throughput through the shikimate node when production goals demand high chorismate-derived flux.

3.3 2024 microbial production (other hosts) reinforces aroB as a common engineering lever

In 2024, Corynebacterium crenatum was engineered for L-tyrosine production by, among other interventions, overexpressing aroB/aroD/aroE. Reported titers reached 6.42 g/L in shake flasks and 34.6 g/L in fed-batch fermentation, with the mixed carbon-source strategy giving 16.9% higher L-tyrosine than glucose alone in flask conditions. (yang2024metabolicengineeringof pages 1-3). While not P. putida, this is relevant as a recent demonstration of real-world pathway design where increasing flux through the DAHP→DHQ→shikimate segment (including DHQS) is part of achieving industrially relevant aromatic amino acid titers.

4) Current applications and real-world implementations

4.1 P. putida KT2440 as a chassis for shikimate-derived bioproducts

A prominent implementation is muconic acid production in engineered P. putida KT2440. In this Nature Communications study (published Aug 2022; still actively cited as an implementation benchmark), the authors state that overexpression of aroB encoding the native 3-dehydroquinate synthase enabled efficient muconic acid production from mixed sugars. (ling2022muconicacidproduction pages 1-2). They report 33.7 g/L muconate, 0.18 g/L/h productivity, and 46% molar yield (reported as 92% of maximum theoretical yield). (ling2022muconicacidproduction pages 1-2). These data demonstrate industrially relevant titers/productivities and position aroB as a concrete intervention point in deployed strain designs.

4.2 pABA as an industrial intermediate

pABA is described as a widely used industrial intermediate, and the 2024 DoE study in P. putida provides a quantitative example showing that systematic tuning of shikimate-pathway gene expression can move pABA titers across two orders of magnitude, with aroB identified as a key bottleneck. (camposmagana2024combinatorialengineeringreveals pages 1-4).

5) Expert opinions and analysis (authoritative synthesis)

Authoritative reviews emphasize two main reasons DHQS/AroB is strategically important:
1. Pathway control point: DHQS forms the first carbocyclic intermediate and sits early in the pathway, so it can become rate-limiting when high flux is required for aromatic amino acids or downstream specialized metabolites. (shende2024theshikimatepathway pages 3-4, camposmagana2024combinatorialengineeringreveals pages 1-4).
2. Mechanistic complexity enabling inhibitor design: DHQS has well-defined cofactor and metal requirements and multi-step chemistry; high-resolution structures with cofactors/inhibitors have informed mechanistic models, making DHQS a tractable enzyme for mechanistic interrogation and (in broader contexts) inhibitor development. (shende2024theshikimatepathway pages 3-4, derrer2013theshikimatepathway pages 4-5).

A practical interpretation consistent with recent P. putida engineering results is that aroB expression/capacity often needs explicit optimization in overproduction settings (e.g., pABA), and can be part of high-performing strain architectures (e.g., muconate). (camposmagana2024combinatorialengineeringreveals pages 1-4, ling2022muconicacidproduction pages 1-2).

6) Key statistics and quantitative data from studies (selected)

  • Muconate production in engineered P. putida KT2440 with aroB overexpression in the design: 33.7 g/L, 0.18 g/L/h, 46% molar yield (92% of theoretical maximum). Publication date: Aug 2022. URL: https://doi.org/10.1038/s41467-022-32296-y (ling2022muconicacidproduction pages 1-2).
  • pABA production in P. putida DoE study: initial titers 2–186.2 mg/L, improved to 232.1 mg/L after a second engineering round; aroB identified as bottleneck. Preprint posted Jun 17, 2024. URL: https://doi.org/10.1101/2024.06.17.599342 (camposmagana2024combinatorialengineeringreveals pages 1-4).
  • Comparative bacterial DHQS kinetics (M. tuberculosis): KM(DAHP) 6.3 µM, KM(NAD+) 70 µM, kcat ~0.63 s−1; metal dependence with best restoration by Co2+. Publication date: Mar 2020. URL: https://doi.org/10.3390/molecules25061259 (nunes2020mycobacteriumtuberculosisshikimate pages 10-12).

7) Summary of functional annotation (for database-style use)

  • Gene / locus / accession: aroB / PP_5078 / UniProt Q88CV2 (verified in KT2440 context). (lopezlara2020influenceofrehydration pages 2-4)
  • Enzyme name / EC: 3-dehydroquinate synthase (DHQS) / EC 4.2.3.4. (derrer2013theshikimatepathway pages 3-4, maeda2012theshikimatepathway pages 7-8)
  • Primary reaction: DAHP → 3-dehydroquinate (DHQ) (second shikimate step; first carbocyclic product). (derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4)
  • Cofactors: catalytic NAD+ and a divalent metal (commonly Co2+ or Zn2+). (dev2012structureandfunction pages 2-4, derrer2013theshikimatepathway pages 4-5)
  • Pathway: shikimate pathway → chorismate → aromatic amino acids and diverse aromatic metabolites. (shende2024theshikimatepathway pages 3-4)
  • Cellular location (best-supported): intracellular/cytosolic enzyme acting on soluble metabolites; direct KT2440 localization experiment not found in retrieved sources. (derrer2013theshikimatepathway pages 3-4, ling2022muconicacidproduction pages 1-2)

8) Evidence table (compiled)

Evidence type Organism/system Key finding Quantitative data (if any) Publication (authors, year) URL/DOI
identity Pseudomonas putida KT2440 PP_5078 is explicitly annotated as aroB, encoding 3-dehydroquinate synthase; reported as upregulated during recovery from desiccation/rehydration, supporting the locus-to-function mapping for the target gene/protein (lopezlara2020influenceofrehydration pages 2-4) Microarray study with triplicate analyses; numeric fold-change not given in excerpt (lopezlara2020influenceofrehydration pages 2-4) López-Lara et al., 2020 https://doi.org/10.1186/s13213-020-01596-3
identity/pathway role Pseudomonas putida KT2440 engineered for muconate production The native aroB in P. putida KT2440 is identified as encoding 3-dehydroquinate synthase; overexpression improved flux to shikimate-pathway-derived muconic acid (ling2022muconicacidproduction pages 1-2) Muconate production reached 33.7 g/L, 0.18 g/L/h, 46% molar yield (92% of theoretical maximum) in engineered strains that included aroB overexpression (ling2022muconicacidproduction pages 1-2) Ling et al., 2022 https://doi.org/10.1038/s41467-022-32296-y
mechanism Bacterial DHQS/AroB (reviewed across taxa) DHQS/AroB (EC 4.2.3.4) catalyzes conversion of DAHP to 3-dehydroquinate (DHQ), the second step of the shikimate pathway and first carbocyclic intermediate-forming reaction; belongs to the sugar phosphate cyclase superfamily (derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4) Reaction-level annotation; no organism-specific kinetic value in these excerpts (derrer2013theshikimatepathway pages 3-4, shende2024theshikimatepathway pages 3-4) Derrer et al., 2013; Shende et al., 2024 https://doi.org/10.2741/4155; https://doi.org/10.1039/d3np00037k
mechanism/cofactors Bacterial DHQS/AroB (reviewed across taxa) Catalysis requires NAD+ and a divalent metal ion; Co2+ and Zn2+ are the principal reported cofactors. Mechanism proceeds through oxidation, phosphate elimination, reduction, ring opening, and intramolecular aldol cyclization (derrer2013theshikimatepathway pages 3-4, dev2012structureandfunction pages 2-4, maeda2012theshikimatepathway pages 7-8, derrer2013theshikimatepathway pages 4-5) Mechanistic sequence of ~5 chemical steps in one active site; catalytic NAD+ use noted (dev2012structureandfunction pages 2-4) Dev et al., 2012; Derrer et al., 2013; Maeda & Dudareva, 2012 https://doi.org/10.2174/157489312803900983; https://doi.org/10.2741/4155; https://doi.org/10.1146/annurev-arplant-042811-105439
mechanism/kinetics Mycobacterium tuberculosis DHQS (comparative authoritative source for conserved AroB chemistry) AroB/DHQS performs multiple transformations in one active site; kinetics and metal dependence support the conserved bacterial mechanism for DHQS enzymes (nunes2020mycobacteriumtuberculosisshikimate pages 10-12) KM(DAHP) = 6.3 µM, KM(NAD+) = 70 µM, kcat ≈ 0.63 s⁻¹; activity abolished by EDTA and best restored by Co2+ (nunes2020mycobacteriumtuberculosisshikimate pages 10-12) Nunes et al., 2020 https://doi.org/10.3390/molecules25061259
application Pseudomonas putida pABA pathway engineering Combinatorial engineering identified aroB expression as a significant limiting factor / bottleneck for para-aminobenzoic acid production in P. putida, making DHQS a practical intervention point in strain design (camposmagana2024combinatorialengineeringreveals pages 1-4, camposmagana2024combinatorialengineeringreveals pages 11-14) Initial strain set produced 2–186.2 mg/L pABA; second round reached 232.1 mg/L (reported in related preprint/peer-reviewed continuation) (camposmagana2024combinatorialengineeringreveals pages 1-4, camposmagana2025combinatorialengineeringpinpoints pages 1-2) Campos-Magaña et al., 2024; Campos-Magaña et al., 2025 https://doi.org/10.1101/2024.06.17.599342; https://doi.org/10.1186/s13036-025-00553-5
application Corynebacterium crenatum metabolic engineering Overexpression of aroB with aroD/aroE was part of a successful aromatic amino acid engineering strategy, showing broader real-world use of DHQS tuning in microbial biomanufacturing (yang2024metabolicengineeringof pages 1-3) 6.42 g/L L-tyrosine in shake flask; 34.6 g/L in fed-batch; 16.9% higher at 3:1 glucose:mannitol vs glucose alone (yang2024metabolicengineeringof pages 1-3) Yang et al., 2024 https://doi.org/10.1186/s12934-024-02564-1
omics Pseudomonas putida KT2440 under desiccation/rehydration aroB/PP_5078 was included among genes in an overrepresented amino-acid biosynthesis response, consistent with its role in aromatic amino acid precursor supply through the shikimate pathway (lopezlara2020influenceofrehydration pages 2-4) Upregulation reported after 18 days desiccation followed by 20 min or 24 h rehydration sampling; no fold-change in excerpt (lopezlara2020influenceofrehydration pages 2-4) López-Lara et al., 2020 https://doi.org/10.1186/s13213-020-01596-3

Table: This table compiles key evidence verifying that Pseudomonas putida KT2440 PP_5078 is aroB encoding 3-dehydroquinate synthase, and summarizes its conserved reaction chemistry, cofactors, pathway role, omics support, and quantitative metabolic-engineering applications.

9) Limitations of the current evidence set

  • Direct subcellular localization (e.g., experimental cytosolic fractionation or fluorescent tagging in KT2440) was not identified in the retrieved texts; localization here is inferred from biochemical role and general bacterial shikimate pathway architecture. (derrer2013theshikimatepathway pages 3-4, ling2022muconicacidproduction pages 1-2)
  • Domain architecture / InterPro/Pfam features for Q88CV2 were not directly extracted from primary literature in this run; however, the enzyme’s superfamily placement (sugar phosphate cyclase) and conserved mechanistic features are strongly supported by authoritative reviews. (derrer2013theshikimatepathway pages 3-4, derrer2013theshikimatepathway pages 4-5)

References

  1. (lopezlara2020influenceofrehydration pages 2-4): Lilia I. López-Lara, Laura A. Pazos-Rojas, Lesther E. López-Cruz, Yolanda E. Morales-García, Verónica Quintero-Hernández, Jesús de la Torre, Pieter van Dillewijn, Jesús Muñoz-Rojas, and Antonino Baez. Influence of rehydration on transcriptome during resuscitation of desiccated pseudomonas putida kt2440. Annals of Microbiology, Sep 2020. URL: https://doi.org/10.1186/s13213-020-01596-3, doi:10.1186/s13213-020-01596-3. This article has 16 citations and is from a peer-reviewed journal.

  2. (ling2022muconicacidproduction pages 1-2): Chen Ling, George L. Peabody, Davinia Salvachúa, Young-Mo Kim, Colin M. Kneucker, Christopher H. Calvey, Michela A. Monninger, Nathalie Munoz Munoz, Brenton C. Poirier, Kelsey J. Ramirez, Peter C. St. John, Sean P. Woodworth, Jon K. Magnuson, Kristin E. Burnum-Johnson, Adam M. Guss, Christopher W. Johnson, and Gregg T. Beckham. Muconic acid production from glucose and xylose in pseudomonas putida via evolution and metabolic engineering. Nature Communications, Aug 2022. URL: https://doi.org/10.1038/s41467-022-32296-y, doi:10.1038/s41467-022-32296-y. This article has 141 citations and is from a highest quality peer-reviewed journal.

  3. (derrer2013theshikimatepathway pages 3-4): Bianca Derrer, P. Macheroux, and B. Kappes. The shikimate pathway in apicomplexan parasites: implications for drug development. Frontiers in bioscience, 18:944-69, Jun 2013. URL: https://doi.org/10.2741/4155, doi:10.2741/4155. This article has 37 citations and is from a peer-reviewed journal.

  4. (derrer2013theshikimatepathway pages 4-5): Bianca Derrer, P. Macheroux, and B. Kappes. The shikimate pathway in apicomplexan parasites: implications for drug development. Frontiers in bioscience, 18:944-69, Jun 2013. URL: https://doi.org/10.2741/4155, doi:10.2741/4155. This article has 37 citations and is from a peer-reviewed journal.

  5. (shende2024theshikimatepathway pages 3-4): Vikram V. Shende, Katherine D. Bauman, and Bradley S. Moore. The shikimate pathway: gateway to metabolic diversity. Natural product reports, 41:604-648, Jan 2024. URL: https://doi.org/10.1039/d3np00037k, doi:10.1039/d3np00037k. This article has 173 citations and is from a peer-reviewed journal.

  6. (maeda2012theshikimatepathway pages 7-8): Hiroshi Maeda and Natalia Dudareva. The shikimate pathway and aromatic amino acid biosynthesis in plants. Jun 2012. URL: https://doi.org/10.1146/annurev-arplant-042811-105439, doi:10.1146/annurev-arplant-042811-105439. This article has 1905 citations and is from a domain leading peer-reviewed journal.

  7. (dev2012structureandfunction pages 2-4): Aditya Dev, Satya Tapas, Shivendra Pratap, and Pravindra Kumar. Structure and function of enzymes of shikimate pathway. Current Bioinformatics, 7:374-391, Nov 2012. URL: https://doi.org/10.2174/157489312803900983, doi:10.2174/157489312803900983. This article has 20 citations and is from a peer-reviewed journal.

  8. (nunes2020mycobacteriumtuberculosisshikimate pages 10-12): José E. S. Nunes, Mario A. Duque, Talita F. de Freitas, Luiza Galina, Luis F. S. M. Timmers, Cristiano V. Bizarro, Pablo Machado, Luiz A. Basso, and Rodrigo G. Ducati. Mycobacterium tuberculosis shikimate pathway enzymes as targets for the rational design of anti-tuberculosis drugs. Molecules, 25:1259, Mar 2020. URL: https://doi.org/10.3390/molecules25061259, doi:10.3390/molecules25061259. This article has 74 citations.

  9. (camposmagana2024combinatorialengineeringreveals pages 1-4): Marco A Campos-Magaña, Sara Moreno-Paz, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering reveals shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. bioRxiv, Jun 2024. URL: https://doi.org/10.1101/2024.06.17.599342, doi:10.1101/2024.06.17.599342. This article has 0 citations.

  10. (camposmagana2025combinatorialengineeringpinpoints pages 1-2): Marco A Campos-Magaña, Sara Moreno-Paz, Maria Martin-Pascual, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering pinpoints shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. Journal of Biological Engineering, Sep 2025. URL: https://doi.org/10.1186/s13036-025-00553-5, doi:10.1186/s13036-025-00553-5. This article has 0 citations and is from a peer-reviewed journal.

  11. (yang2024metabolicengineeringof pages 1-3): Gang Yang, Sicheng Xiong, Mingzhu Huang, Bin Liu, Yanna Shao, and Xuelan Chen. Metabolic engineering of corynebacterium crenatum for enhanced l-tyrosine production from mannitol and glucose. Microbial Cell Factories, Oct 2024. URL: https://doi.org/10.1186/s12934-024-02564-1, doi:10.1186/s12934-024-02564-1. This article has 7 citations and is from a peer-reviewed journal.

  12. (camposmagana2024combinatorialengineeringreveals pages 11-14): Marco A Campos-Magaña, Sara Moreno-Paz, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering reveals shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. bioRxiv, Jun 2024. URL: https://doi.org/10.1101/2024.06.17.599342, doi:10.1101/2024.06.17.599342. This article has 0 citations.

Artifacts

Citations

  1. lopezlara2020influenceofrehydration pages 2-4
  2. ling2022muconicacidproduction pages 1-2
  3. shende2024theshikimatepathway pages 3-4
  4. nunes2020mycobacteriumtuberculosisshikimate pages 10-12
  5. camposmagana2024combinatorialengineeringreveals pages 1-4
  6. camposmagana2025combinatorialengineeringpinpoints pages 1-2
  7. yang2024metabolicengineeringof pages 1-3
  8. dev2012structureandfunction pages 2-4
  9. derrer2013theshikimatepathway pages 3-4
  10. derrer2013theshikimatepathway pages 4-5
  11. maeda2012theshikimatepathway pages 7-8
  12. camposmagana2024combinatorialengineeringreveals pages 11-14
  13. https://doi.org/10.1038/s41467-022-32296-y
  14. https://doi.org/10.1101/2024.06.17.599342
  15. https://doi.org/10.3390/molecules25061259
  16. https://doi.org/10.1186/s13213-020-01596-3
  17. https://doi.org/10.2741/4155;
  18. https://doi.org/10.1039/d3np00037k
  19. https://doi.org/10.2174/157489312803900983;
  20. https://doi.org/10.1146/annurev-arplant-042811-105439
  21. https://doi.org/10.1101/2024.06.17.599342;
  22. https://doi.org/10.1186/s13036-025-00553-5
  23. https://doi.org/10.1186/s12934-024-02564-1
  24. https://doi.org/10.1186/s13213-020-01596-3,
  25. https://doi.org/10.1038/s41467-022-32296-y,
  26. https://doi.org/10.2741/4155,
  27. https://doi.org/10.1039/d3np00037k,
  28. https://doi.org/10.1146/annurev-arplant-042811-105439,
  29. https://doi.org/10.2174/157489312803900983,
  30. https://doi.org/10.3390/molecules25061259,
  31. https://doi.org/10.1101/2024.06.17.599342,
  32. https://doi.org/10.1186/s13036-025-00553-5,
  33. https://doi.org/10.1186/s12934-024-02564-1,

📄 View Raw YAML

id: Q88CV2
gene_symbol: aroB
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: aroB (PP_5078) encodes 3-dehydroquinate synthase (DHQS; EC 4.2.3.4), a cytoplasmic enzyme that catalyzes the second step of the shikimate pathway. It converts 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP) into 3-dehydroquinate (DHQ), the first carbocyclic intermediate of the pathway, with release of inorganic phosphate. Catalysis requires a tightly bound NAD(+) used catalytically (transiently reduced and reoxidized) and a divalent metal cation (Co2+ or Zn2+) per subunit, driving a multi-step cascade of alcohol oxidation, phosphate elimination, carbonyl reduction, ring opening, and intramolecular aldol cyclization within a single active site. The shikimate pathway proceeds through chorismate, the branch-point precursor for the aromatic amino acids (phenylalanine, tyrosine, tryptophan) and other aromatic metabolites such as folate and ubiquinone. The pathway is present in bacteria, fungi, algae, and plants but absent in animals. The protein belongs to the sugar phosphate cyclases superfamily, dehydroquinate synthase family.
existing_annotations:
- term:
    id: GO:0003856
    label: 3-dehydroquinate synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: Core molecular function. aroB/PP_5078 is the 3-dehydroquinate synthase of P. putida KT2440 (EC 4.2.3.4, RHEA:21968), supported by UniProt/HAMAP rule MF_00110, conserved domain architecture (TIGR01357 aroB, Pfam PF01761), and KT2440 literature mapping the locus to this activity.
    action: ACCEPT
- term:
    id: GO:0005737
    label: cytoplasm
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: located_in
  review:
    summary: DHQS acts on soluble cytosolic metabolites (DAHP, NAD+, divalent cation) and is a soluble intracellular enzyme of central metabolism. Consistent with UniProt subcellular location (cytoplasm) and the general architecture of the bacterial shikimate pathway. No experimental KT2440 localization assay exists, but the inference is well supported.
    action: ACCEPT
- term:
    id: GO:0009073
    label: aromatic amino acid biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: DHQS provides the second step of the shikimate pathway that supplies chorismate, the precursor of phenylalanine, tyrosine, and tryptophan. This biological process annotation is correct and represents a core role of the gene.
    action: ACCEPT
- term:
    id: GO:0009423
    label: chorismate biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: DHQS catalyzes step 2 of 7 in chorismate biosynthesis from D-erythrose 4-phosphate and phosphoenolpyruvate (UniPathway UPA00053). This is the most precise biological-process term for the gene and is well supported.
    action: ACCEPT
- term:
    id: GO:0016838
    label: carbon-oxygen lyase activity, acting on phosphates
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: InterPro2GO-derived MF term that is a more general parent of the specific 3-dehydroquinate synthase activity (GO:0003856) already annotated. DHQS is classified under EC 4.2.3.- (carbon-oxygen lyases acting on phosphates), so the term is not wrong, but it is redundant with and less informative than the specific child term.
    action: MARK_AS_OVER_ANNOTATED
    reason: The precise activity is already captured by GO:0003856. This broad lyase grouping term adds no information beyond the specific annotation and is a less informative restatement of the same molecular function.
core_functions:
- description: 3-dehydroquinate synthase catalyzing the second step of the shikimate pathway, converting DAHP to 3-dehydroquinate with release of phosphate, using catalytic NAD(+) and a divalent metal cofactor
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: aroB/PP_5078 annotated as 3-dehydroquinate synthase activity (GO:0003856) via UniRule UR000001254 / HAMAP MF_00110, EC 4.2.3.4, RHEA:21968.
    full_text_unavailable: true
  - reference_id: PMID:12534463
    supporting_text: KT2440 genome annotation assigns PP_5078 (aroB) as 3-dehydroquinate synthase; deep research (aroB-deep-research-falcon.md) corroborates the PP_5078 to aroB / DHQS mapping and the DAHP to DHQ reaction.
    full_text_unavailable: true
  molecular_function:
    id: GO:0003856
    label: 3-dehydroquinate synthase activity
  directly_involved_in:
  - id: GO:0009423
    label: chorismate biosynthetic process
  in_complex:
  substrates:
  - id: CHEBI:58394
    label: 7-phospho-2-dehydro-3-deoxy-D-arabino-heptonate (DAHP)
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:12534463
  title: Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.
  findings:
  - statement: Genome sequence of P. putida KT2440 in which the aroB/PP_5078 locus (Q88CV2) is annotated as 3-dehydroquinate synthase.
    reference_section_type: RESULTS
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: PubMed-verified KT2440 genome paper; the source of the locus tag and gene assignment. Supports identity, not detailed enzymology.