aroC

UniProt ID: Q88LU7
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

aroC encodes chorismate synthase (CS; EC 4.2.3.5), the enzyme that catalyzes the seventh and final step of the shikimate pathway. It converts 5-enolpyruvylshikimate-3-phosphate (EPSP) into chorismate via an anti-1,4 (trans) elimination of the C-3 phosphate and the C-6 proR hydrogen, introducing a second double bond into the ring system. The reaction requires reduced flavin mononucleotide (FMNH2) as an essential cofactor, even though the overall transformation involves no net change in redox state; like other bacterial monofunctional chorismate synthases, the enzyme relies on an external supply of reduced FMN rather than reducing FMN itself. Chorismate is the central branch-point metabolite of aromatic metabolism, serving as the common precursor for the aromatic amino acids phenylalanine, tyrosine, and tryptophan, as well as for folate (via para-aminobenzoate), ubiquinone/menaquinone, and other aromatic metabolites. The protein belongs to the chorismate synthase family, assembles as a homotetramer, and acts as a soluble cytosolic enzyme of central metabolism.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0004107 chorismate synthase activity
IEA
GO_REF:0000120
ACCEPT
Summary: Core molecular function. AroC is a chorismate synthase (EC 4.2.3.5), as established by the HAMAP rule MF_00300, conserved family membership (IPR000453, PF01264, TIGR00033), and mapping to RHEA:21020. This directly represents the gene's primary catalytic activity.
GO:0005829 cytosol
IEA
GO_REF:0000118
KEEP AS NON CORE
Summary: Chorismate synthase catalyzes a soluble step of central aromatic metabolism and has no membrane-targeting features, consistent with a cytosolic localization. This is a reasonable IEA/TreeGrafter inference for a bacterial metabolic enzyme. Bacteria lack the GO cytosol/cytoplasm distinction relevant in eukaryotes, but the term is acceptable as-is; it is not a core functional annotation.
GO:0009073 aromatic amino acid family biosynthetic process
IEA
GO_REF:0000120
KEEP AS NON CORE
Summary: Chorismate produced by AroC is the direct precursor of phenylalanine, tyrosine, and tryptophan, so AroC is correctly placed in aromatic amino acid biosynthesis. This is a valid biological-process annotation, though it is broader than the gene's most specific role (chorismate biosynthesis), and chorismate also feeds non-amino-acid pathways (folate, ubiquinone). Retained as a supporting/non-core process annotation. (Note: GOA label text "aromatic amino acid biosynthetic process" is the curated synonym; the canonical GO:0009073 label is "aromatic amino acid family biosynthetic process".)
GO:0009423 chorismate biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: Most specific and accurate biological-process annotation: AroC catalyzes the terminal (step 7/7) reaction of chorismate biosynthesis (UniPathway UPA00053, UER00090). This is the core process for the gene.
GO:0010181 FMN binding
IEA
GO_REF:0000118
ACCEPT
Summary: Chorismate synthase requires and binds reduced FMN (FMNH2) as an essential cofactor for catalysis; the UniProt record annotates multiple FMN-binding residues (125-127, 237-238, 277, 292-296, 318). FMN binding is well supported and integral to the catalytic mechanism, so this is accepted as a genuine (though accessory-to-the-MF) molecular function annotation.

Core Functions

Catalyzes the terminal step of the shikimate pathway, the FMNH2-dependent anti-1,4-elimination of phosphate from 5-enolpyruvylshikimate-3-phosphate (EPSP) to form chorismate, supplying the branch-point precursor for aromatic amino acid and other aromatic metabolite biosynthesis.

Supporting Evidence:
  • GO_REF:0000120
    AroC mapped to chorismate synthase (EC 4.2.3.5; RHEA:21020) by HAMAP rule MF_00300 and conserved family membership.
  • file:PSEPK/aroC/aroC-deep-research-falcon.md
    Chorismate synthase (AroC; EC 4.2.3.5) catalyzes the terminal step of the shikimate pathway, converting EPSP into chorismate with elimination of phosphate via a 1,4-trans elimination, and requires reduced FMN (FMNH2) for catalysis.

References

TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.
  • Genome sequencing of P. putida KT2440 identified the aroC gene (PP_1830) encoding chorismate synthase as part of the organism's metabolic gene complement.

Deep Research

Asta

(aroC-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:20:43.747209

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.696)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[2] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.687)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[3] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.687)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[4] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.660)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[5] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[6] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[7] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.655)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[8] The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize

  • Authors: Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al.
  • Year: 2016
  • Venue: BMC Systems Biology
  • URL: https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  • DOI: 10.1186/s12918-016-0369-x
  • PMID: 27899149
  • PMCID: 5129634
  • Citations: 11
  • Influential citations: 1
  • Summary: CornCyc’s computational predictions are more accurate than those in MaizeCyc when compared to experimentally determined function assignments, demonstrating the relative strength of the enzymatic function assignment pipeline used to generate CornCyc.
  • Evidence snippets:
  • Snippet 1 (score: 0.652)
    > A gold standard set of protein functional annotations was generated by extracting data from UniProt [16] and BRENDA [17]. We extracted all protein sequence and annotation data from UniProt (release 2016_05) for the organism Zea mays, keeping the EC annotations only from the manually reviewed component of UniProt, while removing those annotations that had not undergone manual review. We also extracted experimentally verified protein annotations for Zea mays from BRENDA (release 2016.1). The UniProt and BRENDA annotations were then merged by matching proteins based on the database crosslinks provided by BRENDA, resulting in the union of the reviewed annotations from UniProt and the experimentally verified annotations of BRENDA with duplicates removed. The merged protein annotations were then matched to the B73 RefGen_v2 translated gene models using BLASTP based on a sequence identity cutoff of 96% and an e-value cutoff of 1e-20. We selected the top scoring hit for each protein which resulted in matches to 1,815 unique maize proteins. EC annotations for alternate isoforms were consolidated at the gene level, resulting in 1,475 experimentally verified or manually reviewed protein functional annotations across 1,450 maize genes.

[9] Prediction of Horizontally and Widely Transferred Genes in Prokaryotes

  • Authors: Yoji Nakamura
  • Year: 2018
  • Venue: Evolutionary Bioinformatics Online
  • URL: https://www.semanticscholar.org/paper/af7bde229b96609907924d1c68ea463af656efb2
  • DOI: 10.1177/1176934318810785
  • PMID: 30546254
  • PMCID: 6287321
  • Citations: 5
  • Influential citations: 1
  • Summary: A data-driven approach using massive sequence data may contribute to a broader understanding of HGT in prokaryotes, and predicted that six as-yet-uncharacterized genes were widely distributed HT genes, and therefore, will be interesting targets for evolutionary studies.
  • Evidence snippets:
  • Snippet 1 (score: 0.646)
    > Focusing on uncharacterized HT genes in the COG and KEGG databases, four gene groups (COG3209, COG1479, COG3291, and COG3791) were classified into "general function prediction only" (COG category code = R) or "function unknown" (S) according to the COG annotation. By adding two gene groups of "uncharacterized protein" (K08998 and K07062) from the KEGG annotation, a total of six were obtained as functionally ambiguous HT gene groups despite being distributed among more than 300 species. With reference to COG3209 (uncharacterized conserved protein RhaS, contains 28 RHS repeats), the genes encoded in E. coli have been suggested to be horizontally transferred from another organism. 43 Recently, RHS repeat-containing genes are reported to be involved in toxins against competitors 44 ; therefore, the genes in COG3209 could be considered as defense system genes. The functions of COG1479 (uncharacterized conserved protein, contains ParB-like and HNH nuclease domains), COG3291 (PKD repeat), and COG3791 (uncharacterized conserved protein) have yet to be examined. According to InterPro analysis, COG3791 genes have a domain of glutathione-dependent formaldehyde-activating enzyme (IPR006913); therefore, these genes might be related to formaldehyde detoxification. 45 The KO group, K08998, corresponds to COG0759 (membraneanchored protein YidD) of category M in the COG database that has previously been reported to be involved in the protein insertion process. 46 K07062 corresponds to COG1487 and is considered a toxic protein. 47 As a whole, it has to be said that the evolutionary significance of these six gene groups has not been fully realized. Conversely, these genes might be good targets for evolutionary studies in the context of HGT, providing an example of data-driven approaches from massive sequence data. 48

[10] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.646)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[11] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.646)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[12] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.645)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[13] Role of histone-lysine N-methyltransferase 2D (KMT2D) in MEK-ERK signaling-mediated epigenetic regulation: a phosphoproteomics perspective

  • Authors: Sreeshma Ravindran Kammarambath, Leona Dcunha, Athira Perunelly Gopalakrishnan, Amal Fahma, N. Krishna et al.
  • Year: 2025
  • Venue: Frontiers in Bioinformatics
  • URL: https://www.semanticscholar.org/paper/0ac0729148aff3d839e6a15984e11532e9e740f9
  • DOI: 10.3389/fbinf.2025.1683469
  • PMID: 41341998
  • PMCID: 12669113
  • Citations: 3
  • Summary: The phosphoregulatory network of Histone-lysine N-methyltransferase 2D is delineated, positioning it as a dynamic epigenetic effector modulated by MEK-ERK signaling, with broader implications for cancer and developmental disorders.
  • Evidence snippets:
  • Snippet 1 (score: 0.644)
    > Each protein was mapped to its corresponding gene symbol based on the HGNC (downloaded on 30.05.2023) and to its corresponding UniProt (13.04.2023) (UniProt, 2023) accessions using our in-built mapping tool to ensure consistent and standardized annotation. We conducted the analysis using the methodologies outlined in (Sanjeev et al., 2024). The overall workflow used in this study is outlined in Figure 1.

[14] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.643)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[15] A Genome-Wide Association Study Identifying Novel Genetic Markers of Response to Treatment with Interleukin-23 Inhibitors in Psoriasis

  • Authors: Sophia Zachari, K. Liadaki, Angeliki Planaki, E. Zafiriou, Olga Kouvarou et al.
  • Year: 2025
  • Venue: Genes
  • URL: https://www.semanticscholar.org/paper/d5f656311b54e222e7487ea32a061869b30178a1
  • DOI: 10.3390/genes16101195
  • PMID: 41153410
  • PMCID: 12564705
  • Summary: These findings provide promising pharmacogenetic markers which, upon validation in larger, independent cohorts, will enable the translation of a patient’s genotype into a response phenotype, thereby guiding clinical decisions and improving drug effectiveness.
  • Evidence snippets:
  • Snippet 1 (score: 0.638)
    > The UniProt knowledgebase (www.uniprot.org/uniprotkb/), (accessed on 20 June 2025), the central hub for the collection of functional information on proteins, with accurate and rich annotation [33], was used to retrieve the approved human gene and protein names and symbols.

[16] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.637)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[17] A customized Web portal for the genome of the ctenophore Mnemiopsis leidyi

  • Authors: R. Moreland, A. Nguyen, Joseph F. Ryan, Joseph F. Ryan, C. Schnitzler et al.
  • Year: 2014
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/327b13b7b70d702a34b70b6ef01188d8167601d8
  • DOI: 10.1186/1471-2164-15-316
  • PMID: 24773765
  • PMCID: 4234515
  • Citations: 41
  • Summary: This sequencing effort has produced the first set of whole-genome sequencing data on any ctenophore species and is amongst the first wave of projects to sequence an animal genome de novo solely using next-generation sequencing technologies.
  • Evidence snippets:
  • Snippet 1 (score: 0.633)
    > In an effort to engage the collective expertise of the scientific community, we have implemented a collaborative wiki (MediaWiki version 1.19.11) for the Mnemiopsis gene complement. The Mnemiopsis Gene Wiki is accessible from the left sidebar of most pages and is searchable either by selecting a Mnemiopsis gene identifier (e.g., ML00011a) from the drop-down menu or by manually entering an identifier in the appropriate search box. Users can also access these pages by clicking on a gene in the 2.2 track of the genome browser. Each record in the Gene Wiki represents a single Mnemiopsis gene and provides the following annotation: nucleotide and protein sequences, coding exonic genomic coordinates, pre-computed BLAST hits from numerous organisms displaying the top hits for each protein, the top non-self BLAST hit to Mnemiopsis, Pfam-A domains, Gene Ontology (GO) functional annotations, human disease genes from Online Mendelian Inheritance in Man (OMIM), and a table of ortholog clusters formed by phylogenetically informed clustering methods [4] (Figure 5). In addition, controlled editable sections have been included that permit (and encourage) the scientific community to provide further gene annotation for isoforms, in situ images, references, and other notes for each gene. Users interested in supplementing our gene model annotation at the Mnemiopsis Gene Wiki pages must first create an account and log in prior to submitting their contributions. In-house subject matter expert data curators are notified by e-mail following the creation of a new user account or an edit to an existing Gene Wiki record. Any content changes or additions to the Gene Wiki are thoroughly evaluated by these data curators and are made public subject to their approval.
    > Pre-compiled BLAST hits are enumerated in tabular form. Each Mnemiopsis protein was compared to the UniProt and NCBI non-redundant protein databases (nr) using BLASTP. The results display the hit number, the accession numbers, E-values, and brief descriptions of the top four hits (lowest E-values). Accession numbers are linked to relevant corresponding entries at UniProt and GenBank.

[18] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.633)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[19] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.631)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].
  • Authors: Jaap van der Heijden, Asanda Mazubane, Marko Sallisalmi, E. Vorontsov, J. Tenhunen et al.
  • Year: 2025
  • Venue: Clinical Proteomics
  • URL: https://www.semanticscholar.org/paper/ae917565de913693a2713b46c3eda672d9d01c7f
  • DOI: 10.1186/s12014-025-09556-2
  • PMID: 40885913
  • PMCID: 12398169
  • Summary: The altered proteomic profile of hyaluronan-related proteins as reflected by the GO terms indicates a complex dysregulation not only in hyaluronan metabolism and extracellular matrix, but also in the regulation of several proteolytic enzymes.
  • Evidence snippets:
  • Snippet 1 (score: 0.631)
    > The identification of hyaluronan-associated genes was performed using Python 3.10.12. The UniProt REST Web Application Programming Interface (API) was used to query and retrieve gene annotations that met specific criteria (UniProt Consortium). A keyword-based query, utilizing the terms "hyaluronan, " "hyaluronic acid, " "hyaluronidase, " "hyaluronic acid synthase, " "hyaluronate binding protein, " "hyaluronan oligosaccharides, " "hyaluronan synthase, " and "hyaluronan receptor, " was executed across our dataset of 663 genes. Gene symbols were returned as "hits" when the query keywords were found in the annotations, including functional comments, Gene Ontology (GO) terms, and cross-referenced databases, formatted in JSON.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  2. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  3. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  4. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  5. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  6. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  7. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  8. Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al. (2016). The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize. BMC Systems Biology. https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  9. Yoji Nakamura (2018). Prediction of Horizontally and Widely Transferred Genes in Prokaryotes. Evolutionary Bioinformatics Online. https://www.semanticscholar.org/paper/af7bde229b96609907924d1c68ea463af656efb2
  10. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  11. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  12. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  13. Sreeshma Ravindran Kammarambath, Leona Dcunha, Athira Perunelly Gopalakrishnan, Amal Fahma, N. Krishna et al. (2025). Role of histone-lysine N-methyltransferase 2D (KMT2D) in MEK-ERK signaling-mediated epigenetic regulation: a phosphoproteomics perspective. Frontiers in Bioinformatics. https://www.semanticscholar.org/paper/0ac0729148aff3d839e6a15984e11532e9e740f9
  14. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  15. Sophia Zachari, K. Liadaki, Angeliki Planaki, E. Zafiriou, Olga Kouvarou et al. (2025). A Genome-Wide Association Study Identifying Novel Genetic Markers of Response to Treatment with Interleukin-23 Inhibitors in Psoriasis. Genes. https://www.semanticscholar.org/paper/d5f656311b54e222e7487ea32a061869b30178a1
  16. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  17. R. Moreland, A. Nguyen, Joseph F. Ryan, Joseph F. Ryan, C. Schnitzler et al. (2014). A customized Web portal for the genome of the ctenophore Mnemiopsis leidyi. BMC Genomics. https://www.semanticscholar.org/paper/327b13b7b70d702a34b70b6ef01188d8167601d8
  18. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  19. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  20. Jaap van der Heijden, Asanda Mazubane, Marko Sallisalmi, E. Vorontsov, J. Tenhunen et al. (2025). Plasma proteomics in septic shock and alcohol-related pancreatitis: a hyaluronan-centered approach. Clinical Proteomics. https://www.semanticscholar.org/paper/ae917565de913693a2713b46c3eda672d9d01c7f

Falcon

(aroC-deep-research-falcon.md)
Research Report: Functional Annotation of **aroC** (Chorismate synthase) in *Pseudomonas putida* KT2440 (UniProt Q88LU7; PP_1830) Falcon Edison Scientific Literature 30 citations 1 artifacts 2026-06-11T20:49:16.310727

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research Report: Functional Annotation of aroC (Chorismate synthase) in Pseudomonas putida KT2440 (UniProt Q88LU7; PP_1830)

1. Target verification (gene/protein identity)

The research target is AroC (chorismate synthase; EC 4.2.3.5) encoded by aroC and mapped (per the provided UniProt record) to ordered locus PP_1830 in Pseudomonas putida strain KT2440 (UniProt Q88LU7). The retrieved literature that mentions “aroC” in Pseudomonas metabolic engineering uses it consistently as a shikimate-pathway gene encoding chorismate synthase (e.g., included as a shikimate-pathway overexpression target in P. putida KT2440) and not as an unrelated gene, supporting correct symbol usage in this organism context (camposmagana2024combinatorialengineeringreveals pages 4-7).

Limitation: within the retrieved full texts, there is little direct primary characterization explicitly citing PP_1830/Q88LU7; therefore, organism-specific biochemical details for this exact protein are inferred from conserved AroC enzymology (family-level evidence) and P. putida KT2440 pathway-engineering studies.

2. Key concepts and definitions (current understanding)

2.1 Chorismate synthase (AroC): definition and core biochemical role

Chorismate synthase (AroC; EC 4.2.3.5) catalyzes the terminal step of the shikimate pathway, converting 5-enolpyruvylshikimate-3-phosphate (EPSP) into chorismate with elimination of inorganic phosphate (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13). The enzymatic transformation is described as a 1,4-trans elimination of the phosphate group from EPSP, which introduces a second double bond into the six-membered ring system to yield chorismate (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13).

Because chorismate is the key branch-point metabolite at the end of the shikimate pathway, AroC function is tightly tied to biosynthesis of aromatic amino acids and other chorismate-derived products (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13).

2.2 Substrate specificity and reaction stoichiometry

The canonical substrate is EPSP, and the product is chorismate + phosphate (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13). Structural/biochemical work in other taxa (e.g., fungal chorismate synthase) explicitly models/binds EPSP together with FMNH2 at the active site and uses EPSP in enzyme assays, reinforcing EPSP as the functional substrate in chorismate synthases across organisms (rodriguesvendramini2019promisingnewantifungal pages 2-5).

2.3 Cofactor requirements and mechanistic features

Although the overall EPSP→chorismate conversion is an elimination with no net redox change, chorismate synthase requires reduced flavin mononucleotide (FMNH2) for catalysis (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13). A 2024 authoritative review summarizes that monofunctional chorismate synthases (e.g., from E. coli and higher plants) are unable to reduce FMN themselves and therefore depend on externally supplied reduced FMN, whereas some organisms (e.g., fungi/protozoa) have bifunctional enzymes capable of reducing FMN using NADPH, changing how FMNH2 is supplied (shende2024theshikimatepathway pages 11-13).

2.4 Quaternary structure and enzyme family context

Chorismate synthase is often tetrameric. For example, a mycobacterial chorismate synthase structure was solved with bound FMN and described as a tetramer (dimer of dimers) (nunes2020mycobacteriumtuberculosisshikimate pages 20-23). A fungal chorismate synthase was modeled as a homotetramer for docking/MD stability and ligand binding (rodriguesvendramini2019promisingnewantifungal pages 2-5). This supports a common structural theme that is relevant for conserved function.

3. Biological processes, pathways, and cellular location (functional annotation)

3.1 Pathway context: shikimate pathway and chorismate branch point

The shikimate pathway comprises seven steps that convert central carbon precursors (classically PEP and E4P) into chorismate, which then branches into multiple essential biosynthetic routes (nunes2020mycobacteriumtuberculosisshikimate pages 3-7). The product chorismate is a central node feeding the biosynthesis of phenylalanine, tyrosine, and tryptophan, and also contributes to other pathways, such as PABA/folate-related metabolism and quinone-related metabolites in microbes (guida2024aminoacidbiosynthesis pages 1-2, nunes2020mycobacteriumtuberculosisshikimate pages 3-7).

In P. putida KT2440 specifically, multiple engineering studies rely on increasing shikimate/chorismate supply to produce aromatic chemicals, demonstrating that chorismate availability (and therefore AroC activity upstream) is a key determinant of metabolic output in this chassis (yu2016metabolicengineeringof pages 1-3, dias2023fromdegraderto pages 8-11, camposmagana2024combinatorialengineeringreveals pages 7-11).

3.2 Likely subcellular localization

No direct experimental localization for AroC (PP_1830/Q88LU7) was retrieved from the available texts. However, given that chorismate synthase catalyzes a soluble step of central metabolism and the retrieved mechanistic/structural literature treats it as a soluble enzyme (with defined quaternary structure and bound flavin), the most consistent working annotation is that AroC functions as a cytosolic enzyme in bacteria.

Evidence limitation: because no explicit localization measurement for this P. putida enzyme was retrieved, the cytosolic localization should be treated as a reasonable inference rather than directly demonstrated in KT2440 within the cited sources.

4. Recent developments and latest research (prioritizing 2023–2024)

4.1 2024 state-of-the-art synthesis of shikimate-pathway enzymology

A 2024 Natural Product Reports review (covering 1997–2023) provides an up-to-date synthesis of shikimate pathway enzymology and highlights that chorismate synthase has been interrogated using kinetic isotope and structural biology methods, reiterating (i) the EPSP→chorismate phosphate-elimination chemistry and (ii) the essential requirement for FMNH2 (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13). This is an authoritative, high-citation review source suitable for current “textbook-level” mechanistic statements.

4.2 2024 P. putida combinatorial tuning of shikimate genes (aroC included)

A 2024 bioRxiv preprint used a design-of-experiments (Plackett–Burman) + linear modeling approach to tune expression of nine factors in P. putida shikimate and pABA biosynthesis, explicitly including aroC among the genes amplified from KT2440 and tested at two expression levels (camposmagana2024combinatorialengineeringreveals pages 4-7). The authors report pABA titers ranging from ~2 mg/L to ~232 mg/L across their engineering cycle, with a reported best yield of 0.024 mol/mol glucose (camposmagana2024combinatorialengineeringreveals pages 7-11). Critically for aroC functional annotation in a metabolic-engineering context, they report that high aroC expression can be unfavorable and conclude that mild aroC overexpression is preferable—consistent with aroC not necessarily being the dominant bottleneck compared with other steps like aroB (camposmagana2024combinatorialengineeringreveals pages 7-11).

4.3 2023 P. putida repurposing to produce gallic acid from glycerol

A 2023 study engineered P. putida KT2440 to produce gallic acid from glycerol and reported 346.7 ± 0.004 mg/L final concentration after 72 h and an observed yield of 0.12 g gallic acid per g glycerol (dias2023fromdegraderto pages 8-11). While the work did not directly manipulate aroC, it demonstrates the importance of driving flux through the shikimate/chorismate node and highlights NADPH-related considerations for high-yield routes (dias2023fromdegraderto pages 8-11).

5. Current applications and real-world implementations

5.1 Industrial biotechnology: routing chorismate to commodity and specialty aromatics

Pseudomonas putida KT2440 is used as a robust host for producing aromatic chemicals via chorismate-derived routes. A prominent example is production of para-hydroxybenzoic acid (PHBA) by expressing E. coli ubiC (chorismate lyase) to convert chorismate into PHBA, combined with a feedback-resistant DAHP synthase (aroG D146N) and deletion of competing chorismate-consuming/degrading pathways. This strategy achieved 1.73 g/L PHBA with a carbon yield of 18.1% (C-mol/C-mol) in fed-batch fermentation (yu2016metabolicengineeringof pages 1-3). These data show that maintaining sufficient chorismate supply (which depends upstream on AroC) is central to industrial aromatic production in KT2440.

5.2 Synthetic biology “tuning” view of aroC

The 2024 DoE study implies that, at least for pABA production in KT2440, aroC expression requires balance: overexpression beyond a mild level may reduce performance (camposmagana2024combinatorialengineeringreveals pages 7-11). This kind of result is useful for functional annotation because it indicates aroC is embedded in a network where enzyme overabundance can create cofactor demands (e.g., reduced FMN supply), imbalances, or metabolic burden, even if the enzyme is essential for producing chorismate.

6. Expert opinions and analysis from authoritative sources

6.1 Why chorismate/shikimate metabolism is a key microbial vulnerability

A 2024 TB-focused review reiterates the selective-targeting rationale for shikimate pathway enzymes: the pathway is absent in humans and essential in certain pathogens such as Mycobacterium tuberculosis, and chorismate is a metabolic node for aromatic amino acids and other essential metabolites (guida2024aminoacidbiosynthesis pages 1-2). This expert framing supports the broad interpretation that enzymes upstream of chorismate (including AroC) underpin essential biosynthetic capacity in many microbes.

A 2024 Applied and Environmental Microbiology paper further validates the idea of targeting chorismate-utilizing enzymes (anthranilate synthase and aminodeoxychorismate synthase) as antibiotic targets, emphasizing desirable target features such as essentiality and lack of human homologs, and proposing that disrupting conserved protein–protein interactions could provide robust antibacterial strategies (funke2024validationofaminodeoxychorismate pages 1-2). While not about AroC directly, it illustrates the ongoing 2024 research emphasis on the chorismate node as an antimicrobial intervention point.

7. Relevant statistics and quantitative data from recent studies

7.1 Metabolic engineering performance metrics in P. putida KT2440

  • PHBA production (chorismate-derived): 1.73 g/L titer; 18.1% C-mol/C-mol yield in non-optimized fed-batch (2016; still a widely cited benchmark for KT2440) (yu2016metabolicengineeringof pages 1-3).
  • pABA production: ~2–232 mg/L titers reported in 2024 DoE study; best yield 0.024 mol/mol glucose; high aroC overexpression reported as unfavorable with recommendation for mild tuning (camposmagana2024combinatorialengineeringreveals pages 7-11, camposmagana2024combinatorialengineeringreveals pages 4-7).
  • Gallic acid production: 346.7 ± 0.004 mg/L; observed yield 0.12 g/g glycerol (2023) (dias2023fromdegraderto pages 8-11).

7.2 Infectious disease / AMR burden statistics referenced in 2024 literature

  • A 2024 TB drug-discovery review reports 10.6 million TB diagnoses and 1.30 million deaths in 2022, and ~410,000 MDR/RR-TB cases in 2022, providing the clinical motivation for pursuing new antibacterial targets including pathways like shikimate metabolism (guida2024aminoacidbiosynthesis pages 1-2).
  • A 2024 MRSA-focused shikimate-kinase inhibitor paper cites 1.27 million deaths in 2019 attributed to antimicrobial resistance (AMR) (riossoto2024inhibitionofshikimate pages 1-2).

8. Functional annotation summary for UniProt Q88LU7 (aroC; PP_1830)

Primary molecular function: Chorismate synthase (AroC) catalyzes EPSP → chorismate + phosphate via 1,4-trans elimination; requires reduced FMN (FMNH2) for catalysis (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13).

Biological process context: terminal step of the shikimate pathway supplying chorismate, which serves as the precursor branch point for aromatic amino acids and other essential aromatic metabolites (guida2024aminoacidbiosynthesis pages 1-2, nunes2020mycobacteriumtuberculosisshikimate pages 3-7).

Cellular location (best-supported annotation from available evidence): likely cytosolic, but direct localization evidence for PP_1830/Q88LU7 in KT2440 was not retrieved in the available texts.

Application relevance: In P. putida KT2440, engineering of chorismate supply is central to bioproduction of aromatic products such as PHBA, pABA, and gallic acid; studies indicate that balanced expression of shikimate enzymes (including aroC) can matter for yield and titer (yu2016metabolicengineeringof pages 1-3, camposmagana2024combinatorialengineeringreveals pages 7-11, dias2023fromdegraderto pages 8-11).


Summary table

Aspect Summary
Gene/protein identity aroC; UniProt Q88LU7; ordered locus PP_1830; organism Pseudomonas putida strain KT2440; annotated as a chorismate synthase family protein in the user-supplied UniProt record. Direct organism-specific primary literature on PP_1830 itself appears limited, so some functional details below are inferred from conserved chorismate synthase biochemistry and family-level evidence plus P. putida pathway-engineering studies (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13, camposmagana2024combinatorialengineeringreveals pages 4-7).
Enzyme name / EC Chorismate synthase; EC 4.2.3.5; terminal enzyme of the shikimate pathway in bacteria (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13).
Reaction catalyzed 5-enolpyruvylshikimate-3-phosphate (EPSP) → chorismate + phosphate; mechanistically described as a 1,4-trans elimination of phosphate from EPSP, creating the second double bond in the ring system of chorismate (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13, rodriguesvendramini2019promisingnewantifungal pages 2-5).
Required cofactor(s) Requires reduced FMN (FMNH2) for activity even though the overall reaction is not net redox. In many bacterial/plant monofunctional enzymes, FMN is reduced externally (often by a separate flavin reductase or reduced flavin supply), whereas bifunctional enzymes in some fungi/protozoa can reduce FMN themselves with NADPH (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13).
Pathway role Catalyzes the terminal step from EPSP to chorismate in the shikimate pathway. Chorismate is the major branch-point precursor for phenylalanine, tyrosine, tryptophan, and additional metabolites such as PABA/folate, ubiquinone/menaquinone-related products, and other aromatic metabolites (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13, guida2024aminoacidbiosynthesis pages 1-2, nunes2020mycobacteriumtuberculosisshikimate pages 3-7).
Quaternary structure notes Chorismate synthase is commonly reported as a tetramer; structural work on non-P. putida orthologs describes a homotetramer or tetrameric assembly important for ligand binding/stability (rodriguesvendramini2019promisingnewantifungal pages 2-5, nunes2020mycobacteriumtuberculosisshikimate pages 20-23).
Likely localization No direct localization evidence for PP_1830/Q88LU7 was retrieved here. Given its role in central aromatic biosynthesis and lack of membrane-targeting evidence in the gathered sources, the enzyme is most reasonably treated as a cytosolic bacterial metabolic enzyme, but this point should be considered an inference rather than a directly sourced P. putida measurement.
P. putida application: PHBA production In P. putida KT2440, shikimate/chorismate flux was engineered toward para-hydroxybenzoic acid (PHBA) by expressing E. coli ubiC and feedback-resistant aroG D146N, with deletions of pobA, pheA, trpE, and hexR. Best reported performance: 1.73 g/L PHBA and 18.1% C-mol/C-mol carbon yield in non-optimized fed-batch fermentation (yu2016metabolicengineeringof pages 5-6, yu2016metabolicengineeringof pages 1-3).
P. putida application: pABA production / aroC tuning A 2024 combinatorial-expression study in P. putida found pABA titers spanning ~2–232 mg/L. High aroC overexpression was reported to have an unfavorable or only mildly beneficial effect; the authors concluded mild rather than maximal aroC expression was preferable, while aroB emerged as a stronger bottleneck (camposmagana2024combinatorialengineeringreveals pages 7-11, camposmagana2024combinatorialengineeringreveals pages 4-7).
P. putida application: gallic acid production Engineering of shikimate-pathway flux in P. putida KT2440 for gallic acid production from glycerol achieved 346.7 ± 0.004 mg/L gallic acid and an observed yield of 0.12 g/g glycerol; this study did not specifically manipulate aroC, but it demonstrates practical importance of chorismate-pathway flux control in this host (dias2023fromdegraderto pages 8-11).

Table: This table summarizes the verified identity, core enzymology, pathway context, and recent Pseudomonas putida engineering relevance of aroC/chorismate synthase. It is useful as a compact evidence-backed reference for the final research report.

URLs and publication dates for key cited sources

  • Shende VV, Bauman KD, Moore BS. “The shikimate pathway: gateway to metabolic diversity.” Natural Product Reports (Jan 2024). https://doi.org/10.1039/d3np00037k (shende2024theshikimatepathway pages 10-11, shende2024theshikimatepathway pages 11-13)
  • Campos-Magaña MA et al. “Combinatorial engineering reveals shikimate pathway bottlenecks in para-aminobenzoic acid production in Pseudomonas putida.” bioRxiv (Jun 2024). https://doi.org/10.1101/2024.06.17.599342 (camposmagana2024combinatorialengineeringreveals pages 7-11, camposmagana2024combinatorialengineeringreveals pages 4-7)
  • Dias FMS et al. “From degrader to producer: reversing the gallic acid metabolism of Pseudomonas putida KT2440.” International Microbiology (Nov 2023). https://doi.org/10.1007/s10123-022-00282-5 (dias2023fromdegraderto pages 8-11)
  • Funke FJ, Schlee S, Sterner R. “Validation of aminodeoxychorismate synthase and anthranilate synthase as novel targets…” Applied and Environmental Microbiology (May 2024). https://doi.org/10.1128/aem.00572-24 (funke2024validationofaminodeoxychorismate pages 1-2)
  • Guida M et al. “Amino Acid Biosynthesis Inhibitors in Tuberculosis Drug Discovery.” Pharmaceutics (May 2024). https://doi.org/10.3390/pharmaceutics16060725 (guida2024aminoacidbiosynthesis pages 1-2, guida2024aminoacidbiosynthesis pages 2-4)
  • Rios-Soto L et al. “Inhibition of Shikimate Kinase from MRSA by Benzimidazole Derivatives…” International Journal of Molecular Sciences (May 2024). https://doi.org/10.3390/ijms25105077 (riossoto2024inhibitionofshikimate pages 1-2)
  • Yu S et al. “Metabolic Engineering of Pseudomonas putida KT2440 for the Production of para-Hydroxy Benzoic Acid.” Frontiers in Bioengineering and Biotechnology (Nov 2016). https://doi.org/10.3389/fbioe.2016.00090 (yu2016metabolicengineeringof pages 1-3, yu2016metabolicengineeringof pages 5-6)

References

  1. (camposmagana2024combinatorialengineeringreveals pages 4-7): Marco A Campos-Magaña, Sara Moreno-Paz, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering reveals shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. bioRxiv, Jun 2024. URL: https://doi.org/10.1101/2024.06.17.599342, doi:10.1101/2024.06.17.599342. This article has 0 citations.

  2. (shende2024theshikimatepathway pages 10-11): Vikram V. Shende, Katherine D. Bauman, and Bradley S. Moore. The shikimate pathway: gateway to metabolic diversity. Natural product reports, 41:604-648, Jan 2024. URL: https://doi.org/10.1039/d3np00037k, doi:10.1039/d3np00037k. This article has 173 citations and is from a peer-reviewed journal.

  3. (shende2024theshikimatepathway pages 11-13): Vikram V. Shende, Katherine D. Bauman, and Bradley S. Moore. The shikimate pathway: gateway to metabolic diversity. Natural product reports, 41:604-648, Jan 2024. URL: https://doi.org/10.1039/d3np00037k, doi:10.1039/d3np00037k. This article has 173 citations and is from a peer-reviewed journal.

  4. (rodriguesvendramini2019promisingnewantifungal pages 2-5): Franciele Abigail Vilugron Rodrigues-Vendramini, Cidnei Marschalk, Marina Toplak, Peter Macheroux, Patricia de Souza Bonfim-Mendonça, Terezinha Inez Estivalet Svidzinski, Flavio Augusto Vicente Seixas, and Erika Seki Kioshima. Promising new antifungal treatment targeting chorismate synthase from paracoccidioides brasiliensis. Antimicrobial Agents and Chemotherapy, Jan 2019. URL: https://doi.org/10.1128/aac.01097-18, doi:10.1128/aac.01097-18. This article has 30 citations and is from a highest quality peer-reviewed journal.

  5. (nunes2020mycobacteriumtuberculosisshikimate pages 20-23): José E. S. Nunes, Mario A. Duque, Talita F. de Freitas, Luiza Galina, Luis F. S. M. Timmers, Cristiano V. Bizarro, Pablo Machado, Luiz A. Basso, and Rodrigo G. Ducati. Mycobacterium tuberculosis shikimate pathway enzymes as targets for the rational design of anti-tuberculosis drugs. Molecules, 25:1259, Mar 2020. URL: https://doi.org/10.3390/molecules25061259, doi:10.3390/molecules25061259. This article has 74 citations.

  6. (nunes2020mycobacteriumtuberculosisshikimate pages 3-7): José E. S. Nunes, Mario A. Duque, Talita F. de Freitas, Luiza Galina, Luis F. S. M. Timmers, Cristiano V. Bizarro, Pablo Machado, Luiz A. Basso, and Rodrigo G. Ducati. Mycobacterium tuberculosis shikimate pathway enzymes as targets for the rational design of anti-tuberculosis drugs. Molecules, 25:1259, Mar 2020. URL: https://doi.org/10.3390/molecules25061259, doi:10.3390/molecules25061259. This article has 74 citations.

  7. (guida2024aminoacidbiosynthesis pages 1-2): Michela Guida, Chiara Tammaro, Miriana Quaranta, Benedetta Salvucci, Mariangela Biava, Giovanna Poce, and Sara Consalvi. Amino acid biosynthesis inhibitors in tuberculosis drug discovery. Pharmaceutics, 16:725, May 2024. URL: https://doi.org/10.3390/pharmaceutics16060725, doi:10.3390/pharmaceutics16060725. This article has 3 citations.

  8. (yu2016metabolicengineeringof pages 1-3): Shiqin Yu, Manuel R. Plan, Gal Winter, and Jens O. Krömer. Metabolic engineering of pseudomonas putida kt2440 for the production of para-hydroxy benzoic acid. Frontiers in Bioengineering and Biotechnology, Nov 2016. URL: https://doi.org/10.3389/fbioe.2016.00090, doi:10.3389/fbioe.2016.00090. This article has 76 citations.

  9. (dias2023fromdegraderto pages 8-11): Felipe M. S. Dias, Raoní K. Pantoja, José Gregório C. Gomez, and Luiziana F. Silva. From degrader to producer: reversing the gallic acid metabolism of pseudomonas putida kt2440. International Microbiology, 26:243-255, Nov 2023. URL: https://doi.org/10.1007/s10123-022-00282-5, doi:10.1007/s10123-022-00282-5. This article has 7 citations and is from a peer-reviewed journal.

  10. (camposmagana2024combinatorialengineeringreveals pages 7-11): Marco A Campos-Magaña, Sara Moreno-Paz, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering reveals shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. bioRxiv, Jun 2024. URL: https://doi.org/10.1101/2024.06.17.599342, doi:10.1101/2024.06.17.599342. This article has 0 citations.

  11. (funke2024validationofaminodeoxychorismate pages 1-2): Franziska Jasmin Funke, Sandra Schlee, and Reinhard Sterner. Validation of aminodeoxychorismate synthase and anthranilate synthase as novel targets for bispecific antibiotics inhibiting conserved protein-protein interactions. Applied and Environmental Microbiology, May 2024. URL: https://doi.org/10.1128/aem.00572-24, doi:10.1128/aem.00572-24. This article has 5 citations and is from a peer-reviewed journal.

  12. (riossoto2024inhibitionofshikimate pages 1-2): Lluvia Rios-Soto, Alicia Hernández-Campos, David Tovar-Escobar, Rafael Castillo, Erick Sierra-Campos, Mónica Valdez-Solana, Alfredo Téllez-Valencia, and Claudia Avitia-Domínguez. Inhibition of shikimate kinase from methicillin-resistant staphylococcus aureus by benzimidazole derivatives. kinetic, computational, toxicological, and biological activity studies. International Journal of Molecular Sciences, 25:5077, May 2024. URL: https://doi.org/10.3390/ijms25105077, doi:10.3390/ijms25105077. This article has 8 citations.

  13. (yu2016metabolicengineeringof pages 5-6): Shiqin Yu, Manuel R. Plan, Gal Winter, and Jens O. Krömer. Metabolic engineering of pseudomonas putida kt2440 for the production of para-hydroxy benzoic acid. Frontiers in Bioengineering and Biotechnology, Nov 2016. URL: https://doi.org/10.3389/fbioe.2016.00090, doi:10.3389/fbioe.2016.00090. This article has 76 citations.

  14. (guida2024aminoacidbiosynthesis pages 2-4): Michela Guida, Chiara Tammaro, Miriana Quaranta, Benedetta Salvucci, Mariangela Biava, Giovanna Poce, and Sara Consalvi. Amino acid biosynthesis inhibitors in tuberculosis drug discovery. Pharmaceutics, 16:725, May 2024. URL: https://doi.org/10.3390/pharmaceutics16060725, doi:10.3390/pharmaceutics16060725. This article has 3 citations.

Artifacts

Citations

  1. camposmagana2024combinatorialengineeringreveals pages 4-7
  2. rodriguesvendramini2019promisingnewantifungal pages 2-5
  3. shende2024theshikimatepathway pages 11-13
  4. nunes2020mycobacteriumtuberculosisshikimate pages 20-23
  5. nunes2020mycobacteriumtuberculosisshikimate pages 3-7
  6. camposmagana2024combinatorialengineeringreveals pages 7-11
  7. dias2023fromdegraderto pages 8-11
  8. yu2016metabolicengineeringof pages 1-3
  9. guida2024aminoacidbiosynthesis pages 1-2
  10. funke2024validationofaminodeoxychorismate pages 1-2
  11. riossoto2024inhibitionofshikimate pages 1-2
  12. shende2024theshikimatepathway pages 10-11
  13. yu2016metabolicengineeringof pages 5-6
  14. guida2024aminoacidbiosynthesis pages 2-4
  15. https://doi.org/10.1039/d3np00037k
  16. https://doi.org/10.1101/2024.06.17.599342
  17. https://doi.org/10.1007/s10123-022-00282-5
  18. https://doi.org/10.1128/aem.00572-24
  19. https://doi.org/10.3390/pharmaceutics16060725
  20. https://doi.org/10.3390/ijms25105077
  21. https://doi.org/10.3389/fbioe.2016.00090
  22. https://doi.org/10.1101/2024.06.17.599342,
  23. https://doi.org/10.1039/d3np00037k,
  24. https://doi.org/10.1128/aac.01097-18,
  25. https://doi.org/10.3390/molecules25061259,
  26. https://doi.org/10.3390/pharmaceutics16060725,
  27. https://doi.org/10.3389/fbioe.2016.00090,
  28. https://doi.org/10.1007/s10123-022-00282-5,
  29. https://doi.org/10.1128/aem.00572-24,
  30. https://doi.org/10.3390/ijms25105077,

📄 View Raw YAML

id: Q88LU7
gene_symbol: aroC
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: >-
  aroC encodes chorismate synthase (CS; EC 4.2.3.5), the enzyme that catalyzes
  the seventh and final step of the shikimate pathway. It converts
  5-enolpyruvylshikimate-3-phosphate (EPSP) into chorismate via an anti-1,4
  (trans) elimination of the C-3 phosphate and the C-6 proR hydrogen, introducing
  a second double bond into the ring system. The reaction requires reduced flavin
  mononucleotide (FMNH2) as an essential cofactor, even though the overall
  transformation involves no net change in redox state; like other bacterial
  monofunctional chorismate synthases, the enzyme relies on an external supply of
  reduced FMN rather than reducing FMN itself. Chorismate is the central
  branch-point metabolite of aromatic metabolism, serving as the common precursor
  for the aromatic amino acids phenylalanine, tyrosine, and tryptophan, as well as
  for folate (via para-aminobenzoate), ubiquinone/menaquinone, and other aromatic
  metabolites. The protein belongs to the chorismate synthase family, assembles as
  a homotetramer, and acts as a soluble cytosolic enzyme of central metabolism.
existing_annotations:
- term:
    id: GO:0004107
    label: chorismate synthase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: >-
      Core molecular function. AroC is a chorismate synthase (EC 4.2.3.5), as
      established by the HAMAP rule MF_00300, conserved family membership
      (IPR000453, PF01264, TIGR00033), and mapping to RHEA:21020. This directly
      represents the gene's primary catalytic activity.
    action: ACCEPT
- term:
    id: GO:0005829
    label: cytosol
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: >-
      Chorismate synthase catalyzes a soluble step of central aromatic
      metabolism and has no membrane-targeting features, consistent with a
      cytosolic localization. This is a reasonable IEA/TreeGrafter inference for a
      bacterial metabolic enzyme. Bacteria lack the GO cytosol/cytoplasm
      distinction relevant in eukaryotes, but the term is acceptable as-is; it is
      not a core functional annotation.
    action: KEEP_AS_NON_CORE
- term:
    id: GO:0009073
    label: aromatic amino acid family biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: >-
      Chorismate produced by AroC is the direct precursor of phenylalanine,
      tyrosine, and tryptophan, so AroC is correctly placed in aromatic amino acid
      biosynthesis. This is a valid biological-process annotation, though it is
      broader than the gene's most specific role (chorismate biosynthesis), and
      chorismate also feeds non-amino-acid pathways (folate, ubiquinone). Retained
      as a supporting/non-core process annotation. (Note: GOA label text
      "aromatic amino acid biosynthetic process" is the curated synonym; the
      canonical GO:0009073 label is "aromatic amino acid family biosynthetic
      process".)
    action: KEEP_AS_NON_CORE
- term:
    id: GO:0009423
    label: chorismate biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: >-
      Most specific and accurate biological-process annotation: AroC catalyzes the
      terminal (step 7/7) reaction of chorismate biosynthesis (UniPathway
      UPA00053, UER00090). This is the core process for the gene.
    action: ACCEPT
- term:
    id: GO:0010181
    label: FMN binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: enables
  review:
    summary: >-
      Chorismate synthase requires and binds reduced FMN (FMNH2) as an essential
      cofactor for catalysis; the UniProt record annotates multiple FMN-binding
      residues (125-127, 237-238, 277, 292-296, 318). FMN binding is well
      supported and integral to the catalytic mechanism, so this is accepted as a
      genuine (though accessory-to-the-MF) molecular function annotation.
    action: ACCEPT
core_functions:
- description: >-
    Catalyzes the terminal step of the shikimate pathway, the FMNH2-dependent
    anti-1,4-elimination of phosphate from 5-enolpyruvylshikimate-3-phosphate
    (EPSP) to form chorismate, supplying the branch-point precursor for aromatic
    amino acid and other aromatic metabolite biosynthesis.
  molecular_function:
    id: GO:0004107
    label: chorismate synthase activity
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: >-
      AroC mapped to chorismate synthase (EC 4.2.3.5; RHEA:21020) by HAMAP rule
      MF_00300 and conserved family membership.
  - reference_id: file:PSEPK/aroC/aroC-deep-research-falcon.md
    supporting_text: >-
      Chorismate synthase (AroC; EC 4.2.3.5) catalyzes the terminal step of the
      shikimate pathway, converting EPSP into chorismate with elimination of
      phosphate via a 1,4-trans elimination, and requires reduced FMN (FMNH2) for
      catalysis.
  directly_involved_in:
  - id: GO:0009423
    label: chorismate biosynthetic process
references:
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:12534463
  title: >-
    Complete genome sequence and comparative analysis of the metabolically
    versatile Pseudomonas putida KT2440.
  findings:
  - statement: >-
      Genome sequencing of P. putida KT2440 identified the aroC gene (PP_1830)
      encoding chorismate synthase as part of the organism's metabolic gene
      complement.
    reference_section_type: RESULTS
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: >-
      PubMed-verified genome paper for P. putida KT2440; source of the
      PP_1830/aroC gene assignment. Does not provide direct biochemical
      characterization of the AroC protein.