aroK

UniProt ID: Q88CV1
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

AroK is the shikimate kinase (EC 2.7.1.71) of Pseudomonas putida KT2440. It catalyzes the fifth step of the shikimate pathway, the ATP-dependent phosphorylation of the 3-hydroxyl group of shikimate to form shikimate 3-phosphate (3-phosphoshikimate), with release of ADP and a proton. The enzyme requires a Mg(2+) cofactor (one ion bound per subunit) and adopts the P-loop (Walker A) NTPase fold characteristic of the shikimate kinase family, with a glycine-rich nucleotide-binding loop near the N-terminus. As a cytoplasmic monomer, AroK provides shikimate 3-phosphate as the substrate for the downstream EPSP synthase (AroA) and ultimately chorismate, the branch-point precursor for aromatic amino acids (Phe, Tyr, Trp), folate, ubiquinone, and other aromatic metabolites. In bacteria, shikimate kinase activity can be encoded by two isoenzymes, AroK (shikimate kinase I) and AroL (shikimate kinase II); this step is a recognized flux-control point in shikimate-pathway metabolic engineering.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000287 magnesium ion binding
IEA
GO_REF:0000104
ACCEPT
Summary: Shikimate kinase requires a divalent Mg(2+) cofactor for catalysis; UniProt records binding of one Mg(2+) ion per subunit, coordinated near the ATP-binding P-loop. This is a well-supported, family-conserved molecular function.
Reason: Magnesium binding is intrinsic to shikimate kinase catalysis and is supported by the HAMAP rule and conserved metal-binding residue; consistent with the enzymology of the shikimate kinase family.
GO:0004765 shikimate kinase activity
IEA
GO_REF:0000120
ACCEPT
Summary: This is the core molecular function of AroK, supported by the EC 2.7.1.71 assignment, the RHEA:13121 catalytic-activity mapping, HAMAP rule MF_00109, and strong family/domain evidence (Pfam SKI, shikimate kinase signature).
Reason: Directly represents the defining enzymatic activity of the gene product and is the central core function.
GO:0005524 ATP binding
IEA
GO_REF:0000104
ACCEPT
Summary: AroK is an ATP-dependent kinase that uses ATP as the phosphoryl donor; it contains a P-loop/Walker A motif (residues 11-16) and additional ATP-contacting residues. ATP binding is a required, well-supported molecular function.
Reason: ATP binding is an obligatory part of the shikimate kinase reaction and is supported by the conserved P-loop nucleotide-binding fold.
GO:0005737 cytoplasm
IEA
GO_REF:0000120
ACCEPT
Summary: AroK is a soluble cytoplasmic enzyme acting in central metabolism, consistent with UniProt subcellular location and the general biology of bacterial shikimate-pathway enzymes.
Reason: Correct localization for a cytoplasmic biosynthetic enzyme; consistent with the cytosol annotation below.
GO:0005829 cytosol
IEA
GO_REF:0000118
ACCEPT
Summary: TreeGrafter assigns the cytosol component, the more specific child of cytoplasm, consistent with a soluble bacterial shikimate kinase.
Reason: Accurate and slightly more specific localization than GO:0005737; both are appropriate for this enzyme.
GO:0009423 chorismate biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: AroK catalyzes step 5 of 7 in chorismate biosynthesis (the shikimate pathway), converting shikimate to shikimate 3-phosphate en route to chorismate. This is the precise biological process for the gene.
Reason: Directly and specifically captures the pathway role of AroK; UniPathway UPA00053/UER00088 corroborates the chorismate-biosynthesis placement.

Core Functions

ATP- and Mg(2+)-dependent shikimate kinase that phosphorylates the 3-hydroxyl of shikimate to form shikimate 3-phosphate, catalyzing step 5 of the shikimate pathway in chorismate biosynthesis.

Molecular Function:
shikimate kinase activity
Directly Involved In:
Supporting Evidence:
  • GO_REF:0000120
    EC 2.7.1.71; RHEA:13121 shikimate + ATP = 3-phosphoshikimate + ADP + H(+); HAMAP-Rule MF_00109.

References

Electronic Gene Ontology annotations created by transferring manual GO annotations between related proteins based on shared sequence features
TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.

Suggested Questions for Experts

Q: Does P. putida KT2440 encode a second shikimate kinase isoenzyme (AroL/shikimate kinase II), and what is the relative contribution of AroK versus any paralog to in vivo shikimate kinase flux?

Suggested Experiments

Experiment: Determine steady-state kinetic parameters (Km/kcat for shikimate and ATP, Mg(2+) dependence) of recombinant KT2440 AroK to confirm substrate specificity and rule out promiscuous quinate kinase activity.

Experiment: Assess essentiality/fitness of an aroK deletion in KT2440 on minimal medium with and without aromatic amino acid supplementation to test for aromatic auxotrophy.

Deep Research

Asta

(aroK-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:27:54.488312

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Synthesis, characterization, and computational evaluation of some synthesized xanthone derivatives: focus on kinase target network and biomedical properties

  • Authors: Wisam Taher Muslim, L. J. Mohammad, Munaf M. Naji, Isaac Karimi, Matheel D. Al-Sabti et al.
  • Year: 2025
  • Venue: Frontiers in Pharmacology
  • URL: https://www.semanticscholar.org/paper/659ab502877a1d6b5ab7ce45fa51f0f9a13dcf24
  • DOI: 10.3389/fphar.2024.1511627
  • PMID: 39830340
  • PMCID: 11738930
  • Summary: Acute leukemic T-cells were one of the top predicted tumor cell lines for these ligands and the possible antileukemic effects of synthesized xanthone derivatives are potentially very interesting and warrant further studies.
  • Evidence snippets:
  • Snippet 1 (score: 0.719)
    > The UniProt accession identification of target kinases was converted to gene symbols for humans using the SynGO gene set analysis tool (Koopmans et al., 2019), and pooled together, and submitted to GeneMANIA to construct target kinase network. GeneMANIA is a handy web interface for acquiring gene ontology, scrutinizing gene lists, and highlighting genes for functional assays (Warde-Farley et al., 2010). After choosing Homo sapiens from the list of optional organisms, the genes of interest in the previous step were entered into the search bar and the results were collated and high-scored genes were culled for further discussion. Moreover, the protein-protein network was also constructed in STRING ver. 12 launched at https://string-db.org, and submitted to Cytoscape ver. 3.10.2 for network analysis using a novel Cytoscape plugin cytoHubba and visualization (Shannon et al. , 2003).

[2] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.717)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[3] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.716)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[4] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[5] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.704)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[6] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.700)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[7] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.694)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[8] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.693)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[9] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.672)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[10] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[11] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[12] MultiLoc2: integrating phylogeny and Gene Ontology terms improves subcellular protein localization prediction

  • Authors: Torsten Blum, S. Briesemeister, O. Kohlbacher
  • Year: 2009
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/c2f00f9a94fe72eeeadc54a37a816731f329bfa4
  • DOI: 10.1186/1471-2105-10-274
  • PMID: 19723330
  • PMCID: 2745392
  • Citations: 294
  • Influential citations: 36
  • Summary: MultiLoc2 is an extensive high-performance subcellular protein localization prediction system that outperforms other prediction systems in two benchmarks studies and yields higher accuracies compared to its previous version.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > The Gene Ontology (GO) is a controlled vocabulary for uniformly describing gene products in terms of biological processes, cellular components and molecular function across all organisms [46]. It has been shown that GO terms can be used to improve the performance of subcellular protein localization prediction methods [47,48]. In the literature to date, there are three possibilities for obtaining GO annotation terms for a query sequence. If the UniProt [49] accession number is known, one can simply extract the GO annotation from the UniProt database [50]. However, this procedure fails for novel proteins without accession number. Another possibility is to search for homologous proteins annotated with GO terms using BLAST [28,29]. This becomes difficult in cases where proteins have no close homolog or proteins have many homologs, because no GO term can be obtained or GO terms might be ambiguous. A further method of inferring GO terms is InterProScan [51] used, for example, by Chou and Cai [52]. Given a protein sequence, the tool scans against various pattern and signature data sources collected by the InterPro project [53]. InterPro also provides a mapping of the detected protein domains and functional sites to GO terms.
    > Our subpredictor GOLoc is based on GO terms calculated using InterProScan. Since the GO terms are derived directly from the query sequence, we avoid the drawbacks of using accession numbers or BLAST. The input of GOLoc is a binary-coded vector which represents all GO terms of the training sequences (see Fig. 2). GO terms present in the query sequence are set to 1 in the vector and to 0 otherwise [see Additional file 1].

[13] Telomere-to-telomere genome assembly of Phoxinus lagowskii

  • Authors: Yanfeng Zhou, Chunhai Chen, Di’an Fang, Chenhe Wang, Yajuan Peng et al.
  • Year: 2025
  • Venue: Scientific Data
  • URL: https://www.semanticscholar.org/paper/23f573aff23769df3b733d508b46f721cadb94ac
  • DOI: 10.1038/s41597-025-05367-0
  • PMID: 40533492
  • PMCID: 12177035
  • Citations: 1
  • Summary: A T2T (Telomere-to-telomere) genome for P. lagowskii with chromosome-level is reported, serving as an invaluable resource for studies in evolution, comparative genomics, fish breeding applications, and ecological research.
  • Evidence snippets:
  • Snippet 1 (score: 0.661)
    > The gene set of 24,610 genes were functionally annotated using diamond v0.8.23
    > with an E-value threshold of 1E-5 based on the five databases including NR (NCBI nonredundant protein), TrEMB (http://www.uniprot.org), KOG 42 , KEEG (Kyoto Encyclopedia of Genes and Genomes, http://www. genome.jp/kegg/), Swiss-Prot (http://www.gpmaw.com/html/swiss-prot.html). The protein motifs and domains were identified using the InterProScan with InterPro 93.0 43 . The GO Ontology (GO) was classified from the results of InterProScan 44 . The annotation of 24,599 predicted genes (99.96%) out of the total 24,610 genes can be found by at least one database (Fig. 3b and Table 6). Of these functional proteins, 19,367 genes (~78.7%) were supported by five databases (Fig. 3a).

[14] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.657)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[15] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.655)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[16] RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications

  • Authors: Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al.
  • Year: 2025
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  • DOI: 10.1093/nar/gkaf1206
  • PMID: 41206768
  • PMCID: 12807602
  • Citations: 2
  • Summary: The new RM2Target v2.0 will serve as a foundational resource for exploring RNA epitranscriptomic regulation, enabling investigations into cross-talk among modifications, underlying molecular mechanisms, and disease connections, thereby facilitating both basic research and translational applications in RNA epigenetics.
  • Evidence snippets:
  • Snippet 1 (score: 0.653)
    > To obtain basic information on WERs and their target genes, such as official gene symbols, gene IDs, gene types, and genomic locations, gene annotations were downloaded from the GENCODE project [ 44 ] for human and mouse, and from NCBI [ 45 ] and Ensembl [ 46 ] for the other species. Genomic locations were extracted from the corresponding GTF annotation files. Gene symbols were primarily standardized based on the NCBI Gene database [ 45 ] for mRNAs and lncRNAs, GtR-NAdb [ 47 ] for tRNAs, miRbase [ 48 ] for microRNAs, and cir-cBase [ 49 ] for circRNAs. Deprecated or substituted versions of genes were filtered out. The LiftOver [ 50 ] program was employed to convert and unify genomic coordinates across different genome assembly versions.
    > The functional descriptions of WERs were compiled based on the UniProt database [ 51 ] and further supplemented with evidence from relevant publications, with particular emphasis on their functions as RNA modification regulatory proteins.

[17] HomoKinase: A Curated Database of Human Protein Kinases

  • Authors: S. Subramani, Saranya Jayapalan, R. Kalpana, J. Natarajan
  • Year: 2013
  • Venue: International Scholarly Research Notices
  • URL: https://www.semanticscholar.org/paper/81a21403fc772c6b9c00915e1124fad2cfb1894b
  • DOI: 10.1155/2013/417634
  • Citations: 18
  • Summary: The present version of theomoKinase database contains 498 curated human protein kinases and links to other popular databases.
  • Evidence snippets:
  • Snippet 1 (score: 0.653)
    > The HGNC approved human genes, which satisfy all these three GO criteria, were classified as human protein kinases and used to build the database.
    > The predicted list of protein kinases were further divided into groups, families, subfamilies, and domains. The group classifications were done using the PhosphoSite database [13], whereas the superfamily, family, subfamily, and domain level classifications were retrieved from UniProt [14]. In addition, various biological information such as official symbol, full name, biological IDs, other known aliases, amino acid sequences, functional domain, gene ontology, pathway assignments, and drug compounds were extracted from various biological databases such as (i) NCBI, (ii) UniProt, (iii) Amigo Go, (iv) KEGG, and (v) DrugBank. Figure 2 depicts a schematic summary of the HomoKinase data warehouse creation process.
    > The curated human protein kinase names and their related information retrieved from other databases were used to develop the HomoKinase database. The HomoKinase database is implemented as client/server architecture with easy-to-use web interface. The server is made of MySQL database, and the web client and programs for the human protein kinase retrieval, annotation, and query interface were designed using PHP programming language. Entrez Gene stores information on 1,93,709 genes specific to Homo sapiens (as on October 2012). We retrieved 33,489 human genes/proteins specific to our query term "(Homo sapiens [Organism] AND HGNC). " On further comparison with HGNC database, only the 19,026 genes have official HGNC gene symbol, and the remaining were 8399 pseudogenes, 4230 noncoding RNAs, 707 phenotype, and 1127 other genes.
    > The 19,032 HGNC approved human genes were further classified into protein kinases by checking the presence of three GO annotation terms (i) ATP binding property, (ii) kinase activity, and (iii) protein phosphorylation property. The HGNC approved genes fulfilling the above three GO properties (e.g., CDK1, MARK1) were classified as protein kinases and included in the database.

[18] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.652)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[19] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.652)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[20] GOnet: a tool for interactive Gene Ontology analysis

  • Authors: M. Pomaznoy, Brendan Ha, Bjoern Peters
  • Year: 2018
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  • DOI: 10.1186/s12859-018-2533-3
  • PMID: 30526489
  • PMCID: 6286514
  • Citations: 247
  • Influential citations: 17
  • Summary: The open-source GOnet web-application is created, which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis and provides insight into the functional interconnection of the submitted entries.
  • Evidence snippets:
  • Snippet 1 (score: 0.648)
    > In a basic workflow, the GOnet application receives a list of gene symbols, protein symbols, or protein IDs (UniProt IDs) as an input, and outputs a graph (an example given in Fig. 1). There are various input parameters which will affect the actual structure of the graph visualized and its appearance. The first main user choice is which GO terms the genes are annotated against:
    > 1. GO terms statistically significantly over-represented in the gene list submitted. 2. A predefined subset (also known as 'GO slim'), or a user-supplied list of terms.
    > In the first case the analysis will be referred to as an 'enrichment' analysis, in the second as an 'annotation' analysis.
    > Input parameters 1) Gene list. A mandatory input parameter containing the genes/proteins of interest. Currently human and mouse data is supported. An example of a human gene list might look like this:
    > Fig. 1 Sample network output generated by GOnet application. Gene differentially expressed in CD4 Bulk Memory T cells in Latent TB patients compared to healthy controls were used as an example [22] The gene list can also be accompanied with a contrast value. For example, This contrast value can be any decimal number, such as the log-fold change of gene expression between two conditions. This is merely a visualization enhancement. If the value is supplied it can be used later to differentially color specific genes in the graph (note different colors of gene nodes in Fig. 1), and visually indicate up-or down-regulation of specific genes and gene clusters.
    > The application can process common gene symbols (like in the example above), UniProt IDs, and MGI Accession IDs (mouse only). The former type of ID (gene symbols), although is the most human friendly, can unfortunately be ambiguous. For example, AIM1 can mean 'absent in melanoma' (also called CRYBG1) or 'Aurora and Ipl1-like midbody-associated protein' (also known as AURKB). Due to this ambiguity UniProt IDs or MGI accession IDs (for mouse) are preferred.
    > 2) GO namespace. Can be any of 'biological process', 'molecular function' or 'cellular component'.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Wisam Taher Muslim, L. J. Mohammad, Munaf M. Naji, Isaac Karimi, Matheel D. Al-Sabti et al. (2025). Synthesis, characterization, and computational evaluation of some synthesized xanthone derivatives: focus on kinase target network and biomedical properties. Frontiers in Pharmacology. https://www.semanticscholar.org/paper/659ab502877a1d6b5ab7ce45fa51f0f9a13dcf24
  2. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  3. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  4. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  5. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  6. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  7. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  8. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  9. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  10. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  11. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  12. Torsten Blum, S. Briesemeister, O. Kohlbacher (2009). MultiLoc2: integrating phylogeny and Gene Ontology terms improves subcellular protein localization prediction. BMC Bioinformatics. https://www.semanticscholar.org/paper/c2f00f9a94fe72eeeadc54a37a816731f329bfa4
  13. Yanfeng Zhou, Chunhai Chen, Di’an Fang, Chenhe Wang, Yajuan Peng et al. (2025). Telomere-to-telomere genome assembly of Phoxinus lagowskii. Scientific Data. https://www.semanticscholar.org/paper/23f573aff23769df3b733d508b46f721cadb94ac
  14. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  15. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  16. Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al. (2025). RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications. Nucleic Acids Research. https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  17. S. Subramani, Saranya Jayapalan, R. Kalpana, J. Natarajan (2013). HomoKinase: A Curated Database of Human Protein Kinases. International Scholarly Research Notices. https://www.semanticscholar.org/paper/81a21403fc772c6b9c00915e1124fad2cfb1894b
  18. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  19. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  20. M. Pomaznoy, Brendan Ha, Bjoern Peters (2018). GOnet: a tool for interactive Gene Ontology analysis. BMC Bioinformatics. https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9

Falcon

(aroK-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 23 citations 2 artifacts 2026-06-11T21:09:29.533954

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: Functional annotation of aroK (UniProt Q88CV1, locus PP_5079) in Pseudomonas putida KT2440

0) Gene/protein identity verification (critical)

The target provided (UniProt Q88CV1; gene aroK; ordered locus PP_5079) is annotated as shikimate kinase (SK), EC 2.7.1.71, belonging to the shikimate kinase family with P-loop NTPase/“shikimate/glucono_kinase” domain architecture (per user-supplied UniProt context). In the retrieved literature, aroK consistently refers to a bacterial shikimate kinase I isoenzyme (contrasted with aroL, shikimate kinase II), consistent with the UniProt description and EC assignment. (gu2016tunableswitchmediated pages 1-2, niraula2025aromaticaminoacids pages 8-10)

Limitation: none of the retrieved papers explicitly cross-references UniProt accession Q88CV1 or the KT2440 locus tag PP_5079 in the text excerpts obtained, so KT2440-specific details below are supported by P. putida KT2440 studies that manipulate aroK at the gene level, combined with general shikimate-kinase enzymology. (ling2022muconicacidproduction pages 2-3, gu2016tunableswitchmediated pages 1-2)


1) Key concepts and definitions (current understanding)

1.1 Shikimate pathway and the biochemical role of AroK

The shikimate pathway is the central biosynthetic route from central carbon metabolism toward chorismate, a key branch point precursor for aromatic amino acids and numerous aromatic metabolites. Within this pathway, shikimate kinase (SK; EC 2.7.1.71) catalyzes the ATP-dependent phosphorylation of shikimate to shikimate-3-phosphate (S3P); this step is required to proceed to 5-enolpyruvyl-shikimate-3-phosphate (EPSP) and ultimately chorismate. (niraula2025aromaticaminoacids pages 8-10)

Reaction (core functional annotation):
- Shikimate + ATP → Shikimate-3-phosphate + ADP (SK; EC 2.7.1.71). (niraula2025aromaticaminoacids pages 8-10)

1.2 AroK versus AroL (bacterial isoenzymes)

Many bacteria encode two shikimate kinase isoenzymes:
- AroK: shikimate kinase I
- AroL: shikimate kinase II

This redundancy is important in both physiology and metabolic engineering: engineered “shikimate-overproducer” strains often inactivate/decrease shikimate kinase activity to accumulate shikimate upstream, but complete loss of SK activity can lead to aromatic auxotrophy (requiring supplementation with aromatic compounds). (gu2016tunableswitchmediated pages 1-2)

1.3 Structure/function concepts relevant to substrate specificity

While organism-specific kinetics for P. putida KT2440 AroK were not retrieved in the accessible excerpts, a recent synthesis of bacterial SK structure–function notes that shikimate-binding can be described via subsites (CX, OCORE, OLID) and that specific conserved residues (e.g., positions corresponding to R60 and R140 in one referenced mapping) are functionally critical; mutating strictly conserved residues severely compromises function in engineering contexts, consistent with strong substrate-enzyme specificity requirements for shikimate recognition. (bo2023multiplemetabolicengineering pages 7-9)


2) Functional annotation for P. putida KT2440 aroK (Q88CV1)

2.1 Molecular function (enzyme activity)

AroK is a shikimate kinase (SK; EC 2.7.1.71) catalyzing the ATP-dependent phosphorylation of shikimate to shikimate-3-phosphate in the shikimate pathway. (niraula2025aromaticaminoacids pages 8-10, gu2016tunableswitchmediated pages 1-2)

2.2 Biological process / pathway placement

AroK functions in chorismate biosynthesis via the shikimate pathway, enabling downstream synthesis of aromatic amino acids and diverse chorismate-derived metabolites. (niraula2025aromaticaminoacids pages 8-10)

2.3 Cellular localization (where it acts)

The retrieved sources explicitly contrast chloroplast-localized shikimate kinase isoforms in algae/plants with bacterial SK functioning without organellar targeting; thus, for bacteria including P. putida KT2440, the working assumption supported by this context is that SK is a soluble, intracellular (cytosolic) enzyme acting in central metabolism. (niraula2025aromaticaminoacids pages 8-10)

Limitation: no KT2440-specific localization experiment (e.g., fractionation, fluorescence tagging) was found in the retrieved set.


3) Recent developments and latest research (prioritizing 2023–2024)

3.1 KT2440 as an aromatics/synbio chassis (expert synthesis, 2024)

A 2024 authoritative review/perspective describes P. putida KT2440’s rise as a synthetic biology chassis and emphasizes that its central metabolism delivers high NADPH regeneration (via ED/EDEMP-related architecture), supporting expression of heterologous pathways and tolerance to stress—properties relevant to building shikimate-pathway-derived aromatic production strains. (lorenzo2024pseudomonasputidakt2440 pages 2-4, lorenzo2024pseudomonasputidakt2440 pages 4-7, lorenzo2024pseudomonasputidakt2440 pages 1-2)

This review also highlights KT2440’s native and engineered capabilities in aromatic metabolism and production, citing examples of engineered production of shikimate-derived aromatics such as phenol, vanillate, and anthranilate. (lorenzo2024pseudomonasputidakt2440 pages 4-7)

Source: de Lorenzo et al., 2024-07, Journal of Bacteriology. https://doi.org/10.1128/jb.00136-24 (lorenzo2024pseudomonasputidakt2440 pages 2-4)

3.2 2024 “shikimate pathway-dependent catabolism” (SDC) rewiring to near-theoretical yields

A 2024 study reports a major conceptual advance: reprogramming P. putida so that growth-supporting catabolism becomes dependent on shikimate-pathway-derived reactions, demonstrating substantial metabolic plasticity and enabling very high-yield aromatic production. They report 89% of the maximum theoretical yield for 4-hydroxybenzoate in minimal medium and describe adaptive evolution mutations (e.g., in miaA and mexT) as important contributors. (santos2024shikimatepathwaydependentcatabolism pages 5-7)

Source: dos Santos et al., 2024-08 (preprint server posting). https://doi.org/10.21203/rs.3.rs-4761679/v1 (santos2024shikimatepathwaydependentcatabolism pages 5-7)

3.3 2024 combinatorial pathway engineering to find bottlenecks (pABA)

A 2024 design-of-experiments/combinatorial expression study in P. putida quantitatively mapped how varying shikimate- and pABA-pathway gene expression influences production. It reports wide performance spread (2–186.2 mg/L) and improved strains (~232.1 mg/L), identifying aroB (not aroK) as a key bottleneck for pABA production under tested conditions, while situating the shikimate pathway as a major industrial node (muconate, tryptophan, salicylic acid, vanillin, 2-phenylethanol). (camposmagana2024combinatorialengineeringreveals pages 1-4)

Source: Campos-Magaña et al., 2024-06 (bioRxiv). https://doi.org/10.1101/2024.06.17.599342 (camposmagana2024combinatorialengineeringreveals pages 1-4)


4) Current applications and real-world implementations (with P. putida examples)

4.1 AroK as a practical flux-control lever in KT2440 muconate bioproduction

A key KT2440 implementation is cis,cis-muconate production, a bioprivileged platform chemical. In a high-impact Nature Communications study, KT2440-derived strains were engineered with aroK under the inducible Ptac promoter, either alone (ΔpykF::Ptac:aroK) or together with aroB (ΔpykF::Ptac:aroK:aroB), demonstrating that modulating aroK expression is used in practice to tune shikimate-pathway flux in KT2440-derived production strains. (ling2022muconicacidproduction pages 2-3, ling2022muconicacidproduction media b6590159)

The same study reports 33.7 g/L muconate, 0.18 g/L/h productivity, and 46% molar yield (reported as 92% of maximum theoretical yield) in the rationally engineered strain context. (ling2022muconicacidproduction pages 2-3)

Source: Ling et al., 2022-08, Nature Communications. https://doi.org/10.1038/s41467-022-32296-y (ling2022muconicacidproduction pages 2-3)

4.2 KT2440 production of gallic acid from glycerol (2023)

A 2023 study demonstrates conversion of glycerol to the antioxidant gallic acid in KT2440 by combining heterologous pathway expression with deletions that block product degradation, explicitly leveraging KT2440’s metabolic and chassis advantages. The reported titer is 346.7 ± 0.004 mg/L gallic acid after 72 h (shake flasks). (dias2023fromdegraderto pages 1-2)

Source: Dias et al., 2023-11, International Microbiology. https://doi.org/10.1007/s10123-022-00282-5 (dias2023fromdegraderto pages 1-2)


5) Expert opinions and analysis (authoritative sources)

5.1 Why KT2440 is repeatedly chosen for shikimate-pathway aromatics

Across authoritative sources, KT2440 is characterized as an “industriphilic” synthetic biology chassis and “versatile aromatics cell factory,” with repeated emphasis on:
- High NADPH regeneration and redox-biased central metabolism, supporting reductive biosynthesis and heterologous pathway operation. (lorenzo2024pseudomonasputidakt2440 pages 2-4, lorenzo2024pseudomonasputidakt2440 pages 4-7)
- Intrinsic tolerance to aromatic compounds and possession of oxygenase repertoires to transform them—useful because many aromatics are toxic or require harsh processes. (lorenzo2024pseudomonasputidakt2440 pages 4-7)

These points contextualize aroK’s importance: as a control point in the shikimate pathway, aroK expression/activity influences whether carbon proceeds toward chorismate-derived products (desired aromatics) or accumulates upstream intermediates (e.g., shikimate) (gu2016tunableswitchmediated pages 1-2, ling2022muconicacidproduction pages 2-3).


6) Quantitative statistics and data points from recent and relevant studies

Key quantitative outcomes tied to shikimate-pathway engineering in P. putida include:
- Muconate production (KT2440): 33.7 g/L, 0.18 g/L/h, 46% molar yield (92% of max theoretical); includes strains engineered with Ptac:aroK (flux control through shikimate pathway). (ling2022muconicacidproduction pages 2-3, ling2022muconicacidproduction media b6590159)
- Gallic acid production (KT2440): 346.7 ± 0.004 mg/L after 72 h from 10 g/L glycerol in shake flasks (as described in excerpt). (dias2023fromdegraderto pages 1-2)
- pABA production (P. putida, 2024 DoE): 2–186.2 mg/L across sampled genotypes; improved to ~232.1 mg/L; identified aroB as bottleneck. (camposmagana2024combinatorialengineeringreveals pages 1-4)
- 4-hydroxybenzoate yield (P. putida, 2024 SDC concept): 89% of maximum theoretical yield in minimal medium (yield benchmark claim). (santos2024shikimatepathwaydependentcatabolism pages 5-7)


7) Evidence summary table

Claim/Topic Evidence summary Organism/context Quantitative data Source (with year, DOI URL)
Core enzymatic reaction of shikimate kinase Shikimate kinase (SK; EC 2.7.1.71) phosphorylates shikimate at the C3 hydroxyl using ATP to form shikimate-3-phosphate (S3P), a required step toward EPSP and chorismate in the shikimate pathway (niraula2025aromaticaminoacids pages 8-10) General bacterial/plant/algal shikimate-pathway context Reaction: shikimate + ATP -> shikimate-3-phosphate + ADP Niraula et al., 2025, https://doi.org/10.3390/biotech14010006
Bacterial isoenzymes AroK and AroL In bacteria, two shikimate kinase isoenzymes are commonly recognized: shikimate kinase I (AroK) and shikimate kinase II (AroL). Engineering studies often repress/delete these enzymes to accumulate shikimate, but full loss causes auxotrophy and necessitates aromatic supplementation (gu2016tunableswitchmediated pages 1-2) Engineered bacterial shikimate producers, especially E. coli Example reported for engineered E. coli: 13.15 g/L shikimate in 5-L fed-batch using tunable aroK expression rather than permanent deletion Gu et al., 2016, https://doi.org/10.1038/srep29745
Shikimate-binding determinants in AroK Structural/functional analysis of bacterial AroK identifies shikimate-binding subsites CX, OCORE, and OLID. Conserved residues contacting the substrate include positions corresponding to R60 and R140; mutating these severely impaired function/growth in engineering contexts (bo2023multiplemetabolicengineering pages 7-9) Bacterial AroK, with cited structural work including Helicobacter pylori and other bacterial SKs No single kinetic value in snippet; severe growth inhibition observed for key conserved-site mutants Bo et al., 2023, https://doi.org/10.3390/metabo13060747
KT2440-specific aroK manipulation in muconate engineering In Pseudomonas putida KT2440-derived strains for muconate production, aroK was deliberately overexpressed from Ptac, either alone or together with aroB (e.g., ΔpykF::Ptac:aroK and ΔpykF::Ptac:aroK:aroB), indicating aroK is a practical flux-control point in the native shikimate pathway (ling2022muconicacidproduction pages 2-3) Pseudomonas putida KT2440 metabolic engineering for muconic acid Muconate process in study reached 33.7 g/L muconate, 0.18 g/L/h, 46% molar yield (92% of maximum theoretical yield); snippet specifically highlights Ptac:aroK strain designs Ling et al., 2022, https://doi.org/10.1038/s41467-022-32296-y
2024 pABA combinatorial engineering identifies pathway bottleneck A 2024 DoE study in P. putida varied shikimate- and pABA-pathway gene expression across 14 representative strains sampled from 512 possible combinations. The analysis identified aroB, not aroK, as a significant bottleneck for pABA production, while noting aroK overexpression had been beneficial in related production contexts (camposmagana2024combinatorialengineeringreveals pages 1-4) Pseudomonas putida shikimate/pABA pathway engineering Titers ranged from 2 to 186.2 mg/L initially; second-round designs reached ~232.1 mg/L Campos-Magaña et al., 2024, https://doi.org/10.1101/2024.06.17.599342
2024 shikimate pathway-dependent catabolism (SDC) A 2024 P. putida study rewired metabolism so shikimate-pathway-derived reactions supplied pyruvate, demonstrating strong metabolic plasticity of the pathway relevant to enzymes such as AroK that control flux into downstream aromatic intermediates (santos2024shikimatepathwaydependentcatabolism pages 5-7) Pseudomonas putida engineered for aromatic production 89% of maximum theoretical yield for 4-hydroxybenzoate in minimal medium; adapted strains also produced salicylate and 3-hydroxybenzoate at high yields dos Santos et al., 2024, https://doi.org/10.21203/rs.3.rs-4761679/v1
2023 real-world aromatic production application in KT2440 P. putida KT2440 was engineered to produce gallic acid from glycerol using a synthetic operon (aroG4, quiC, pobA*) plus deletions blocking degradation (pcaHG and galTAPR). This shows the practical biotechnological value of controlling shikimate-pathway flux in KT2440, even though aroK itself was not directly manipulated in the snippet (dias2023fromdegraderto pages 1-2) Pseudomonas putida KT2440 producing gallic acid from glycerol 346.7 ± 0.004 mg/L gallic acid after 72 h in shake flasks Dias et al., 2023, https://doi.org/10.1007/s10123-022-00282-5

Table: This table summarizes the most relevant gathered evidence for annotating Pseudomonas putida KT2440 aroK as shikimate kinase, including core reaction chemistry, bacterial isoenzyme context, and KT2440-specific engineering evidence. It is useful for separating direct functional evidence from broader pathway and application-level evidence.


aroK (UniProt Q88CV1; PP_5079) encodes shikimate kinase (EC 2.7.1.71), a cytosolic ATP-dependent kinase that converts shikimate to shikimate-3-phosphate in the shikimate pathway, enabling downstream EPSP/chorismate formation and thereby aromatic amino acid and aromatic metabolite biosynthesis. (niraula2025aromaticaminoacids pages 8-10, gu2016tunableswitchmediated pages 1-2)

In P. putida KT2440, aroK is repeatedly treated as a flux-control point in industrially motivated shikimate-pathway engineering, including muconate production strains where aroK is overexpressed from Ptac in defined genomic backgrounds. (ling2022muconicacidproduction pages 2-3, ling2022muconicacidproduction media b6590159)


The current evidence set is strong for pathway placement and engineering relevance, but limited for P. putida KT2440-specific mechanistic details.

Most impactful missing items would be:
1) Direct biochemical characterization (Km/kcat for shikimate/ATP; Mg2+ dependence; substrate scope such as quinate) for KT2440 AroK.
2) Genetic essentiality/fitness measurements for KT2440 aroK under defined media (e.g., transposon insertion sequencing) and rescue by aromatic supplementation.
3) Subcellular localization experiments (fractionation or fluorescence tagging) in KT2440.

These gaps were not recoverable within the retrieved corpus and are therefore not asserted beyond supported inference.

References

  1. (gu2016tunableswitchmediated pages 1-2): Pengfei Gu, Tianyuan Su, Qian Wang, Quanfeng Liang, and Qingsheng Qi. Tunable switch mediated shikimate biosynthesis in an engineered non-auxotrophic escherichia coli. Scientific Reports, Jul 2016. URL: https://doi.org/10.1038/srep29745, doi:10.1038/srep29745. This article has 43 citations and is from a peer-reviewed journal.

  2. (niraula2025aromaticaminoacids pages 8-10): Archana Niraula, Amir Danesh, Natacha Merindol, Fatma Meddeb-Mouelhi, and Isabel Desgagné-Penix. Aromatic amino acids: exploring microalgae as a potential biofactory. BioTech, 14:6, Jan 2025. URL: https://doi.org/10.3390/biotech14010006, doi:10.3390/biotech14010006. This article has 9 citations.

  3. (ling2022muconicacidproduction pages 2-3): Chen Ling, George L. Peabody, Davinia Salvachúa, Young-Mo Kim, Colin M. Kneucker, Christopher H. Calvey, Michela A. Monninger, Nathalie Munoz Munoz, Brenton C. Poirier, Kelsey J. Ramirez, Peter C. St. John, Sean P. Woodworth, Jon K. Magnuson, Kristin E. Burnum-Johnson, Adam M. Guss, Christopher W. Johnson, and Gregg T. Beckham. Muconic acid production from glucose and xylose in pseudomonas putida via evolution and metabolic engineering. Nature Communications, Aug 2022. URL: https://doi.org/10.1038/s41467-022-32296-y, doi:10.1038/s41467-022-32296-y. This article has 141 citations and is from a highest quality peer-reviewed journal.

  4. (bo2023multiplemetabolicengineering pages 7-9): Taidong Bo, Chen Wu, Zeting Wang, Hao Jiang, Fei-Yong Wang, N. Chen, and Yanjun Li. Multiple metabolic engineering strategies to improve shikimate titer in escherichia coli. Metabolites, 13:747, Jun 2023. URL: https://doi.org/10.3390/metabo13060747, doi:10.3390/metabo13060747. This article has 12 citations.

  5. (lorenzo2024pseudomonasputidakt2440 pages 2-4): Victor de Lorenzo, Danilo Pérez-Pantoja, and Pablo I. Nikel. pseudomonas putida kt2440: the long journey of a soil-dweller to become a synthetic biology chassis. Journal of Bacteriology, Jul 2024. URL: https://doi.org/10.1128/jb.00136-24, doi:10.1128/jb.00136-24. This article has 78 citations and is from a peer-reviewed journal.

  6. (lorenzo2024pseudomonasputidakt2440 pages 4-7): Victor de Lorenzo, Danilo Pérez-Pantoja, and Pablo I. Nikel. pseudomonas putida kt2440: the long journey of a soil-dweller to become a synthetic biology chassis. Journal of Bacteriology, Jul 2024. URL: https://doi.org/10.1128/jb.00136-24, doi:10.1128/jb.00136-24. This article has 78 citations and is from a peer-reviewed journal.

  7. (lorenzo2024pseudomonasputidakt2440 pages 1-2): Victor de Lorenzo, Danilo Pérez-Pantoja, and Pablo I. Nikel. pseudomonas putida kt2440: the long journey of a soil-dweller to become a synthetic biology chassis. Journal of Bacteriology, Jul 2024. URL: https://doi.org/10.1128/jb.00136-24, doi:10.1128/jb.00136-24. This article has 78 citations and is from a peer-reviewed journal.

  8. (santos2024shikimatepathwaydependentcatabolism pages 5-7): Vitor Martins dos Santos, Lyon Bruinsma, Christos Batianis, Sara Moreno-Paz, Kesi Kurnia, Job Dirkmaat, Alexandra Müller, Jose Juncosa Nuñez, and Ruud Weusthuis. Shikimate pathway-dependent catabolism: enabling near-to-maximum production yield of aromatics. Unknown journal, Aug 2024. URL: https://doi.org/10.21203/rs.3.rs-4761679/v1, doi:10.21203/rs.3.rs-4761679/v1.

  9. (camposmagana2024combinatorialengineeringreveals pages 1-4): Marco A Campos-Magaña, Sara Moreno-Paz, Vitor AP Martins dos Santos, Luis Garcia-Morales, and Maria Suarez-Diez. Combinatorial engineering reveals shikimate pathway bottlenecks in para-aminobenzoic acid production in pseudomonas putida. bioRxiv, Jun 2024. URL: https://doi.org/10.1101/2024.06.17.599342, doi:10.1101/2024.06.17.599342. This article has 0 citations.

  10. (ling2022muconicacidproduction media b6590159): Chen Ling, George L. Peabody, Davinia Salvachúa, Young-Mo Kim, Colin M. Kneucker, Christopher H. Calvey, Michela A. Monninger, Nathalie Munoz Munoz, Brenton C. Poirier, Kelsey J. Ramirez, Peter C. St. John, Sean P. Woodworth, Jon K. Magnuson, Kristin E. Burnum-Johnson, Adam M. Guss, Christopher W. Johnson, and Gregg T. Beckham. Muconic acid production from glucose and xylose in pseudomonas putida via evolution and metabolic engineering. Nature Communications, Aug 2022. URL: https://doi.org/10.1038/s41467-022-32296-y, doi:10.1038/s41467-022-32296-y. This article has 141 citations and is from a highest quality peer-reviewed journal.

  11. (dias2023fromdegraderto pages 1-2): Felipe M. S. Dias, Raoní K. Pantoja, José Gregório C. Gomez, and Luiziana F. Silva. From degrader to producer: reversing the gallic acid metabolism of pseudomonas putida kt2440. International Microbiology, 26:243-255, Nov 2023. URL: https://doi.org/10.1007/s10123-022-00282-5, doi:10.1007/s10123-022-00282-5. This article has 7 citations and is from a peer-reviewed journal.

Artifacts

Citations

  1. niraula2025aromaticaminoacids pages 8-10
  2. gu2016tunableswitchmediated pages 1-2
  3. bo2023multiplemetabolicengineering pages 7-9
  4. santos2024shikimatepathwaydependentcatabolism pages 5-7
  5. camposmagana2024combinatorialengineeringreveals pages 1-4
  6. ling2022muconicacidproduction pages 2-3
  7. dias2023fromdegraderto pages 1-2
  8. https://doi.org/10.1128/jb.00136-24
  9. https://doi.org/10.21203/rs.3.rs-4761679/v1
  10. https://doi.org/10.1101/2024.06.17.599342
  11. https://doi.org/10.1038/s41467-022-32296-y
  12. https://doi.org/10.1007/s10123-022-00282-5
  13. https://doi.org/10.3390/biotech14010006
  14. https://doi.org/10.1038/srep29745
  15. https://doi.org/10.3390/metabo13060747
  16. https://doi.org/10.1038/srep29745,
  17. https://doi.org/10.3390/biotech14010006,
  18. https://doi.org/10.1038/s41467-022-32296-y,
  19. https://doi.org/10.3390/metabo13060747,
  20. https://doi.org/10.1128/jb.00136-24,
  21. https://doi.org/10.21203/rs.3.rs-4761679/v1,
  22. https://doi.org/10.1101/2024.06.17.599342,
  23. https://doi.org/10.1007/s10123-022-00282-5,

📄 View Raw YAML

id: Q88CV1
gene_symbol: aroK
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: AroK is the shikimate kinase (EC 2.7.1.71) of Pseudomonas putida KT2440. It catalyzes the fifth step of the shikimate pathway, the ATP-dependent phosphorylation of the 3-hydroxyl group of shikimate to form shikimate 3-phosphate (3-phosphoshikimate), with release of ADP and a proton. The enzyme requires a Mg(2+) cofactor (one ion bound per subunit) and adopts the P-loop (Walker A) NTPase fold characteristic of the shikimate kinase family, with a glycine-rich nucleotide-binding loop near the N-terminus. As a cytoplasmic monomer, AroK provides shikimate 3-phosphate as the substrate for the downstream EPSP synthase (AroA) and ultimately chorismate, the branch-point precursor for aromatic amino acids (Phe, Tyr, Trp), folate, ubiquinone, and other aromatic metabolites. In bacteria, shikimate kinase activity can be encoded by two isoenzymes, AroK (shikimate kinase I) and AroL (shikimate kinase II); this step is a recognized flux-control point in shikimate-pathway metabolic engineering.
existing_annotations:
- term:
    id: GO:0000287
    label: magnesium ion binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000104
  qualifier: enables
  review:
    summary: Shikimate kinase requires a divalent Mg(2+) cofactor for catalysis; UniProt records binding of one Mg(2+) ion per subunit, coordinated near the ATP-binding P-loop. This is a well-supported, family-conserved molecular function.
    action: ACCEPT
    reason: Magnesium binding is intrinsic to shikimate kinase catalysis and is supported by the HAMAP rule and conserved metal-binding residue; consistent with the enzymology of the shikimate kinase family.
- term:
    id: GO:0004765
    label: shikimate kinase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: This is the core molecular function of AroK, supported by the EC 2.7.1.71 assignment, the RHEA:13121 catalytic-activity mapping, HAMAP rule MF_00109, and strong family/domain evidence (Pfam SKI, shikimate kinase signature).
    action: ACCEPT
    reason: Directly represents the defining enzymatic activity of the gene product and is the central core function.
- term:
    id: GO:0005524
    label: ATP binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000104
  qualifier: enables
  review:
    summary: AroK is an ATP-dependent kinase that uses ATP as the phosphoryl donor; it contains a P-loop/Walker A motif (residues 11-16) and additional ATP-contacting residues. ATP binding is a required, well-supported molecular function.
    action: ACCEPT
    reason: ATP binding is an obligatory part of the shikimate kinase reaction and is supported by the conserved P-loop nucleotide-binding fold.
- term:
    id: GO:0005737
    label: cytoplasm
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: located_in
  review:
    summary: AroK is a soluble cytoplasmic enzyme acting in central metabolism, consistent with UniProt subcellular location and the general biology of bacterial shikimate-pathway enzymes.
    action: ACCEPT
    reason: Correct localization for a cytoplasmic biosynthetic enzyme; consistent with the cytosol annotation below.
- term:
    id: GO:0005829
    label: cytosol
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: TreeGrafter assigns the cytosol component, the more specific child of cytoplasm, consistent with a soluble bacterial shikimate kinase.
    action: ACCEPT
    reason: Accurate and slightly more specific localization than GO:0005737; both are appropriate for this enzyme.
- term:
    id: GO:0009423
    label: chorismate biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: AroK catalyzes step 5 of 7 in chorismate biosynthesis (the shikimate pathway), converting shikimate to shikimate 3-phosphate en route to chorismate. This is the precise biological process for the gene.
    action: ACCEPT
    reason: Directly and specifically captures the pathway role of AroK; UniPathway UPA00053/UER00088 corroborates the chorismate-biosynthesis placement.
core_functions:
- description: ATP- and Mg(2+)-dependent shikimate kinase that phosphorylates the 3-hydroxyl of shikimate to form shikimate 3-phosphate, catalyzing step 5 of the shikimate pathway in chorismate biosynthesis.
  molecular_function:
    id: GO:0004765
    label: shikimate kinase activity
  supported_by:
  - reference_id: GO_REF:0000120
    supporting_text: EC 2.7.1.71; RHEA:13121 shikimate + ATP = 3-phosphoshikimate + ADP + H(+); HAMAP-Rule MF_00109.
  directly_involved_in:
  - id: GO:0009423
    label: chorismate biosynthetic process
proposed_new_terms: []
suggested_questions:
- question: Does P. putida KT2440 encode a second shikimate kinase isoenzyme (AroL/shikimate kinase II), and what is the relative contribution of AroK versus any paralog to in vivo shikimate kinase flux?
suggested_experiments:
- description: Determine steady-state kinetic parameters (Km/kcat for shikimate and ATP, Mg(2+) dependence) of recombinant KT2440 AroK to confirm substrate specificity and rule out promiscuous quinate kinase activity.
- description: Assess essentiality/fitness of an aroK deletion in KT2440 on minimal medium with and without aromatic amino acid supplementation to test for aromatic auxotrophy.
references:
- id: GO_REF:0000104
  title: Electronic Gene Ontology annotations created by transferring manual GO annotations between related proteins based on shared sequence features
  findings: []
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:12534463
  title: Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.
  findings: []
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: KT2440 genome paper (Nelson et al. 2002); same PMID verified in aroA/aroB/aroC/aroE.