glk

UniProt ID: Q88P42
Organism: Pseudomonas putida KT2440
Review Status: COMPLETE
Aliases:
PP_1011 Glucokinase Glucose kinase
📝 Provide Detailed Feedback

Gene Description

Glk is the cytosolic glucokinase of Pseudomonas putida KT2440. It phosphorylates imported glucose to glucose-6-phosphate using ATP and feeds the phosphorylative branch of glucose assimilation into the Entner-Doudoroff-centered carbohydrate catabolic network. In KT2440, glk is part of the edd-glk operon and mutant and flux analyses indicate that the glucokinase branch is quantitatively important for growth on glucose even though P. putida can also oxidize glucose through periplasmic gluconate and 2-ketogluconate routes.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0004340 glucokinase activity
IEA
GO_REF:0000120
ACCEPT
Summary: This is the core molecular function of glk. UniProt assigns the protein as glucokinase (EC 2.7.1.2), and pathway work in KT2440 places Glk at the ATP-dependent phosphorylation step that converts cytoplasmic glucose to glucose-6-phosphate.
Reason: This term is specific, mechanistically correct, and central to the gene's function.
Supporting Evidence:
file:PSEPK/glk/glk-uniprot.txt
RecName: Full=Glucokinase
file:PSEPK/glk/glk-uniprot.txt
Reaction=D-glucose + ATP = D-glucose 6-phosphate + ADP + H(+);
file:PSEPK/glk/glk-notes.md
In KT2440, glucose imported into the cytoplasm is phosphorylated by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
file:PSEPK/glk/glk-deep-research-falcon.md
The UniProt target **Q88P42** corresponds to **glk** in *Pseudomonas putida* KT2440 (ordered locus **PP_1011**) encoding the cytosolic enzyme **glucokinase (Glk)**, which phosphorylates imported glucose to **glucose‑6‑phosphate (G6P)** as the entry step of the phosphorylative branch of glucose catabolism.
GO:0005524 ATP binding
IEA
GO_REF:0000120
MARK AS OVER ANNOTATED
Summary: ATP binding is mechanistically true for a kinase, but it is much less informative than the specific catalytic term glucokinase activity and is redundant for describing the core function of this enzyme.
Reason: The catalytic activity term already captures the biologically informative function.
Supporting Evidence:
file:PSEPK/glk/glk-uniprot.txt
/ligand="ATP"
file:PSEPK/glk/glk-notes.md
ATP binding and D-glucose binding are mechanistically true but substantially less informative than the specific catalytic term glucokinase activity.
GO:0005536 D-glucose binding
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: Substrate recognition is implicit in glucokinase activity, so this term is not wrong but is less informative than the specific catalytic annotation.
Reason: This binding term adds little beyond the core enzymatic activity term.
Supporting Evidence:
file:PSEPK/glk/glk-notes.md
In KT2440, glucose imported into the cytoplasm is phosphorylated by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
file:PSEPK/glk/glk-notes.md
ATP binding and D-glucose binding are mechanistically true but substantially less informative than the specific catalytic term glucokinase activity.
GO:0005737 cytoplasm
IEA
GO_REF:0000120
MODIFY
Summary: Glk is a soluble intracellular enzyme and cytoplasmic localization is correct, but the more precise term in this context is cytosol.
Reason: Cytoplasm is broader than needed for a soluble bacterial enzyme.
Proposed replacements: cytosol
Supporting Evidence:
file:PSEPK/glk/glk-uniprot.txt
SUBCELLULAR LOCATION: Cytoplasm
file:PSEPK/glk/glk-notes.md
For cellular component, cytosol is the more specific GO term for a soluble bacterial cytoplasmic enzyme, while cytoplasm is acceptable but broader.
GO:0005829 cytosol
IEA
GO_REF:0000118
ACCEPT
Summary: This is the preferred cellular component term for a soluble cytoplasmic enzyme such as Glk and is more precise than the broader term cytoplasm.
Reason: Specific and biologically appropriate localization term.
Supporting Evidence:
file:PSEPK/glk/glk-uniprot.txt
SUBCELLULAR LOCATION: Cytoplasm
file:PSEPK/glk/glk-notes.md
For cellular component, cytosol is the more specific GO term for a soluble bacterial cytoplasmic enzyme, while cytoplasm is acceptable but broader.
GO:0006096 glycolytic process
IEA
GO_REF:0000120
MODIFY
Summary: Glk participates in glucose catabolism to pyruvate, but in P. putida KT2440 that process is routed through the Entner-Doudoroff-centered network rather than a generic undifferentiated glycolysis term.
Reason: A more pathway-specific biological process term is available.
Supporting Evidence:
file:PSEPK/glk/glk-notes.md
glk is physically linked to the Entner-Doudoroff pathway because edd and glk form one operon in KT2440.
file:PSEPK/glk/glk-notes.md
Because of this genomic organization, the glucokinase pathway is co-induced with Entner-Doudoroff genes and is expressed even during growth on gluconate or 2-ketogluconate.
file:PSEPK/glk/glk-deep-research-falcon.md
This physical linkage is consistent with functional coupling: the **glucokinase entry step (Glk)** and the subsequent ED processing step (Edd) are transcriptionally coordinated.
GO:0051156 glucose 6-phosphate metabolic process
IEA
GO_REF:0000002
ACCEPT
Summary: This term accurately reflects the immediate pathway context of Glk, which generates glucose-6-phosphate from glucose and ATP.
Reason: Directly describes the metabolic process in which the enzymatic reaction participates.
Supporting Evidence:
file:PSEPK/glk/glk-uniprot.txt
Reaction=D-glucose + ATP = D-glucose 6-phosphate + ADP + H(+);
file:PSEPK/glk/glk-notes.md
In KT2440, glucose imported into the cytoplasm is phosphorylated by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.

Core Functions

Glk is a soluble cytosolic glucokinase that uses ATP to phosphorylate imported glucose to glucose-6-phosphate, thereby feeding the phosphorylative arm of glucose assimilation into the Entner-Doudoroff-centered glycolytic network of Pseudomonas putida KT2440.

Supporting Evidence:
  • file:PSEPK/glk/glk-uniprot.txt
    Reaction=D-glucose + ATP = D-glucose 6-phosphate + ADP + H(+);
  • file:PSEPK/glk/glk-notes.md
    In KT2440, glucose imported into the cytoplasm is phosphorylated by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
  • file:PSEPK/glk/glk-notes.md
    glk is physically linked to the Entner-Doudoroff pathway because edd and glk form one operon in KT2440.
  • file:PSEPK/glk/glk-deep-research-falcon.md
    The UniProt target **Q88P42** corresponds to **glk** in *Pseudomonas putida* KT2440 (ordered locus **PP_1011**) encoding the cytosolic enzyme **glucokinase (Glk)**, which phosphorylates imported glucose to **glucose‑6‑phosphate (G6P)** as the entry step of the phosphorylative branch of glucose catabolism.

References

Gene Ontology annotation through association of InterPro records with GO terms.
  • InterPro-based annotation captures the conserved glucokinase family assignment and related substrate-level inferences.
TreeGrafter-generated GO annotations
  • TreeGrafter provides a phylogeny-based cellular component inference for this conserved bacterial glucokinase family.
Combined Automated Annotation using Multiple IEA Methods.
  • Automated UniProt and rule-based methods recover the core glucokinase activity and broad glucose catabolic process assignments for glk.
Convergent peripheral pathways catalyze initial glucose catabolism in Pseudomonas putida: genomic and flux analysis.
  • Glucose catabolism in P. putida occurs through three simultaneous pathways converging at 6-phosphogluconate.
    "glucose catabolism in Pseudomonas putida occurs through the simultaneous operation of three pathways that converge at the level of 6-phosphogluconate"
  • Glucose imported into the cytoplasm is phosphorylated by glucokinase to glucose-6-phosphate.
    "Glucose is transported to the cytoplasm in a process mediated by an ABC uptake system encoded by open reading frames PP1015 to PP1018 and is then phosphorylated by glucokinase (encoded by the glk gene) and converted by glucose-6-phosphate dehydrogenase (encoded by the zwf genes) to 6-phosphogluconate."
  • The glucokinase pathway and 2-ketogluconate loop are quantitatively more important than direct gluconate phosphorylation.
    "although all three functioned simultaneously, the glucokinase pathway and the 2-ketogluconate loop were quantitatively more important than the direct phosphorylation of gluconate."
  • The glucokinase pathway is required for growth on glucose in KT2440.
    "It can therefore be concluded that the glucokinase pathway is a sine qua non condition for P. putida to grow with glucose."
Regulation of Glucose Metabolism in Pseudomonas: The Phosphorylative Branch and Entner-Doudoroff Enzymes Are Regulated by a Repressor Containing a Sugar Isomerase Domain.
  • The glucose phosphorylative pathway and Entner-Doudoroff pathway are organized into operons that include an edd-glk-gltR2-gltS operon.
    "In Pseudomonas putida, genes for the glucose phosphorylative pathway and the Entner-Doudoroff pathway are organized in two operons; one made up of the zwf, pgl, and eda genes and another consisting of the edd, glk, gltR2, and gltS genes."
  • Expression from P(zwf), P(edd), and P(gap) is modulated by HexR in response to glucose availability.
    "Expression from P(zwf), P(edd), and P(gap) is modulated by HexR in response to the availability of glucose in the medium."
  • Binding of KDPG to HexR releases the repressor, whereas glucose, glucose 6-phosphate, and 6-phosphogluconate do not.
    "Binding of the Entner-Doudoroff pathway intermediate 2-keto-3-deoxy-6-phosphogluconate to HexR released the repressor from its target operators, whereas other chemicals such as glucose, glucose 6-phosphate, and 6-phosphogluconate did not induce complex dissociation."
file:PSEPK/glk/glk-uniprot.txt
UniProt entry Q88P42
  • glk corresponds to PP_1011 in Pseudomonas putida KT2440.
  • UniProt annotates the protein as glucokinase EC 2.7.1.2.
  • The catalytic reaction is D-glucose plus ATP to D-glucose 6-phosphate plus ADP and H+.
  • UniProt places the protein in the cytoplasm and in the bacterial glucokinase family.
file:PSEPK/glk/glk-notes.md
Curator notes for glk in Pseudomonas putida KT2440
  • The core role of glk is ATP-dependent phosphorylation of glucose to glucose-6-phosphate.
  • glk is physically and transcriptionally linked to the Entner-Doudoroff pathway.
  • Generic binding terms are less informative than glucokinase activity for this enzyme.
file:PSEPK/glk/glk-deep-research-falcon.md
Deep research report for glk generated with Falcon
  • Q88P42 corresponds to glk/PP_1011 and encodes the cytosolic glucokinase of Pseudomonas putida KT2440.
  • Falcon research links glk to the phosphorylative branch of glucose catabolism and the edd-glk operon.
  • Falcon research summarizes inducible glucokinase activity and mutant phenotypes supporting pathway relevance in KT2440.

Suggested Questions for Experts

Q: How does flux through the glucokinase branch versus the periplasmic gluconate and 2-ketogluconate branches change across different glucose concentrations and oxygen or redox conditions in KT2440?

Q: Does Glk have kinetic or regulatory specialization that distinguishes it from glucokinases in other pseudomonads that use different balances of oxidative and phosphorylative glucose uptake?

Q: How strongly does HexR-mediated control of the edd-glk operon constrain mixed-substrate utilization when glucose and gluconate are simultaneously available?

Suggested Experiments

Experiment: Purify PP_1011 and measure steady-state kinetics for glucose and ATP, substrate specificity, and cofactor dependence under physiologically relevant ionic conditions.

Type: Enzyme biochemistry

Experiment: Construct a clean glk deletion and complemented strain, then quantify growth and intracellular carbon flux on glucose, gluconate, and mixed carbon sources using isotopic tracer experiments.

Type: Genetic perturbation and flux analysis

Experiment: Use promoter reporters or RNA-seq in wild type and hexR backgrounds to quantify how glucose, gluconate, 2-ketogluconate, and KDPG-related perturbations regulate the edd-glk-gltR2-gltS operon.

Type: Operon regulation

Deep Research

Asta

(glk-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 19 citations 2026-07-06T05:24:34.603959

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 19
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.755)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.718)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.
  • Snippet 2 (score: 0.656)
    > (I think match should be teal and mismatch -red?)
    > The legend was previously mismatched with the labels. This has been corrected in the new uploaded figure . 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section?
    > The reviewer's presumption is correct; we had stated the date of data retrieval in the caption of Table 1, but we agree it should instead be stated centrally in the Methods. We have now added it to the Methods section as well, for clarity (Lines 696-700) 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copypasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...")
    > We thank the reviewer for catching this accidental insertion. We have now removed the spurious fragment.
    > 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > We have removed this speculation in the revised submission.
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or under-annotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets.

[3] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.702)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[4] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.698)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[5] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[6] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.688)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[7] Role of histone-lysine N-methyltransferase 2D (KMT2D) in MEK-ERK signaling-mediated epigenetic regulation: a phosphoproteomics perspective

  • Authors: Sreeshma Ravindran Kammarambath, Leona Dcunha, Athira Perunelly Gopalakrishnan, Amal Fahma, N. Krishna et al.
  • Year: 2025
  • Venue: Frontiers in Bioinformatics
  • URL: https://www.semanticscholar.org/paper/0ac0729148aff3d839e6a15984e11532e9e740f9
  • DOI: 10.3389/fbinf.2025.1683469
  • PMID: 41341998
  • PMCID: 12669113
  • Citations: 3
  • Summary: The phosphoregulatory network of Histone-lysine N-methyltransferase 2D is delineated, positioning it as a dynamic epigenetic effector modulated by MEK-ERK signaling, with broader implications for cancer and developmental disorders.
  • Evidence snippets:
  • Snippet 1 (score: 0.687)
    > Each protein was mapped to its corresponding gene symbol based on the HGNC (downloaded on 30.05.2023) and to its corresponding UniProt (13.04.2023) (UniProt, 2023) accessions using our in-built mapping tool to ensure consistent and standardized annotation. We conducted the analysis using the methodologies outlined in (Sanjeev et al., 2024). The overall workflow used in this study is outlined in Figure 1.

[8] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.685)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[9] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.683)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[10] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.682)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[11] GOnet: a tool for interactive Gene Ontology analysis

  • Authors: M. Pomaznoy, Brendan Ha, Bjoern Peters
  • Year: 2018
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  • DOI: 10.1186/s12859-018-2533-3
  • PMID: 30526489
  • PMCID: 6286514
  • Citations: 247
  • Influential citations: 17
  • Summary: The open-source GOnet web-application is created, which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis and provides insight into the functional interconnection of the submitted entries.
  • Evidence snippets:
  • Snippet 1 (score: 0.677)
    > In a basic workflow, the GOnet application receives a list of gene symbols, protein symbols, or protein IDs (UniProt IDs) as an input, and outputs a graph (an example given in Fig. 1). There are various input parameters which will affect the actual structure of the graph visualized and its appearance. The first main user choice is which GO terms the genes are annotated against:
    > 1. GO terms statistically significantly over-represented in the gene list submitted. 2. A predefined subset (also known as 'GO slim'), or a user-supplied list of terms.
    > In the first case the analysis will be referred to as an 'enrichment' analysis, in the second as an 'annotation' analysis.
    > Input parameters 1) Gene list. A mandatory input parameter containing the genes/proteins of interest. Currently human and mouse data is supported. An example of a human gene list might look like this:
    > Fig. 1 Sample network output generated by GOnet application. Gene differentially expressed in CD4 Bulk Memory T cells in Latent TB patients compared to healthy controls were used as an example [22] The gene list can also be accompanied with a contrast value. For example, This contrast value can be any decimal number, such as the log-fold change of gene expression between two conditions. This is merely a visualization enhancement. If the value is supplied it can be used later to differentially color specific genes in the graph (note different colors of gene nodes in Fig. 1), and visually indicate up-or down-regulation of specific genes and gene clusters.
    > The application can process common gene symbols (like in the example above), UniProt IDs, and MGI Accession IDs (mouse only). The former type of ID (gene symbols), although is the most human friendly, can unfortunately be ambiguous. For example, AIM1 can mean 'absent in melanoma' (also called CRYBG1) or 'Aurora and Ipl1-like midbody-associated protein' (also known as AURKB). Due to this ambiguity UniProt IDs or MGI accession IDs (for mouse) are preferred.
    > 2) GO namespace. Can be any of 'biological process', 'molecular function' or 'cellular component'.

[12] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.669)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[13] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[14] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[15] Functional annotation of parasitic worm genomes, by assigning protein names and GO terms

  • Authors: Avril Coghlan, M. Berriman
  • Year: 2018
  • Venue: Unknown venue
  • URL: https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  • DOI: 10.1038/protex.2018.055
  • Citations: 1
  • Summary: A computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and Go terms from orthologues in other species.
  • Evidence snippets:
  • Snippet 1 (score: 0.667)
    > Given a set of predicted protein-coding genes for a newly sequenced genome, functional annotation involves assigning putative functions to the predicted genes. Two ways in which this can be done are assigning protein names and Gene Ontology (GO;Gene Ontology Consortium, 2010) terms to the predicted proteins. Here we describe a computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and GO terms from orthologues in other species.
    > When assigning protein names, UniProt protein naming rules (www.uniprot.org/docs/nameprot) are followed where possible. This recommends that a good and stable name for a protein is "as neutral as possible"; that a protein name "should be, as far as possible, unique and attributed to all orthologs"; and a protein name "should not contain a specific characteristic of the protein, and in particular it should not reflect the function or role of the protein, nor its subcellular location, its domain structure, its tissue specificity, its molecular weight or its species of origin".
    > In our protocol, a protein name is assigned to each predicted protein based on curated names in UniProt (Bairoch & Apweiler, 2000) for human, zebrafish, Drosophila melanogaster, Caenorhabditis elegans, and Schistosoma mansoni orthologues identified from a database of gene families (e.g. built using Ensembl Compara; Vilella et al. 2009), or (if no information is found from orthologues) based on InterPro (Hunter et al. 2012) domains. Figure 1 shows an example of using our protein naming pipeline for four Strongyloides ratti genes that belong to the tubulin polyglutamylase family (underlined in pink), where four different protein names were assigned to them (in pink), based on names of their C. elegans or human orthologues.
    > Since each of the S. ratti genes belonged to a different subfamily of the tubulin polyglutamylase family, they were assigned different names.

[16] The alpha-ketoacid dehydrogenase complexes of Drosophila melanogaster.

  • Authors: Steven J. Marygold
  • Year: 2024
  • Venue: microPublication Biology
  • URL: https://www.semanticscholar.org/paper/50942e603e0e14ee9195c0d7cb52db11a521f964
  • DOI: 10.17912/micropub.biology.001209
  • PMID: 38741935
  • PMCID: 11089389
  • Citations: 2
  • Summary: This work identifies and classify the genes encoding all Drosophila AKDHC subunits, update their functional annotations and integrate this work into the FlyBase database.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > Symbol: gene symbol in FlyBase -asterisk (*) indicates a gene with testis-specific expression. CG#: gene annotation ID in FlyBase. Function: component and associated EC number (where available/applicable). Key reference(s) for initial identification or genetic characterization (in a metabolic context): 1. Gruntenko et al. 1998;2. Chen et al. 2008;3. Yoon et al. 2017;4. Yap et al. 2021a;5. Yap et al. 2021b;6. Whittle et al. 2023;7. González Morales et al. 2023;8. Homem et al. 2014;9. Bonnay et al. 2020;10. Ivanova et al. 2004;11. Boyko et al. 2020;12. Liu et al. 2017;13. Li et al. 2020;14. Devilliers et al. 2021;15. Goyal et al. 2022;16. Huang et al. 2022;17. Plaçais et al. 2017;18. Dung et al. 2018;19. Rabah et al. 2023;20. Klenz et al. 1995;21. Katsube et al. 1997;22. Gándara et al. 2019;23. Lambrechts et al. 2019;24. Lee et al. 2022;25. Chen et al. 2006;26. Kim et al. 2023;27. Tsai et al. 2020. Human ortholog: gene symbol at HGNC, with % amino acid identity between the encoded protein and the Drosophila protein. Human disease: OMIM symbol for disease(s) associated with the human gene (Amberger et al. 2019) -see Extended Data File 1 for details.

[17] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.658)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[18] Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser

  • Authors: Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała
  • Year: 2025
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  • DOI: 10.1186/s12864-025-12329-8
  • PMID: 41408139
  • PMCID: 12822150
  • Citations: 2
  • Summary: An R script is developed that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database and provides a single list of updated symbols with annotation about their ambiguity.
  • Evidence snippets:
  • Snippet 1 (score: 0.658)
    > Gene nomenclature contains current official symbols and various numbers of synonyms, which pose a challenge to integrating genomic data and increase the probability that different genes share the same symbol. Therefore, we retrieved identifiers assigned to all protein-coding genes in human, mouse, rat and pig genomes that are available in the Ensembl genome browser (release 113) to assess the number of genes, compare species and identify ambiguous symbols. Results: Our analysis revealed that the total number of symbols, both official symbols and synonyms, used to identify protein-coding genes ranges from 16,600 in pigs to 64,580 in mice. Furthermore, the gene nomenclature is not complete because there are also genes without an assigned symbol, which indicates gaps in understanding protein-coding genes, especially in pigs. We also found a large number of gene symbols that map to more than one gene. These symbols might complicate the identification of about 10% of rat and mouse genes and 18% of human protein-coding genes. A simple solution for this problem is the usage of stable gene IDs assigned by scientific institutions and committees (Ensembl, NCBI, RGD, HGNC and VGNC) provided that the genomic information associated with these IDs is retrieved directly from proprietary databases containing the most accurate data. Finally, although gene symbols may pose a problem with unequivocal identification of genes, there are instances when no other identifiers are available in the literature. Therefore, we have developed an R script performing search of the Ensembl database and integrating data to provide a single list of updated symbols with annotation about their ambiguity. Conclusions: Gene symbols are not always reliable and should be reported together with stable IDs to enable unequivocal identification of genes. Therefore, data containing only gene symbols should be used cautiously to avoid misidentification of genes. A solution for this problem is our R script REgeness that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database.

[19] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  3. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  4. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  5. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  6. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  7. Sreeshma Ravindran Kammarambath, Leona Dcunha, Athira Perunelly Gopalakrishnan, Amal Fahma, N. Krishna et al. (2025). Role of histone-lysine N-methyltransferase 2D (KMT2D) in MEK-ERK signaling-mediated epigenetic regulation: a phosphoproteomics perspective. Frontiers in Bioinformatics. https://www.semanticscholar.org/paper/0ac0729148aff3d839e6a15984e11532e9e740f9
  8. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  9. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  10. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  11. M. Pomaznoy, Brendan Ha, Bjoern Peters (2018). GOnet: a tool for interactive Gene Ontology analysis. BMC Bioinformatics. https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  12. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  13. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  14. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  15. Avril Coghlan, M. Berriman (2018). Functional annotation of parasitic worm genomes, by assigning protein names and GO terms. https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  16. Steven J. Marygold (2024). The alpha-ketoacid dehydrogenase complexes of Drosophila melanogaster.. microPublication Biology. https://www.semanticscholar.org/paper/50942e603e0e14ee9195c0d7cb52db11a521f964
  17. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  18. Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała (2025). Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser. BMC Genomics. https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  19. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508

Falcon

(glk-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 28 citations 2026-03-21T14:35:39.766687

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: functional annotation of glk (UniProt Q88P42; locus PP_1011) in Pseudomonas putida KT2440

Executive summary (verified target identity)

The UniProt target Q88P42 corresponds to glk in Pseudomonas putida KT2440 (ordered locus PP_1011) encoding the cytosolic enzyme glucokinase (Glk), which phosphorylates imported glucose to glucose‑6‑phosphate (G6P) as the entry step of the phosphorylative branch of glucose catabolism. This identity is directly supported by KT2440-focused studies that name glk (PP_1011) as glucokinase and place it in the edd–glk operon/cluster and broader glucose-catabolism gene neighborhood. (castillo2007convergentperipheralpathways pages 5-6, castillo2008asetof pages 1-2, daddaoua2009regulationofglucose pages 1-2, nguyen2024investigatinganaerobicmetabolism pages 27-30)


1) Key concepts and definitions (current understanding)

1.1 Gene product and enzymatic function

Glucokinase (Glk) is a sugar kinase that catalyzes phosphorylation of D-glucose to D-glucose‑6‑phosphate, coupling this to cellular energy metabolism (ATP consumption). In KT2440, glucose that is transported into the cytosol via an ABC transporter is then phosphorylated by Glk to G6P, which is further converted toward 6‑phosphogluconate by glucose‑6‑phosphate dehydrogenase (Zwf) and 6‑phosphogluconolactonase (Pgl). (chen2024gnurrepressesthe pages 1-3, daddaoua2009regulationofglucose pages 1-2)

Functional reaction (as used in KT2440 glucose assimilation):
- D‑glucose + ATP → D‑glucose‑6‑phosphate + ADP (commonly annotated as EC 2.7.1.2; EC number is not explicitly stated in the retrieved excerpts, but the enzyme is explicitly called “glucokinase” and described as catalyzing glucose→G6P). (chen2024gnurrepressesthe pages 1-3, daddaoua2009regulationofglucose pages 1-2)

1.2 KT2440 glucose metabolism context: convergent “peripheral” routes

A defining feature of Pseudomonas glucose metabolism is the co-existence of multiple peripheral routes that converge at 6‑phosphogluconate. In KT2440, glucose may enter central metabolism via:
1) a cytosolic phosphorylative route: glucose import → Glk → G6P → Zwf/Pgl → 6‑phosphogluconate; and
2) periplasmic oxidative routes producing gluconate and 2‑ketogluconate that are subsequently imported and metabolized to 6‑phosphogluconate. (castillo2007convergentperipheralpathways pages 1-2, nikel2015pseudomonasputidakt2440 pages 1-2)

1.3 EDEMP cycle: integration of ED, PP and upper‑EMP reactions

KT2440 does not run a canonical linear Embden–Meyerhof–Parnas (EMP) glycolysis due to missing 6‑phosphofructo‑1‑kinase activity; instead, glucose is processed through an integrated cyclic architecture often described as an EDEMP cycle, combining Entner–Doudoroff (ED) reactions with upper‑EMP/gluconeogenic steps and pentose‑phosphate (PP) reactions. In this framework, Glk supplies G6P that can feed these interconnected modules, and carbon can be recycled from triose phosphates back to hexose phosphates. (nikel2015pseudomonasputidakt2440 pages 1-2, nikel2015pseudomonasputidakt2440 media d74e2f18)


2) KT2440-specific functional evidence for glk (PP_1011)

2.1 Genomic context and operon organization

Multiple sources report that glk is co-transcribed with edd (encoding 6‑phosphogluconate dehydratase), i.e. an edd–glk operon. The operon/cluster extends to include gltR2/gltS (transport-associated regulation), and edd is divergently transcribed relative to gap‑1 in the region. (castillo2007convergentperipheralpathways pages 5-6, daddaoua2009regulationofglucose pages 1-2)

This physical linkage is consistent with functional coupling: the glucokinase entry step (Glk) and the subsequent ED processing step (Edd) are transcriptionally coordinated. (castillo2007convergentperipheralpathways pages 5-6, castillo2008asetof pages 1-2)

2.2 Direct biochemical/physiological evidence: inducible Glk activity

In KT2440, enzymatic assays from cell extracts showed that glucokinase activity is inducible by growth on glucose. Specifically, glucokinase (and gluconokinase) activities were low during growth on citrate but increased ~10‑fold during growth on glucose, supporting that Glk is expressed and catalytically active under glucose conditions (and not only inferred from sequence annotation). (castillo2007convergentperipheralpathways pages 5-6)

2.3 Quantitative physiology and flux: relative importance of the Glk branch

13C flux analysis and physiology in KT2440 indicate that glucose catabolism strongly favors periplasmic oxidation under tested conditions: approximately 90% of consumed glucose was reported to proceed through the oxidative route (via gluconate) into central metabolism as 6‑phosphogluconate, while the remaining fraction enters via other routes, including the cytosolic phosphorylation route that depends on Glk. (nikel2015pseudomonasputidakt2440 pages 1-2)

A complementary flux/physiology study reported a glucose uptake rate ~6 mmol·g⁻¹·h⁻¹, growth rate 0.58 h⁻¹, and biomass yield 0.44 g biomass per g carbon for the wild type under the studied condition set. (castillo2007convergentperipheralpathways pages 1-2)

2.4 Mutant phenotype: impact of perturbing the glucokinase pathway

A KT2440 glk mutant exhibited measurable impairments in carbon consumption and biomass yield in the referenced physiological analyses; one excerpted dataset reports total carbon consumption of ~5.44 mmol·g⁻¹·h⁻¹ and biomass yield of ~0.35 g biomass per g carbon consumed for the glk mutant. (castillo2007convergentperipheralpathways pages 5-6)

Separate work on catabolite repression and pathway integration explicitly used a glk mutant strain to interrogate growth and regulation on mixed substrates (glucose + toluene), further emphasizing that glk is a tractable genetic node with measurable phenotypic consequences. (castillo2007convergentperipheralpathways pages 6-8)


3) Cellular localization and where the reaction occurs

Glk functions in the cytoplasm, acting on glucose that has been transported into the cytosol. This is explicitly described as glucose being transported into the cytoplasm and “subsequently phosphorylated by glucokinase (Glk) to glucose‑6‑phosphate.” (chen2024gnurrepressesthe pages 1-3)

This contrasts with the periplasmic oxidative route (glucose→gluconate/2‑ketogluconate) which is initiated outside the cytosol and later converges internally at phosphorylated intermediates. (nikel2015pseudomonasputidakt2440 pages 1-2)


4) Regulation and pathway control (expert-level synthesis)

4.1 HexR-centered control of ED and glucose assimilation nodes

Mechanistic regulatory work in Pseudomonas identifies HexR as a central transcriptional regulator controlling expression of promoters including those for ED/glucose catabolism (e.g., Pedd and Pzwf/Pgap). HexR binds target promoters with nanomolar-range affinity and is released by the ED intermediate KDPG (2‑keto‑3‑deoxy‑6‑phosphogluconate), linking metabolic state to gene expression. This regulatory logic couples expression of the edd–glk operon to ED pathway flux. (daddaoua2009regulationofglucose pages 1-2)

Consistently, systems-level analysis describing peripheral glucose control reports that HexR controls genes encoding glucokinase and glucose‑6‑phosphate dehydrogenase, integrating the Glk entry step with downstream conversion of G6P toward 6‑phosphogluconate. (castillo2008asetof pages 1-2)

4.2 Multi-regulator architecture controlling glucose/gluconate catabolism

Beyond HexR, the KT2440 glucose catabolic region is described as being under multiple transcriptional regulators, including transport-associated regulation (e.g., GltR/GltS) and regulators of gluconate/2‑ketogluconate utilization (e.g., PtxS, GnuR in related frameworks). (castillo2008asetof pages 1-2, daddaoua2009regulationofglucose pages 1-2)

Recent 2024 development (high priority): A KT2440 study used multi‑omics to define the GnuR regulon, showing that GnuR directly represses catabolic genes involved in the Entner–Doudoroff and peripheral glucose/gluconate metabolism pathways, and that expression of the glucose catabolism genes is induced by both glucose and gluconate. This work modernizes the regulatory map around glucose utilization while maintaining Glk’s core biochemical assignment as the glucose→G6P phosphorylation step. (Publication date: Nov 2024; URL: https://doi.org/10.1111/1751-7915.70059) (chen2024gnurrepressesthe pages 1-3)


5) Recent developments (2023–2024 prioritized)

5.1 2024: Systems-level regulatory dissection (GnuR)

The 2024 study on GnuR provides a KT2440-focused, contemporary update: it reports that glucose and gluconate induce expression of glucose catabolism genes and identifies direct repression targets of GnuR within ED/peripheral metabolism, improving the mechanistic basis for engineering glucose catabolism. (Nov 2024; https://doi.org/10.1111/1751-7915.70059) (chen2024gnurrepressesthe pages 1-3)

5.2 2024: Bioelectrochemical / electro-fermentation conditions and uptake-route testing

A 2024 study of anaerobic glucose uptake in a bioelectrochemical system explicitly frames energy accounting around “sugar phosphorylation in the upper part of metabolism (at the level of glucokinase…)” and tests uptake-route mutants to determine whether uptake stoichiometry limits electro-fermentation performance. In the studied setup, deletion of uptake routes affected sugar consumption/current output, but secreted acetate levels were not significantly altered among strains—suggesting other constraints (e.g., energy limitation or regulation) rather than the glucokinase stoichiometry per se. (Publication date: Nov 2024; URL: https://doi.org/10.1111/1751-7915.14375) (nguyen2024investigatinganaerobicmetabolism pages 27-30)

5.3 2024: Metabolic engineering of glucose routing for improved mcl‑PHA production

A 2024 metabolic engineering study in P. putida reported that modifying glucose metabolism and regulatory nodes can markedly improve medium-chain-length PHA titers (e.g., increases in titer reported for combined knockouts and regulatory modifications), reflecting ongoing real-world implementation of glucose catabolism rewiring in Pseudomonas production strains. (Publication date: Nov 2024; URL: https://doi.org/10.3390/cimb46110761) (chen2024gnurrepressesthe pages 1-3)

Note on limitations: Using the provided tool-accessible corpus, I could not retrieve 2023–2024 sources specifically focused on structural mechanism or family-wide biochemical kinetics for PF02685 bacterial glucokinases; therefore, mechanistic/structural claims beyond the demonstrated pathway role are not made here. (nikel2015pseudomonasputidakt2440 pages 1-2, chen2024gnurrepressesthe pages 1-3)


6) Current applications and real-world implementations

6.1 Glk as an engineering control point for sugar metabolism “wiring”

The edd–glk node is repeatedly leveraged as a defined genetic “handle” to rewire upper sugar metabolism. For example, refactoring/expanding sugar utilization (e.g., cellobiose/xylose modules) relies on intracellular glucose being phosphorylated and routed into central metabolism via the native upper network in KT2440 derivatives. (https://doi.org/10.1016/j.ymben.2018.05.019; 2018) (bujdos2021inženýrstvípseudomonasputida pages 40-43)

6.2 Glk in modular glycolysis implantation frameworks

A synthetic biology strategy (“GlucoBricks”) reported measurable glucokinase specific activities in engineered strains and evaluated portability of glycolytic modules to Gram-negative hosts including Pseudomonas. While the excerpted kinetic data are reported for engineered E. coli backgrounds, this work frames Glk as the first committed step enabling engineered glycolytic flux. (https://doi.org/10.1021/acssynbio.6b00230; Feb 2017) (sanchezpascuala2017refactoringtheembden–meyerhof–parnas pages 4-6)

6.3 Bioelectrochemical systems and redox/energy-aware bioprocessing

Under electro-fermentative conditions, selection/engineering of glucose uptake routes (including those requiring cytosolic phosphorylation via glucokinase) is being studied to improve carbon turnover and product formation; route mutants were used to evaluate whether the energetic/redox demands of uptake and phosphorylation are limiting. (https://doi.org/10.1111/1751-7915.14375; Nov 2024) (nguyen2024investigatinganaerobicmetabolism pages 27-30)


7) Relevant statistics and quantitative data (from cited studies)

  • ~10-fold induction of glucokinase activity in glucose-grown vs citrate-grown KT2440 cells (enzyme assays in cell extracts). (https://doi.org/10.1128/JB.00203-07; Jul 2007) (castillo2007convergentperipheralpathways pages 5-6)
  • Wild-type KT2440 physiology under studied conditions: glucose uptake ~6 mmol·g⁻¹·h⁻¹, growth rate 0.58 h⁻¹, biomass yield 0.44 g biomass/g carbon. (https://doi.org/10.1128/JB.00203-07; Jul 2007) (castillo2007convergentperipheralpathways pages 1-2)
  • glk mutant example physiology: total carbon consumption ~5.44 mmol·g⁻¹·h⁻¹, biomass yield ~0.35 g/g carbon consumed. (https://doi.org/10.1128/JB.00203-07; Jul 2007) (castillo2007convergentperipheralpathways pages 5-6)
  • Flux partitioning: ~90% of glucose consumed routed through the periplasmic oxidative pathway (gluconate route) in KT2440 under the studied conditions; EDEMP cycling includes recycling of triose phosphates to hexose phosphates (~10% reported in the described framework). (https://doi.org/10.1074/jbc.M115.687749; Oct 2015) (nikel2015pseudomonasputidakt2440 pages 1-2)

Unavailable in retrieved excerpts: explicit Glk kinetic constants (Km, kcat) and direct substrate-specificity panels for purified P. putida Glk were not accessible from the available text snippets. (nikel2015pseudomonasputidakt2440 pages 17-19)


8) Visual evidence: pathway architecture supporting Glk functional context

The schematic depiction of the EDEMP cycle in KT2440 provides a concise visual explanation of how ED, PP and upper‑EMP reactions integrate in KT2440 glucose metabolism, contextualizing where the Glk-dependent G6P entry step can feed into the broader network. (nikel2015pseudomonasputidakt2440 media d74e2f18)


Gene: glk (PP_1011)

Protein: Glucokinase (Glk)

Cellular location: Cytosol; acts on cytosolic glucose after import. (chen2024gnurrepressesthe pages 1-3)

Primary biochemical function: Phosphorylates glucose to glucose‑6‑phosphate to initiate the cytosolic phosphorylative glucose assimilation route in KT2440; downstream routing proceeds via Zwf/Pgl to 6‑phosphogluconate and ED/EDEMP-associated metabolism. (chen2024gnurrepressesthe pages 1-3, daddaoua2009regulationofglucose pages 1-2, nikel2015pseudomonasputidakt2440 pages 1-2)

Genomic/pathway context: Part of an edd–glk operon/cluster linked to ED entry, co-regulated with glucose transport and ED genes (including gltR2/gltS), embedded within the multi-route peripheral glucose utilization architecture of Pseudomonas. (castillo2007convergentperipheralpathways pages 5-6, daddaoua2009regulationofglucose pages 1-2)

Regulatory context (high confidence): Expression of edd/glk-linked functions is controlled by HexR, with effector KDPG modulating promoter binding; additional regulators (e.g., GnuR) shape global glucose/gluconate catabolism responses in KT2440. (daddaoua2009regulationofglucose pages 1-2, chen2024gnurrepressesthe pages 1-3)


Evidence summary table

Claim / topic Evidence-backed summary Quantitative findings Example application / implementation (year; DOI/URL) Key sources (context IDs)
Gene/protein identity glk in Pseudomonas putida KT2440 encodes the cytosolic glucokinase Glk; locus PP_1011; this matches UniProt Q88P42 and the glucose-phosphorylation branch of sugar catabolism. PP_1011 explicitly linked to Glk in KT2440 pathway/genome context. Used as a defined engineering target in KT2440 derivatives lacking glk to block glucose assimilation (2025; https://doi.org/10.1093/synbio/ysaf012). (nguyen2024investigatinganaerobicmetabolism pages 27-30, escapa2013theroleof pages 37-41, chen2024gnurrepressesthe pages 1-3)
Enzymatic reaction Glk catalyzes ATP-dependent phosphorylation of glucose → glucose-6-phosphate (G6P), feeding upper central metabolism; downstream G6P is oxidized by Zwf/Pgl toward 6-phosphogluconate. Reaction is described qualitatively; no Glk-specific Km/kcat retrieved in the available evidence. Central step in glucose entry retained or rewired in KT2440 metabolic engineering for mixed-sugar use (2018; https://doi.org/10.1016/j.ymben.2018.05.019). (chen2024gnurrepressesthe pages 1-3, daddaoua2009regulationofglucose pages 1-2, nikel2015pseudomonasputidakt2440 pages 1-2)
Pathway role in KT2440 In KT2440, glucose can enter metabolism either by direct cytosolic phosphorylation via Glk or by periplasmic oxidation to gluconate/2-ketogluconate before convergence at 6-phosphogluconate; Glk is part of the phosphorylative branch. Approx. 90% of consumed glucose was reported to proceed through the periplasmic oxidative route, implying a minority enters via direct phosphorylation under tested conditions. Important for quantitative physiology and flux-guided redesign of KT2440 glucose metabolism (2015; https://doi.org/10.1074/jbc.M115.687749). (nikel2015pseudomonasputidakt2440 pages 1-2, castillo2007convergentperipheralpathways pages 1-2)
Operon / genomic context glk is physically and transcriptionally linked with edd (6-phosphogluconate dehydratase), i.e. an edd-glk operon/cluster; broader local context includes gltR2/gltS and nearby glucose-uptake genes. glk and edd were reported to overlap and be transcriptionally coupled. This native organization motivated “GlucoBrick”/upper-glycolysis refactoring efforts in Gram-negative hosts including P. putida (2017; https://doi.org/10.1021/acssynbio.6b00230). (castillo2007convergentperipheralpathways pages 5-6, castillo2008asetof pages 1-2, daddaoua2009regulationofglucose pages 1-2)
Regulation by HexR HexR controls expression of glucose-catabolic genes including the edd-glk branch and the zwf/pgl/eda branch; induction links glucose sensing to ED-pathway entry. HexR binds target promoters with nanomolar-range affinity and is released by the ED intermediate KDPG. Regulatory understanding is used for pathway rewiring and control of carbon flux in KT2440 chassis engineering (2024; https://doi.org/10.3390/cimb46110761). (castillo2008asetof pages 1-2, daddaoua2009regulationofglucose pages 1-2)
Regulation by GnuR / peripheral glucose control Recent multi-omics work shows GnuR directly represses genes of peripheral glucose/gluconate catabolism and ED-pathway functions in KT2440; Glk participates in the induced glucose-response program. Expression of catabolic/TF genes was significantly induced by glucose and gluconate; exact fold change not stated in retrieved excerpt. Provides a modern regulatory framework for improving glucose utilization in KT2440 (2024; https://doi.org/10.1111/1751-7915.70059). (chen2024gnurrepressesthe pages 1-3)
Additional regulatory systems Glucose-catabolic gene expression around glk/edd is integrated with other regulators including GltR/GltS (transport-associated regulation) and PtxS / the kgu branch for 2-ketogluconate metabolism. No Glk-specific kinetic constants provided, but regulatory coupling to transporter/peripheral oxidation branches is well supported. Relevant for engineering strains that favor specific uptake routes under aerobic or electro-fermentative conditions (2024; https://doi.org/10.1111/1751-7915.14375). (castillo2007convergentperipheralpathways pages 6-8, castillo2008asetof pages 1-2, daddaoua2009regulationofglucose pages 1-2)
Enzyme induction / activity evidence Glucokinase activity in KT2440 cell extracts is inducible by glucose, supporting Glk as an active enzyme rather than a purely inferred annotation. ~10-fold increase in glucokinase and gluconokinase activities in glucose-grown cells versus citrate-grown cells. Biochemical activity evidence underpins strain-design choices when redirecting sugar catabolism (2007; https://doi.org/10.1128/JB.00203-07). (castillo2007convergentperipheralpathways pages 5-6)
Mutant phenotype Disrupting the glucokinase branch impairs efficient glucose assimilation; foundational studies concluded the glucokinase pathway is functionally important and in some tested genetic backgrounds essential for growth on glucose. Example values reported for a glk mutant: total carbon consumption about 5.44 mmol g⁻¹ h⁻¹ and biomass yield about 0.35 g biomass/g carbon consumed. Δglk strains are deliberately used as chassis components to eliminate glucose growth in synthetic consortia or substrate-partitioning designs (2025; https://doi.org/10.1093/synbio/ysaf012). (castillo2007convergentperipheralpathways pages 5-6, castillo2007convergentperipheralpathways pages 1-2)
Flux architecture / EDEMP cycle Glk supplies G6P to the unusual EDEMP cycle of KT2440, where ED, upper EMP/gluconeogenic reactions, and PP-pathway enzymes are integrated to support NADPH generation and flexible carbon redistribution. About 10% of triose phosphates were estimated to recycle back to hexose phosphates in the EDEMP architecture. Core concept for systems biology, modeling, and redox-aware chassis engineering in KT2440 (2015; https://doi.org/10.1074/jbc.M115.687749). (nikel2015pseudomonasputidakt2440 pages 1-2, nikel2015pseudomonasputidakt2440 media d74e2f18, nikel2015pseudomonasputidakt2440 media d533aa4f)
Engineering for sugar co-utilization Native glk (PP_1011) is exploited in refactoring upper sugar metabolism so intracellular glucose released from cellobiose or imported sugars can be funneled into central metabolism. Engineering studies identify PP_1011 as the native glucokinase node supporting growth on glucose-containing mixtures. Co-utilization of cellobiose, xylose, and glucose in engineered KT2440/EM42 (2018; https://doi.org/10.1016/j.ymben.2018.05.019). (bujdos2021inženýrstvípseudomonasputida pages 40-43, sanchezpascuala2017refactoringtheembden–meyerhof–parnas pages 4-6)
Engineering for product formation Manipulating glucose catabolic routing upstream/downstream of Glk improves production phenotypes such as medium-chain-length PHA in P. putida. One 2024 study reported 33.7% higher mcl-PHA titer in a ΔgcdΔgltA mutant and up to 117.5% increase in an additional regulatory mutant background, illustrating the leverage of glucose-routing interventions around the Glk branch. Glucose-pathway modification for enhanced PHA synthesis (2024; https://doi.org/10.3390/cimb46110761). (chen2024gnurrepressesthe pages 1-3)
Bioelectrochemical / anaerobic relevance Under electro-fermentative conditions, the Glk-dependent phosphorylation route remains part of the energetic accounting for glucose use, and route-specific mutants helped test whether uptake stoichiometry limits anaerobic turnover. Deletion of individual sugar-uptake routes altered sugar consumption/current output, but secreted acetate concentrations were not significantly different among strains in the reported setup. Anaerobic glucose uptake in a bioelectrochemical system (2024; https://doi.org/10.1111/1751-7915.14375). (nguyen2024investigatinganaerobicmetabolism pages 27-30)

Table: This table summarizes evidence-backed functional annotation for Pseudomonas putida KT2440 glucokinase Glk (PP_1011; UniProt Q88P42), including reaction, pathway role, regulation, quantitative findings, and representative engineering applications. It is useful as a compact reference for assigning function and pathway context to the gene.

References

  1. (castillo2007convergentperipheralpathways pages 5-6): Teresa del Castillo, Juan L. Ramos, José J. Rodríguez-Herva, Tobias Fuhrer, Uwe Sauer, and Estrella Duque. Convergent peripheral pathways catalyze initial glucose catabolism inpseudomonas putida: genomic and flux analysis. Jul 2007. URL: https://doi.org/10.1128/jb.00203-07, doi:10.1128/jb.00203-07. This article has 310 citations and is from a peer-reviewed journal.

  2. (castillo2008asetof pages 1-2): Teresa del Castillo, Estrella Duque, and Juan L. Ramos. A set of activators and repressors control peripheral glucose pathways in pseudomonas putida to yield a common central intermediate. Journal of Bacteriology, 190:2331-2339, Apr 2008. URL: https://doi.org/10.1128/jb.01726-07, doi:10.1128/jb.01726-07. This article has 128 citations and is from a peer-reviewed journal.

  3. (daddaoua2009regulationofglucose pages 1-2): A. Daddaoua, T. Krell, and J. Ramos. Regulation of glucose metabolism in pseudomonas. The Journal of Biological Chemistry, 284:21360-21368, Jun 2009. URL: https://doi.org/10.1074/jbc.m109.014555, doi:10.1074/jbc.m109.014555. This article has 106 citations.

  4. (nguyen2024investigatinganaerobicmetabolism pages 27-30): HAV Nguyen. Investigating anaerobic metabolism of pseudomonas putida using bioelectrochemical cultivation. Unknown journal, 2024.

  5. (chen2024gnurrepressesthe pages 1-3): Wenbo Chen, Rao Ma, Yong Feng, Yunzhu Xiao, Agnieszka Sekowska, Antoine Danchin, and Conghui You. Gnur represses the expression of glucose and gluconate catabolism in pseudomonas putida kt2440. Microbial Biotechnology, Nov 2024. URL: https://doi.org/10.1111/1751-7915.70059, doi:10.1111/1751-7915.70059. This article has 2 citations and is from a peer-reviewed journal.

  6. (castillo2007convergentperipheralpathways pages 1-2): Teresa del Castillo, Juan L. Ramos, José J. Rodríguez-Herva, Tobias Fuhrer, Uwe Sauer, and Estrella Duque. Convergent peripheral pathways catalyze initial glucose catabolism inpseudomonas putida: genomic and flux analysis. Jul 2007. URL: https://doi.org/10.1128/jb.00203-07, doi:10.1128/jb.00203-07. This article has 310 citations and is from a peer-reviewed journal.

  7. (nikel2015pseudomonasputidakt2440 pages 1-2): Pablo I. Nikel, Max Chavarría, Tobias Fuhrer, Uwe Sauer, and Víctor de Lorenzo. Pseudomonas putida kt2440 strain metabolizes glucose through a cycle formed by enzymes of the entner-doudoroff, embden-meyerhof-parnas, and pentose phosphate pathways. Journal of Biological Chemistry, 290:25920-25932, Oct 2015. URL: https://doi.org/10.1074/jbc.m115.687749, doi:10.1074/jbc.m115.687749. This article has 427 citations and is from a domain leading peer-reviewed journal.

  8. (nikel2015pseudomonasputidakt2440 media d74e2f18): Pablo I. Nikel, Max Chavarría, Tobias Fuhrer, Uwe Sauer, and Víctor de Lorenzo. Pseudomonas putida kt2440 strain metabolizes glucose through a cycle formed by enzymes of the entner-doudoroff, embden-meyerhof-parnas, and pentose phosphate pathways. Journal of Biological Chemistry, 290:25920-25932, Oct 2015. URL: https://doi.org/10.1074/jbc.m115.687749, doi:10.1074/jbc.m115.687749. This article has 427 citations and is from a domain leading peer-reviewed journal.

  9. (castillo2007convergentperipheralpathways pages 6-8): Teresa del Castillo, Juan L. Ramos, José J. Rodríguez-Herva, Tobias Fuhrer, Uwe Sauer, and Estrella Duque. Convergent peripheral pathways catalyze initial glucose catabolism inpseudomonas putida: genomic and flux analysis. Jul 2007. URL: https://doi.org/10.1128/jb.00203-07, doi:10.1128/jb.00203-07. This article has 310 citations and is from a peer-reviewed journal.

  10. (bujdos2021inženýrstvípseudomonasputida pages 40-43): D Bujdoš. Inženýrství pseudomonas putida pro ko-utilizaci a zužitkování celobiózy s glukózou. Unknown journal, 2021.

  11. (sanchezpascuala2017refactoringtheembden–meyerhof–parnas pages 4-6): Alberto Sánchez-Pascuala, Víctor de Lorenzo, and Pablo I. Nikel. Refactoring the embden–meyerhof–parnas pathway as a whole of portable glucobricks for implantation of glycolytic modules in gram-negative bacteria. Feb 2017. URL: https://doi.org/10.1021/acssynbio.6b00230, doi:10.1021/acssynbio.6b00230. This article has 74 citations and is from a domain leading peer-reviewed journal.

  12. (nikel2015pseudomonasputidakt2440 pages 17-19): Pablo I. Nikel, Max Chavarría, Tobias Fuhrer, Uwe Sauer, and Víctor de Lorenzo. Pseudomonas putida kt2440 strain metabolizes glucose through a cycle formed by enzymes of the entner-doudoroff, embden-meyerhof-parnas, and pentose phosphate pathways. Journal of Biological Chemistry, 290:25920-25932, Oct 2015. URL: https://doi.org/10.1074/jbc.m115.687749, doi:10.1074/jbc.m115.687749. This article has 427 citations and is from a domain leading peer-reviewed journal.

  13. (escapa2013theroleof pages 37-41): I. F. Escapa, C. del Cerro, J. L. García, and M. A. Prieto. The role of glpr repressor in pseudomonas putida kt2440 growth and pha production from glycerol. Environmental microbiology, 15 1:93-110, May 2013. URL: https://doi.org/10.1111/j.1462-2920.2012.02790.x, doi:10.1111/j.1462-2920.2012.02790.x. This article has 131 citations and is from a domain leading peer-reviewed journal.

  14. (nikel2015pseudomonasputidakt2440 media d533aa4f): Pablo I. Nikel, Max Chavarría, Tobias Fuhrer, Uwe Sauer, and Víctor de Lorenzo. Pseudomonas putida kt2440 strain metabolizes glucose through a cycle formed by enzymes of the entner-doudoroff, embden-meyerhof-parnas, and pentose phosphate pathways. Journal of Biological Chemistry, 290:25920-25932, Oct 2015. URL: https://doi.org/10.1074/jbc.m115.687749, doi:10.1074/jbc.m115.687749. This article has 427 citations and is from a domain leading peer-reviewed journal.

Citations

  1. castillo2007convergentperipheralpathways pages 5-6
  2. castillo2007convergentperipheralpathways pages 1-2
  3. castillo2007convergentperipheralpathways pages 6-8
  4. chen2024gnurrepressesthe pages 1-3
  5. daddaoua2009regulationofglucose pages 1-2
  6. castillo2008asetof pages 1-2
  7. nguyen2024investigatinganaerobicmetabolism pages 27-30
  8. escapa2013theroleof pages 37-41
  9. https://doi.org/10.1111/1751-7915.70059
  10. https://doi.org/10.1111/1751-7915.14375
  11. https://doi.org/10.3390/cimb46110761
  12. https://doi.org/10.1016/j.ymben.2018.05.019;
  13. https://doi.org/10.1021/acssynbio.6b00230;
  14. https://doi.org/10.1111/1751-7915.14375;
  15. https://doi.org/10.1128/JB.00203-07;
  16. https://doi.org/10.1074/jbc.M115.687749;
  17. https://doi.org/10.1093/synbio/ysaf012
  18. https://doi.org/10.1016/j.ymben.2018.05.019
  19. https://doi.org/10.1074/jbc.M115.687749
  20. https://doi.org/10.1021/acssynbio.6b00230
  21. https://doi.org/10.1128/JB.00203-07
  22. https://doi.org/10.1128/jb.00203-07,
  23. https://doi.org/10.1128/jb.01726-07,
  24. https://doi.org/10.1074/jbc.m109.014555,
  25. https://doi.org/10.1111/1751-7915.70059,
  26. https://doi.org/10.1074/jbc.m115.687749,
  27. https://doi.org/10.1021/acssynbio.6b00230,
  28. https://doi.org/10.1111/j.1462-2920.2012.02790.x,

📚 Additional Documentation

Notes

(glk-notes.md)

glk Gene Research Notes

Identity

  • glk in Pseudomonas putida KT2440 corresponds to UniProt Q88P42 and locus PP_1011, and UniProt annotates the protein as glucokinase (EC 2.7.1.2) in the bacterial glucokinase family with cytoplasmic localization. [file:PSEPK/glk/glk-uniprot.txt "RecName: Full=Glucokinase"; "EC=2.7.1.2"; "SUBCELLULAR LOCATION: Cytoplasm"; "Belongs to the bacterial glucokinase family."]

Core Biochemistry

  • In KT2440, glucose imported into the cytoplasm is phosphorylated by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate. [PMID:17483213 Convergent peripheral pathways catalyze initial glucose catabolism in Pseudomonas putida: genomic and flux analysis, "Glucose is transported to the cytoplasm in a process mediated by an ABC uptake system encoded by open reading frames PP1015 to PP1018 and is then phosphorylated by glucokinase (encoded by the glk gene) and converted by glucose-6-phosphate dehydrogenase (encoded by the zwf genes) to 6-phosphogluconate."]

  • P. putida catabolizes glucose through three simultaneous routes that converge at 6-phosphogluconate, and the glucokinase branch is quantitatively important within that network. [PMID:17483213 Convergent peripheral pathways catalyze initial glucose catabolism in Pseudomonas putida: genomic and flux analysis, "glucose catabolism in Pseudomonas putida occurs through the simultaneous operation of three pathways that converge at the level of 6-phosphogluconate"; "although all three functioned simultaneously, the glucokinase pathway and the 2-ketogluconate loop were quantitatively more important than the direct phosphorylation of gluconate."]

  • The 2007 mutant and flux study concluded that the glucokinase pathway is required for growth on glucose in KT2440. [PMID:17483213 Convergent peripheral pathways catalyze initial glucose catabolism in Pseudomonas putida: genomic and flux analysis, "It can therefore be concluded that the glucokinase pathway is a sine qua non condition for P. putida to grow with glucose."]

Pathway Context And Regulation

  • glk is physically linked to the Entner-Doudoroff pathway because edd and glk form one operon in KT2440. [PMID:19506074 Regulation of Glucose Metabolism in Pseudomonas: The Phosphorylative Branch and Entner-Doudoroff Enzymes Are Regulated by a Repressor Containing a Sugar Isomerase Domain, "The edd and glk genes form another operon that encodes, respectively, 6-phosphogluconate dehydratase (the first enzyme of the Entner-Doudoroff pathway) and glucokinase (an enzyme of the glucose phosphorylative pathway)."]

  • Because of this genomic organization, the glucokinase pathway is co-induced with Entner-Doudoroff genes and is expressed even during growth on gluconate or 2-ketogluconate. [PMID:19506074 Regulation of Glucose Metabolism in Pseudomonas: The Phosphorylative Branch and Entner-Doudoroff Enzymes Are Regulated by a Repressor Containing a Sugar Isomerase Domain, "the glucokinase pathway genes are co-transcribed with Entner-Doudoroff pathway enzymes"; "the glucokinase pathway is induced when bacteria are exposed to gluconate and 2-ketogluconate."]

  • HexR controls the relevant promoters through sensing KDPG rather than glucose or glucose-6-phosphate, placing glk in a broader glucose-responsive Entner-Doudoroff regulatory module. [PMID:19506074 Regulation of Glucose Metabolism in Pseudomonas: The Phosphorylative Branch and Entner-Doudoroff Enzymes Are Regulated by a Repressor Containing a Sugar Isomerase Domain, "Binding of the Entner-Doudoroff pathway intermediate 2-keto-3-deoxy-6-phosphogluconate to HexR released the repressor from its target operators, whereas other chemicals such as glucose, glucose 6-phosphate, and 6-phosphogluconate did not induce complex dissociation."]

Annotation Implications

  • The strongest core annotation for glk is glucokinase activity coupled to glucose catabolism through the Entner-Doudoroff-centered network.

  • ATP binding and D-glucose binding are mechanistically true but substantially less informative than the specific catalytic term glucokinase activity.

  • For cellular component, cytosol is the more specific GO term for a soluble bacterial cytoplasmic enzyme, while cytoplasm is acceptable but broader.

📄 View Raw YAML

id: Q88P42
gene_symbol: glk
product_type: PROTEIN
status: COMPLETE
aliases:
- PP_1011
- Glucokinase
- Glucose kinase
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida KT2440
description: Glk is the cytosolic glucokinase of Pseudomonas putida KT2440. It
  phosphorylates imported glucose to glucose-6-phosphate using ATP and feeds the
  phosphorylative branch of glucose assimilation into the Entner-Doudoroff-centered
  carbohydrate catabolic network. In KT2440, glk is part of the edd-glk operon and
  mutant and flux analyses indicate that the glucokinase branch is quantitatively
  important for growth on glucose even though P. putida can also oxidize glucose
  through periplasmic gluconate and 2-ketogluconate routes.
existing_annotations:
- term:
    id: GO:0004340
    label: glucokinase activity
  qualifier: enables
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  review:
    summary: This is the core molecular function of glk. UniProt assigns the protein
      as glucokinase (EC 2.7.1.2), and pathway work in KT2440 places Glk at the
      ATP-dependent phosphorylation step that converts cytoplasmic glucose to glucose-6-phosphate.
    action: ACCEPT
    reason: This term is specific, mechanistically correct, and central to the gene's
      function.
    supported_by:
    - reference_id: file:PSEPK/glk/glk-uniprot.txt
      supporting_text: 'RecName: Full=Glucokinase'
    - reference_id: file:PSEPK/glk/glk-uniprot.txt
      supporting_text: 'Reaction=D-glucose + ATP = D-glucose 6-phosphate + ADP + H(+);'
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: In KT2440, glucose imported into the cytoplasm is phosphorylated
        by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
    - reference_id: file:PSEPK/glk/glk-deep-research-falcon.md
      supporting_text: The UniProt target **Q88P42** corresponds to **glk** in *Pseudomonas
        putida* KT2440 (ordered locus **PP_1011**) encoding the cytosolic enzyme
        **glucokinase (Glk)**, which phosphorylates imported glucose to **glucose‑6‑phosphate
        (G6P)** as the entry step of the phosphorylative branch of glucose catabolism.
- term:
    id: GO:0005524
    label: ATP binding
  qualifier: enables
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  review:
    summary: ATP binding is mechanistically true for a kinase, but it is much less
      informative than the specific catalytic term glucokinase activity and is redundant
      for describing the core function of this enzyme.
    action: MARK_AS_OVER_ANNOTATED
    reason: The catalytic activity term already captures the biologically informative
      function.
    supported_by:
    - reference_id: file:PSEPK/glk/glk-uniprot.txt
      supporting_text: /ligand="ATP"
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: ATP binding and D-glucose binding are mechanistically true
        but substantially less informative than the specific catalytic term glucokinase
        activity.
- term:
    id: GO:0005536
    label: D-glucose binding
  qualifier: enables
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  review:
    summary: Substrate recognition is implicit in glucokinase activity, so this term
      is not wrong but is less informative than the specific catalytic annotation.
    action: MARK_AS_OVER_ANNOTATED
    reason: This binding term adds little beyond the core enzymatic activity term.
    supported_by:
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: In KT2440, glucose imported into the cytoplasm is phosphorylated
        by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: ATP binding and D-glucose binding are mechanistically true
        but substantially less informative than the specific catalytic term glucokinase
        activity.
- term:
    id: GO:0005737
    label: cytoplasm
  qualifier: located_in
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  review:
    summary: Glk is a soluble intracellular enzyme and cytoplasmic localization is
      correct, but the more precise term in this context is cytosol.
    action: MODIFY
    reason: Cytoplasm is broader than needed for a soluble bacterial enzyme.
    proposed_replacement_terms:
    - id: GO:0005829
      label: cytosol
    supported_by:
    - reference_id: file:PSEPK/glk/glk-uniprot.txt
      supporting_text: 'SUBCELLULAR LOCATION: Cytoplasm'
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: For cellular component, cytosol is the more specific GO term
        for a soluble bacterial cytoplasmic enzyme, while cytoplasm is acceptable
        but broader.
- term:
    id: GO:0005829
    label: cytosol
  qualifier: located_in
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  review:
    summary: This is the preferred cellular component term for a soluble cytoplasmic
      enzyme such as Glk and is more precise than the broader term cytoplasm.
    action: ACCEPT
    reason: Specific and biologically appropriate localization term.
    supported_by:
    - reference_id: file:PSEPK/glk/glk-uniprot.txt
      supporting_text: 'SUBCELLULAR LOCATION: Cytoplasm'
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: For cellular component, cytosol is the more specific GO term
        for a soluble bacterial cytoplasmic enzyme, while cytoplasm is acceptable
        but broader.
- term:
    id: GO:0006096
    label: glycolytic process
  qualifier: involved_in
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  review:
    summary: Glk participates in glucose catabolism to pyruvate, but in P. putida
      KT2440 that process is routed through the Entner-Doudoroff-centered network
      rather than a generic undifferentiated glycolysis term.
    action: MODIFY
    reason: A more pathway-specific biological process term is available.
    proposed_replacement_terms:
    - id: GO:0061688
      label: glycolytic process via Entner-Doudoroff Pathway
    supported_by:
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: glk is physically linked to the Entner-Doudoroff pathway because
        edd and glk form one operon in KT2440.
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: Because of this genomic organization, the glucokinase pathway
        is co-induced with Entner-Doudoroff genes and is expressed even during growth
        on gluconate or 2-ketogluconate.
    - reference_id: file:PSEPK/glk/glk-deep-research-falcon.md
      supporting_text: 'This physical linkage is consistent with functional coupling:
        the **glucokinase entry step (Glk)** and the subsequent ED processing step
        (Edd) are transcriptionally coordinated.'
- term:
    id: GO:0051156
    label: glucose 6-phosphate metabolic process
  qualifier: involved_in
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  review:
    summary: This term accurately reflects the immediate pathway context of Glk, which
      generates glucose-6-phosphate from glucose and ATP.
    action: ACCEPT
    reason: Directly describes the metabolic process in which the enzymatic reaction
      participates.
    supported_by:
    - reference_id: file:PSEPK/glk/glk-uniprot.txt
      supporting_text: 'Reaction=D-glucose + ATP = D-glucose 6-phosphate + ADP + H(+);'
    - reference_id: file:PSEPK/glk/glk-notes.md
      supporting_text: In KT2440, glucose imported into the cytoplasm is phosphorylated
        by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO
    terms.
  findings:
  - statement: InterPro-based annotation captures the conserved glucokinase family
      assignment and related substrate-level inferences.
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings:
  - statement: TreeGrafter provides a phylogeny-based cellular component inference
      for this conserved bacterial glucokinase family.
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods.
  findings:
  - statement: Automated UniProt and rule-based methods recover the core glucokinase
      activity and broad glucose catabolic process assignments for glk.
- id: PMID:17483213
  title: "Convergent peripheral pathways catalyze initial glucose catabolism in Pseudomonas
    putida: genomic and flux analysis."
  findings:
  - statement: Glucose catabolism in P. putida occurs through three simultaneous pathways
      converging at 6-phosphogluconate.
    supporting_text: glucose catabolism in Pseudomonas putida occurs through the
      simultaneous operation of three pathways that converge at the level of 6-phosphogluconate
  - statement: Glucose imported into the cytoplasm is phosphorylated by glucokinase
      to glucose-6-phosphate.
    supporting_text: Glucose is transported to the cytoplasm in a process mediated
      by an ABC uptake system encoded by open reading frames PP1015 to PP1018 and
      is then phosphorylated by glucokinase (encoded by the glk gene) and converted
      by glucose-6-phosphate dehydrogenase (encoded by the zwf genes) to 6-phosphogluconate.
  - statement: The glucokinase pathway and 2-ketogluconate loop are quantitatively
      more important than direct gluconate phosphorylation.
    supporting_text: although all three functioned simultaneously, the glucokinase
      pathway and the 2-ketogluconate loop were quantitatively more important than
      the direct phosphorylation of gluconate.
  - statement: The glucokinase pathway is required for growth on glucose in KT2440.
    supporting_text: It can therefore be concluded that the glucokinase pathway is
      a sine qua non condition for P. putida to grow with glucose.
- id: PMID:19506074
  title: "Regulation of Glucose Metabolism in Pseudomonas: The Phosphorylative
    Branch and Entner-Doudoroff Enzymes Are Regulated by a Repressor Containing
    a Sugar Isomerase Domain."
  findings:
  - statement: The glucose phosphorylative pathway and Entner-Doudoroff pathway
      are organized into operons that include an edd-glk-gltR2-gltS operon.
    supporting_text: In Pseudomonas putida, genes for the glucose phosphorylative
      pathway and the Entner-Doudoroff pathway are organized in two operons; one
      made up of the zwf, pgl, and eda genes and another consisting of the edd,
      glk, gltR2, and gltS genes.
  - statement: Expression from P(zwf), P(edd), and P(gap) is modulated by HexR
      in response to glucose availability.
    supporting_text: Expression from P(zwf), P(edd), and P(gap) is modulated by
      HexR in response to the availability of glucose in the medium.
  - statement: Binding of KDPG to HexR releases the repressor, whereas glucose,
      glucose 6-phosphate, and 6-phosphogluconate do not.
    supporting_text: Binding of the Entner-Doudoroff pathway intermediate 2-keto-3-deoxy-6-phosphogluconate
      to HexR released the repressor from its target operators, whereas other chemicals
      such as glucose, glucose 6-phosphate, and 6-phosphogluconate did not induce
      complex dissociation.
- id: file:PSEPK/glk/glk-uniprot.txt
  title: UniProt entry Q88P42
  findings:
  - statement: glk corresponds to PP_1011 in Pseudomonas putida KT2440.
  - statement: UniProt annotates the protein as glucokinase EC 2.7.1.2.
  - statement: The catalytic reaction is D-glucose plus ATP to D-glucose 6-phosphate
      plus ADP and H+.
  - statement: UniProt places the protein in the cytoplasm and in the bacterial glucokinase
      family.
- id: file:PSEPK/glk/glk-notes.md
  title: Curator notes for glk in Pseudomonas putida KT2440
  findings:
  - statement: The core role of glk is ATP-dependent phosphorylation of glucose to
      glucose-6-phosphate.
  - statement: glk is physically and transcriptionally linked to the Entner-Doudoroff
      pathway.
  - statement: Generic binding terms are less informative than glucokinase activity
      for this enzyme.
- id: file:PSEPK/glk/glk-deep-research-falcon.md
  title: Deep research report for glk generated with Falcon
  findings:
  - statement: Q88P42 corresponds to glk/PP_1011 and encodes the cytosolic glucokinase
      of Pseudomonas putida KT2440.
  - statement: Falcon research links glk to the phosphorylative branch of glucose
      catabolism and the edd-glk operon.
  - statement: Falcon research summarizes inducible glucokinase activity and mutant
      phenotypes supporting pathway relevance in KT2440.
core_functions:
- description: Glk is a soluble cytosolic glucokinase that uses ATP to phosphorylate
    imported glucose to glucose-6-phosphate, thereby feeding the phosphorylative
    arm of glucose assimilation into the Entner-Doudoroff-centered glycolytic network
    of Pseudomonas putida KT2440.
  molecular_function:
    id: GO:0004340
    label: glucokinase activity
  directly_involved_in:
  - id: GO:0051156
    label: glucose 6-phosphate metabolic process
  - id: GO:0061688
    label: glycolytic process via Entner-Doudoroff Pathway
  locations:
  - id: GO:0005829
    label: cytosol
  supported_by:
  - reference_id: file:PSEPK/glk/glk-uniprot.txt
    supporting_text: 'Reaction=D-glucose + ATP = D-glucose 6-phosphate + ADP + H(+);'
  - reference_id: file:PSEPK/glk/glk-notes.md
    supporting_text: In KT2440, glucose imported into the cytoplasm is phosphorylated
      by glucokinase to glucose-6-phosphate before conversion to 6-phosphogluconate.
  - reference_id: file:PSEPK/glk/glk-notes.md
    supporting_text: glk is physically linked to the Entner-Doudoroff pathway because
      edd and glk form one operon in KT2440.
  - reference_id: file:PSEPK/glk/glk-deep-research-falcon.md
    supporting_text: The UniProt target **Q88P42** corresponds to **glk** in *Pseudomonas
      putida* KT2440 (ordered locus **PP_1011**) encoding the cytosolic enzyme **glucokinase
      (Glk)**, which phosphorylates imported glucose to **glucose‑6‑phosphate (G6P)**
      as the entry step of the phosphorylative branch of glucose catabolism.
proposed_new_terms: []
suggested_questions:
- question: How does flux through the glucokinase branch versus the periplasmic
    gluconate and 2-ketogluconate branches change across different glucose concentrations
    and oxygen or redox conditions in KT2440?
- question: Does Glk have kinetic or regulatory specialization that distinguishes
    it from glucokinases in other pseudomonads that use different balances of oxidative
    and phosphorylative glucose uptake?
- question: How strongly does HexR-mediated control of the edd-glk operon constrain
    mixed-substrate utilization when glucose and gluconate are simultaneously available?
suggested_experiments:
- experiment_type: Enzyme biochemistry
  description: Purify PP_1011 and measure steady-state kinetics for glucose and ATP,
    substrate specificity, and cofactor dependence under physiologically relevant
    ionic conditions.
- experiment_type: Genetic perturbation and flux analysis
  description: Construct a clean glk deletion and complemented strain, then quantify
    growth and intracellular carbon flux on glucose, gluconate, and mixed carbon
    sources using isotopic tracer experiments.
- experiment_type: Operon regulation
  description: Use promoter reporters or RNA-seq in wild type and hexR backgrounds
    to quantify how glucose, gluconate, 2-ketogluconate, and KDPG-related perturbations
    regulate the edd-glk-gltR2-gltS operon.