gapA

UniProt ID: Q88P44
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

Glyceraldehyde-3-phosphate dehydrogenase (GAPDH; gapA, locus PP_1009) of Pseudomonas putida KT2440. It is a homotetrameric, NAD-dependent oxidoreductase that catalyzes the reversible oxidative phosphorylation of D-glyceraldehyde-3-phosphate with inorganic phosphate to 1,3-bisphosphoglycerate, reducing NAD+ to NADH. This reaction is the central energy-conserving step of the lower glycolytic (Embden-Meyerhof-Parnas) segment. In P. putida, which catabolizes glucose primarily through the Entner-Doudoroff pathway, GAPDH acts at the node where the triose phosphate produced by the ED and EMP routes is channeled toward 3-phosphoglycerate, pyruvate and the TCA cycle. The enzyme is cytoplasmic and is a member of the NAD-dependent glyceraldehyde-3-phosphate dehydrogenase (type I, GAPDH-I) family, with a Rossmann-fold NAD(P)-binding domain and a catalytic cysteine active site.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0006006 glucose metabolic process
IEA
GO_REF:0000002
ACCEPT
Summary: GAPDH catalyzes a core step of glucose catabolism (lower glycolysis), converting glyceraldehyde-3-phosphate to 1,3-bisphosphoglycerate. In P. putida this node integrates Entner-Doudoroff-derived triose phosphate with lower EMP flux. The annotation is biologically correct and represents a core function.
Reason: Family/domain-based IEA annotation consistent with the well-established role of GAPDH in glucose metabolism; supported by P. putida pathway literature.
Supporting Evidence:
PMID:18245293
PP_1009 (gap-1/gapA) encodes glyceraldehyde-3-phosphate dehydrogenase acting in KT2440 glucose catabolism, repressed by the glucose-catabolism regulator HexR (4.89-fold derepression in a hexR mutant).
GO:0016620 oxidoreductase activity, acting on the aldehyde or oxo group of donors, NAD or NADP as acceptor
IEA
GO_REF:0000002
MODIFY
Summary: This is the correct parent molecular function for GAPDH, which oxidizes the aldehyde group of glyceraldehyde-3-phosphate using NAD as acceptor. A more specific child term exists (glyceraldehyde-3-phosphate dehydrogenase (NAD+) (phosphorylating) activity, GO:0004365), which better captures the precise reaction catalyzed by this enzyme.
Reason: The term is correct but too general. The protein is a canonical phosphorylating, NAD-dependent GAPDH (TIGR01534 GAPDH-I, PROSITE PS00071, with catalytic Cys at position 154 and bound NAD+), so the specific child term is warranted.
GO:0050661 NADP binding
IEA
GO_REF:0000002
REMOVE
Summary: This annotation derives from the broad NAD(P)-binding Rossmann-fold InterPro signature (IPR006424). However, all cofactor-binding residues modeled in the UniProt record bind NAD+ (CHEBI:57540), and the protein is classified as GAPDH-I (TIGR01534), the NAD-specific form of the family. There is no P. putida-specific evidence of NADP usage; NADP-dependent GAPDH is the distinct GapN/GapC/non-phosphorylating class.
Reason: Over-propagated electronic inference from a generic NAD(P)-binding domain signature. The structural evidence (NAD+-only binding sites) and GAPDH-I family membership argue against NADP binding for this specific protein. NAD binding is already captured separately.
GO:0051287 NAD binding
IEA
GO_REF:0000002
ACCEPT
Summary: GAPDH binds NAD+ as its catalytic cofactor. The UniProt record annotates multiple NAD+-binding residues (positions 12-13, 37, 81, 123, 314) via the Rossmann-fold NAD(P)-binding domain. This is correct and a core feature.
Reason: Strongly supported by domain architecture and conserved NAD+-binding residues; consistent with NAD-dependent GAPDH-I family membership.

Core Functions

NAD-dependent, phosphorylating glyceraldehyde-3-phosphate dehydrogenase catalyzing the reversible oxidation/phosphorylation of glyceraldehyde-3-phosphate to 1,3-bisphosphoglycerate, the energy-conserving step of lower glycolysis.

Supporting Evidence:
  • GO_REF:0000002
    InterPro family Glyceraldehyde-3-P_DH_1 (IPR006424) and GAPDH-I signature (TIGR01534); conserved catalytic Cys (ACT_SITE 154) and NAD+-binding residues in the UniProt record.
  • PMID:18245293
    In P. putida KT2440 glucose catabolism, glyceraldehyde-3-phosphate dehydrogenase (PP_1009/gap-1) acts on the product of glucose metabolism to channel carbon toward Krebs cycle intermediates.

References

Gene Ontology annotation through association of InterPro records with GO terms
A set of activators and repressors control peripheral glucose pathways in Pseudomonas putida to yield a common central intermediate
  • PP_1009 (gap-1/gapA) encodes glyceraldehyde-3-phosphate dehydrogenase acting in peripheral/central glucose catabolism in P. putida KT2440 and is repressed by HexR (4.89-fold derepression in a hexR mutant).

Deep Research

Asta

(gapA-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-06T05:24:25.824530

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.778)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.737)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[3] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.720)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[4] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.720)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[5] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.710)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[6] Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea

  • Authors: A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al.
  • Year: 2023
  • Venue: Frontiers in Microbiology
  • URL: https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  • DOI: 10.3389/fmicb.2023.1066096
  • PMID: 36876067
  • PMCID: 9981795
  • Citations: 2
  • Summary: The present study concludes that deciphering the whole genome of F. udum would be instrumental in understanding evolution, virulence determinants, host-pathogen interaction, possible control strategies, ecological behavior, and many other complexities of the pathogen.
  • Evidence snippets:
  • Snippet 1 (score: 0.707)
    > The BLASTx homology search tool, a component of the standalone NCBI-blast-2.3.0+, was used to perform functional annotation of the F. udum genes (Altschul et al., 1990). With a cut-off E value of ≤1e−06 and a similarity of 34%, BLASTx identified the homologous sequences of the genes in the NCBI non-redundant protein database. Gene ontology (GO) analysis was carried out using Blast2GO PRO 4.1.5 (Conesa and Gotz, 2008). In three different mappings, B2G performed as follows: (1) Using two NCBI-provided mapping files, blast result accessions are used to get gene names (symbols; gene info, gene 2 accessions). (2) Blast result GI identifiers were used to retrieve UniProt IDs using a mapping file from PIR (non-redundant reference protein database), which includes PSD, Swiss-Prot, UniProt, TrEMBL, GenPept, RefSeq, and PDB. The names of the identified genes were searched in the species-specific entries of the gene product table of the GO database. With the aid of the KAAS-KEGG Automatic Annotation Server, pathway analyses were carried out. This database provides functional annotation of genes using other data servers (Moriya et al., 2007). Accessions from the blast results were looked for in the DBXRef table of the GO database.

[7] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.693)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[8] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[9] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.690)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[10] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.690)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[11] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.675)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[12] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[13] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.669)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[14] A novel neural response algorithm for protein function prediction

  • Authors: H. Yalamanchili, Quan-Wu Xiao, Junwen Wang
  • Year: 2012
  • Venue: BMC Systems Biology
  • URL: https://www.semanticscholar.org/paper/0ae3a515fb360b8a4f225d623b23f86a63b5659c
  • DOI: 10.1186/1752-0509-6-S1-S19
  • PMID: 23046521
  • PMCID: 3403322
  • Citations: 7
  • Summary: This work designed a novel automated protein functional assignment method based on the neural response algorithm, which simulates the neuronal behavior of the visual cortex in the human brain and gives it an edge over other available methods on annotation accuracy.
  • Evidence snippets:
  • Snippet 1 (score: 0.669)
    > Recent advances in high-throughput sequencing technologies have enabled the scientific community to sequence a large number of genomes. Currently there are 1,390 complete genomes [1] annotated in the KEGG genome repository and many more are in progress. However, experimental functional characterization of these genes cannot match the data production rate. Adding to this, more than 50% of functional annotations are enigmatic [2]. Even the well studied genomes, such as E. coli and C. elegans, have 51.17% and 87.92% ambiguous annotations (putative, probable and unknown) respectively [2]. To fill the gap between the number of sequences and their (quality) annotations, we need fast, yet accurate automated functional annotation methods. Such computational annotation methods are also critical in analyzing, interpreting and characterizing large complex data sets from high-throughput experimental methods, such as protein-protein interactions (PPI) [3] and gene expression data by clustering similar genes and proteins.
    > The definition of biological function itself is enigmatic in biology and highly context dependent [4][5][6]. This is part of the reason why more than 50% of functional annotations are ambiguous. The functional scope of a protein in an organism differs depending on the aspects under consideration. Proteins can be annotated based on their mode of action, i.e. Enzyme Commission (EC) number [7] (physiological aspect) or their association with a disease (phenotypic aspect). The lack of functional coherence increases the complexity of automated functional annotation. Another major barrier is the use of different vocabulary by different annotations. A function can be described differently in different organisms [8]. This problem can be solved by using ontologies, which serve as universal functional definitions. Enzyme Commission (E.C) [9], MIPS Functional Catalogue (FunCat) [10] and Gene Ontology (GO) [11] are such ontologies. With GO being the most recently and widely used, many automated annotation methods use GO for functional annotation.
    > Protein function assignment methods can be divided into two main categories -structure-based methods and sequence-based methods. A protein's function is highly related to its structure. Protein

[15] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[16] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[17] Text-mining and information-retrieval services for molecular biology

  • Authors: Martin Krallinger, A. Valencia
  • Year: 2005
  • Venue: Genome Biology
  • URL: https://www.semanticscholar.org/paper/558a2745d6e1ac99f77dde88d62566237bd3cfad
  • DOI: 10.1186/gb-2005-6-7-224
  • PMID: 15998455
  • PMCID: 1175978
  • Citations: 237
  • Influential citations: 1
  • Summary: A range of text-mining applications have been developed recently that will improve access to knowledge for biologists and database annotators.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > Biological research is name-centered: proteins are referred to in free text by their names or symbols rather than using the unambiguous identifiers provided by annotation databases (such as SwissProt accession numbers [16]). Identifying mentions of proteins and genes unambiguously within free text is a fundamental step for the later extraction of functional attributes of these entities. Unfortunately this is a difficult process, partly because of the complex nature and usage of gene and protein names. Genes and proteins may be referred to in free text in a range of different ways: as full names (for example, porin), as symbols (the Saccharomyces cerevisiae gene POR1), and also through typographical variants (POR-1). Many genes also have several synonyms (such as OMP2 for POR1), or the gene name may be ambiguous [17] and refer to words that also have a different meanings depending on the context (for example, big brain, the full name for the Drosophila melanogaster gene bib, could also be an anatomical description). Furthermore, it has been suggested that errors in gene names might be introduced automatically by certain applications in bioinformatics [18].
    > In the NLP field, the identification of entities in free text is known as named-entity recognition (NER). To identify biological entities such as genes, proteins and drugs automatically and unambiguously within free text, over 50 information-extraction and text-mining tools have recently been implemented, and two community-wide evaluations have been carried out [19,20]. The top left of Figure 1 shows nine existing NER applications for biology that are provided via an online server or are directly downloadable. Note that the average recovery of biological entities from free text by 15 NER tools was 80%, and the results had an accuracy of 80% [21]; these figures are significantly lower than in the case of entities found in documents from fields such as economics, which demonstrates the complex nature of protein names.
    > Proteins and genes are characterized within biological databases through unique identifiers; each identifier is associated with its corresponding protein or nucleotide sequence and functional descriptions.

[18] Synthesis, characterization, and computational evaluation of some synthesized xanthone derivatives: focus on kinase target network and biomedical properties

  • Authors: Wisam Taher Muslim, L. J. Mohammad, Munaf M. Naji, Isaac Karimi, Matheel D. Al-Sabti et al.
  • Year: 2025
  • Venue: Frontiers in Pharmacology
  • URL: https://www.semanticscholar.org/paper/659ab502877a1d6b5ab7ce45fa51f0f9a13dcf24
  • DOI: 10.3389/fphar.2024.1511627
  • PMID: 39830340
  • PMCID: 11738930
  • Summary: Acute leukemic T-cells were one of the top predicted tumor cell lines for these ligands and the possible antileukemic effects of synthesized xanthone derivatives are potentially very interesting and warrant further studies.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > The UniProt accession identification of target kinases was converted to gene symbols for humans using the SynGO gene set analysis tool (Koopmans et al., 2019), and pooled together, and submitted to GeneMANIA to construct target kinase network. GeneMANIA is a handy web interface for acquiring gene ontology, scrutinizing gene lists, and highlighting genes for functional assays (Warde-Farley et al., 2010). After choosing Homo sapiens from the list of optional organisms, the genes of interest in the previous step were entered into the search bar and the results were collated and high-scored genes were culled for further discussion. Moreover, the protein-protein network was also constructed in STRING ver. 12 launched at https://string-db.org, and submitted to Cytoscape ver. 3.10.2 for network analysis using a novel Cytoscape plugin cytoHubba and visualization (Shannon et al. , 2003).

[19] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.665)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[20] Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser

  • Authors: Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała
  • Year: 2025
  • Venue: BMC Genomics
  • URL: https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da
  • DOI: 10.1186/s12864-025-12329-8
  • PMID: 41408139
  • PMCID: 12822150
  • Citations: 2
  • Summary: An R script is developed that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database and provides a single list of updated symbols with annotation about their ambiguity.
  • Evidence snippets:
  • Snippet 1 (score: 0.665)
    > Gene nomenclature contains current official symbols and various numbers of synonyms, which pose a challenge to integrating genomic data and increase the probability that different genes share the same symbol. Therefore, we retrieved identifiers assigned to all protein-coding genes in human, mouse, rat and pig genomes that are available in the Ensembl genome browser (release 113) to assess the number of genes, compare species and identify ambiguous symbols. Results: Our analysis revealed that the total number of symbols, both official symbols and synonyms, used to identify protein-coding genes ranges from 16,600 in pigs to 64,580 in mice. Furthermore, the gene nomenclature is not complete because there are also genes without an assigned symbol, which indicates gaps in understanding protein-coding genes, especially in pigs. We also found a large number of gene symbols that map to more than one gene. These symbols might complicate the identification of about 10% of rat and mouse genes and 18% of human protein-coding genes. A simple solution for this problem is the usage of stable gene IDs assigned by scientific institutions and committees (Ensembl, NCBI, RGD, HGNC and VGNC) provided that the genomic information associated with these IDs is retrieved directly from proprietary databases containing the most accurate data. Finally, although gene symbols may pose a problem with unequivocal identification of genes, there are instances when no other identifiers are available in the literature. Therefore, we have developed an R script performing search of the Ensembl database and integrating data to provide a single list of updated symbols with annotation about their ambiguity. Conclusions: Gene symbols are not always reliable and should be reported together with stable IDs to enable unequivocal identification of genes. Therefore, data containing only gene symbols should be used cautiously to avoid misidentification of genes. A solution for this problem is our R script REgeness that performs a gene symbol update to current official versions combined with identification of ambiguous symbols and retrieval of other IDs from the Ensembl database.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  3. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  4. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  5. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  6. A. Srivastava, R. Srivastava, Jagriti Yadav, Ashutosh Kumar Singh, P. Tiwari et al. (2023). Virulence and pathogenicity determinants in whole genome sequence of Fusarium udum causing wilt of pigeon pea. Frontiers in Microbiology. https://www.semanticscholar.org/paper/ac4c8e1cd07dbd943c544dab0dff140617956e3a
  7. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  8. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  9. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  10. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  11. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  12. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  13. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  14. H. Yalamanchili, Quan-Wu Xiao, Junwen Wang (2012). A novel neural response algorithm for protein function prediction. BMC Systems Biology. https://www.semanticscholar.org/paper/0ae3a515fb360b8a4f225d623b23f86a63b5659c
  15. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  16. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  17. Martin Krallinger, A. Valencia (2005). Text-mining and information-retrieval services for molecular biology. Genome Biology. https://www.semanticscholar.org/paper/558a2745d6e1ac99f77dde88d62566237bd3cfad
  18. Wisam Taher Muslim, L. J. Mohammad, Munaf M. Naji, Isaac Karimi, Matheel D. Al-Sabti et al. (2025). Synthesis, characterization, and computational evaluation of some synthesized xanthone derivatives: focus on kinase target network and biomedical properties. Frontiers in Pharmacology. https://www.semanticscholar.org/paper/659ab502877a1d6b5ab7ce45fa51f0f9a13dcf24
  19. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  20. Grzegorz R. Juszczak, C. Pareek, U. Czarnik, M. Pierzchała (2025). Protein-coding genes in humans and model mammals (mouse, rat and pig): gene identifiers and disambiguation of gene nomenclature retrieved from the Ensembl genome browser. BMC Genomics. https://www.semanticscholar.org/paper/d9089dfc889d790deb49cbc5b4617bda55bcc4da

Falcon

(gapA-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 17 citations 2 artifacts 2026-06-11T17:23:12.712368

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: gapA / PP_1009 (UniProt Q88P44) in Pseudomonas putida KT2440

0) Target verification (mandatory disambiguation)

The UniProt target Q88P44 is annotated as glyceraldehyde-3-phosphate dehydrogenase (GAPDH) in Pseudomonas putida KT2440, gene gapA, ordered locus PP_1009. In primary KT2440 literature, open reading frame PP1009 is explicitly annotated as gap-1 (gapA-like) encoding glyceraldehyde-3-phosphate dehydrogenase, and it is described as monocistronic. This directly matches the UniProt-provided identity and function and rules out ambiguity with unrelated “gapA” symbols in other organisms. (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5)

A schematic in the same work places this locus within the peripheral glucose catabolic gene clusters under central regulators (HexR/PtxS/GnuR), supporting that PP_1009 is part of the canonical glucose catabolic/regulatory landscape in KT2440. (castillo2008asetof media 9c82fbd7, castillo2008asetof media 584b5e53)

1) Key concepts and current functional definition

1.1 Canonical biochemical role (GAPDH reaction)

GapA (GAPDH) catalyzes the oxidative phosphorylation step of lower glycolysis:

  • D-glyceraldehyde-3-phosphate + phosphate → 3-phospho-D-glyceroyl phosphate (1,3-bisphosphoglycerate) + NADH

This explicitly indicates production of NADH (i.e., NAD+ is the electron acceptor) in the described reaction. (geiger2019investigationofrnabased pages 102-105)

A broader enzymology note reports that GAPDH “is known to act slowly on other aldehydes” and that “phosphates can be replaced by thiols,” consistent with known promiscuity/side reactivities described for GAPDH family enzymes (important when interpreting in vitro assays and stress conditions). (geiger2019investigationofrnabased pages 102-105)

1.2 Substrate/cofactor specificity—what is known vs unknown for KT2440

  • Directly supported for KT2440 PP_1009: NAD-linked activity is supported by the explicit NADH product in the reaction description. (geiger2019investigationofrnabased pages 102-105)
  • Not found in retrieved KT2440 sources: quantitative kinetic constants (Km/kcat), experimentally demonstrated NAD vs NADP preference, and measured oligomeric state for KT2440 GapA were not present in the retrieved full texts.

Because of this, any finer-grained claims (e.g., “strictly NAD-specific” vs “dual NAD(P)”) cannot be asserted here for Q88P44 beyond NADH-forming activity, without consulting dedicated biochemical characterization studies.

1.3 Cellular localization

No KT2440-specific experimental localization evidence (e.g., fractionation) was retrieved. However, gapA/GAPDH in bacteria is generally considered a cytosolic enzyme and, consistent with this, a closely related Pseudomonas GapA is predicted cytoplasmic in a comparative analysis; this is supportive but inferential for KT2440 Q88P44. (sun2025theroleof pages 4-6)

2) Pathway context in P. putida KT2440

2.1 ED/EMP integration and the GAP node

P. putida KT2440 is widely described as using a hybrid central carbon architecture in which the Entner–Doudoroff (ED) pathway and portions of EMP (glycolysis) are integrated and interact with the pentose phosphate network. (lorenzo2024pseudomonasputidakt2440 pages 4-7)

Within the KT2440 glucose catabolic program:
- Glyceraldehyde-3-phosphate (GAP) is described as the end product of the ED pathway and is converted to 3-phosphoglycerate (3-PG) by glyceraldehyde-3-phosphate dehydrogenase, encoded by gapA or an isozyme (PP_3443) acting in the “lower EMP” segment. (chen2024gnurrepressesthe pages 4-6)

This positions GapA at a key junction that couples ED-derived triose phosphate to lower glycolytic outputs and ultimately to TCA-cycle entry.

2.2 Role in channeling carbon to the TCA cycle

In a foundational KT2440 study on peripheral glucose pathways and their regulators, glyceraldehyde-3-phosphate dehydrogenase is described as acting on “the final product of glucose metabolism” and helping “to channel glucose to Krebs cycle intermediates,” consistent with its functional placement at the end of the ED-derived glucose breakdown pipeline. (castillo2008asetof pages 5-6)

3) Regulation and control: what controls gapA/PP_1009 in KT2440?

3.1 HexR repression (experimentally supported; quantitative)

In KT2440, HexR is demonstrated to repress PP_1009 (gap-1/gapA-like): in a hexR mutant, expression of PP1009 (gap-1) increased 4.89-fold (P = 0.01). (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5)

This result is also captured in the study’s Table 3 (image evidence). (castillo2008asetof media 9c82fbd7)

The same study provides broader context that HexR coordinates repression of multiple steps that lead to and through ED metabolism, consistent with coordinated control of glucose catabolism. (castillo2008asetof pages 1-2)

3.2 GnuR and other regulators (2024 update)

A 2024 multi-omics study focused on GnuR reports that gapA is among catabolic genes “similarly induced” by both glucose and gluconate, along with ED and PP pathway genes, placing gapA within the substrate-responsive glucose/gluconate catabolic response. (chen2024gnurrepressesthe pages 4-6)

The same work states that the “primary role of GnuR” is to directly repress expression of catabolic genes functioning in ED and peripheral glucose/gluconate metabolism pathways, providing a contemporary regulatory model in which gapA participates as part of a coordinated regulon responding to glucose/gluconate availability. (chen2024gnurrepressesthe pages 3-4)

While the excerpt indicates that “several” genes in these pathways can increase “almost 100-fold” under glucose/gluconate versus succinate, the provided text does not report a gapA-specific fold-change, so this should be interpreted as regulon-level rather than gapA-specific quantitative evidence. (chen2024gnurrepressesthe pages 3-4)

3.3 Genomic organization and regulatory schematics (visual evidence)

The del Castillo et al. work includes a schematic of peripheral glucose catabolic gene clusters and associated regulators (HexR/PtxS/GnuR), and Table 3 lists PP1009 (gap-1) among genes derepressed in the hexR mutant. These visuals support both genomic-context and regulatory claims. (castillo2008asetof media 9c82fbd7, castillo2008asetof media 584b5e53)

4) Recent developments (prioritizing 2023–2024)

4.1 2024: multi-omics regulatory mapping of glucose/gluconate catabolism

Chen et al. (Microbial Biotechnology, Nov 2024) integrate physiological studies and multi-omics to define a regulatory mode for glucose/gluconate catabolism, emphasizing clustered catabolic genes and transcription factors, and explicitly placing gapA among the induced, central catabolic genes connected to ED/PP/EMP functions. (chen2024gnurrepressesthe pages 4-6)

4.2 2024: systems biology—proteomics and flux interpretation at the GAP node

Mendonca et al. (Environmental Science & Technology, Jun 2024) provide proteomics evidence that P. putida has two GAPDH isozymes (GAPA and GAPB) and report condition-dependent regulation:
- GAPA abundance decreased twofold (p < 0.01) in glucose:ferulate mixed-substrate growth versus glucose-only; GAPB did not significantly change (p > 0.392). (mendonca2024disproportionatecarbondioxide pages 6-8)

They discuss functional partitioning (GAPA for lower glycolysis; GAPB for gluconeogenesis) and interpret that even with reduced GAPA, glycolytic flux downstream of GAP can remain favored, highlighting the GAP node as a key interface for coordinating carbon flux between glycolytic and gluconeogenic regimes under complex substrate mixtures. (mendonca2024disproportionatecarbondioxide pages 6-8)

5) Current applications and real-world implementations (with emphasis on 2024)

5.1 KT2440 as an industrial/synthetic biology chassis—why the gapA node matters

A 2024 Journal of Bacteriology synthesis describes KT2440’s central metabolism (ED/EMP hybrid and triose/hexose recycling) as producing high reducing power (NAD(P)H) relative to ATP, a property linked to stress tolerance and suitability for engineering high-redox-demand bioprocesses; GapA is positioned at the triose-phosphate processing node that couples ED output to downstream metabolism and cofactor generation. (lorenzo2024pseudomonasputidakt2440 pages 4-7, lorenzo2024pseudomonasputidakt2440 pages 2-4)

The same 2024 source highlights real-world application areas enabled by this metabolism, including production of polyhydroxyalkanoates (PHAs) and multiple chemicals (e.g., cis,cis-muconate and others) and use in bioremediation of pollutants, as well as integration with bioelectrochemical hardware for biosensing and electron-assisted biodegradation/biocatalysis/CO2 reduction. (lorenzo2024pseudomonasputidakt2440 pages 4-7, lorenzo2024pseudomonasputidakt2440 pages 2-4)

5.2 2024 bioelectrochemical implementation (explicitly dated)

Within the 2024 glucose/gluconate regulation paper’s cited application landscape, a concrete 2024 implementation is referenced: “Anaerobic glucose uptake in Pseudomonas putida KT2440 in a bioelectrochemical system” (Pause et al., 2024), indicating ongoing development of KT2440 in electro-biotechnology contexts where central carbon metabolism—and therefore GAPDH node capacity—can be limiting or targeted. (chen2024gnurrepressesthe pages 12-13, lorenzo2024pseudomonasputidakt2440 pages 12-12)

6) Key quantitative statistics (selected)

  • HexR repression strength: PP1009/gap-1 expression increases 4.89-fold in a hexR mutant (P = 0.01). (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5)
  • Growth/uptake robustness in regulator mutants: regulator mutants (including hexR) show similar growth rates (~0.57–0.60 h−1) and glucose consumption rates (~6.84–9.7 mmol glucose·g cell biomass−1·h−1) to parental strain under tested conditions. (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5)
  • Proteomics regulation at the GAPDH node (2024): GAPA decreases twofold (p < 0.01) in glucose:ferulate vs glucose-only; GAPB unchanged (p > 0.392). (mendonca2024disproportionatecarbondioxide pages 6-8)
  • Regulon-scale induction (2024): multiple glucose catabolism genes can increase “almost 100-fold” under glucose/gluconate vs succinate (note: not gapA-specific in excerpt). (chen2024gnurrepressesthe pages 3-4)

7) Evidence summary table

The following table consolidates key evidence items (identity, reaction, pathway placement, regulation, and quantitative results) with URLs and dates where available:

Aspect Key finding Evidence source (first author year) Publication date URL Citation id(s)
identity/locus The target in Pseudomonas putida KT2440 is PP1009 (PP_1009), annotated as gap-1 / gapA-like, encoding glyceraldehyde-3-phosphate dehydrogenase; it is reported as monocistronic in the PP1009–PP1024 chromosomal region. del Castillo 2008 Apr 2008 https://doi.org/10.1128/jb.01726-07 (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5)
reaction GapA/GAPDH catalyzes oxidation/phosphorylation of D-glyceraldehyde-3-phosphate with inorganic phosphate to 3-phospho-D-glyceroyl phosphate (1,3-bisphosphoglycerate), producing NADH; one source notes the enzyme can act slowly on other aldehydes and thiols can substitute for phosphate. Geiger 2019 2019 not available (geiger2019investigationofrnabased pages 102-105)
cofactor specificity Available evidence supports NAD-linked GAPDH activity through explicit NADH formation in the reaction description. No KT2440-specific experimental evidence for NADP preference was retrieved in the gathered sources, so cofactor specificity beyond NAD-linked activity remains unresolved here. Geiger 2019 2019 not available (geiger2019investigationofrnabased pages 102-105)
pathway role In KT2440 glucose catabolism, GAP is the end product of the Entner–Doudoroff (ED) pathway and is converted to 3-phosphoglycerate (3-PG) by GapA or the isozyme PP_3443 in the lower EMP pathway. HexR-regulated glucose pathways funnel carbon to the central intermediate 6-phosphogluconate and onward to GAP/pyruvate. Chen 2024; del Castillo 2008 Nov 2024; Apr 2008 https://doi.org/10.1111/1751-7915.70059 ; https://doi.org/10.1128/jb.01726-07 (chen2024gnurrepressesthe pages 3-4, chen2024gnurrepressesthe pages 4-6, castillo2008asetof pages 1-2)
regulation HexR represses gapA/gap-1: in a hexR mutant, PP1009/gap-1 expression increased 4.89-fold (P = 0.01). More recent work places gapA among catabolic genes induced by glucose and gluconate; several glucose-catabolism genes increased almost 100-fold under these conditions, although a gapA-specific fold-change was not given in the excerpt. del Castillo 2008; Chen 2024 Apr 2008; Nov 2024 https://doi.org/10.1128/jb.01726-07 ; https://doi.org/10.1111/1751-7915.70059 (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5, chen2024gnurrepressesthe pages 3-4, chen2024gnurrepressesthe pages 4-6, castillo2008asetof media 9c82fbd7)
quantitative data In the hexR mutant, overall physiology remained near wild type under tested conditions despite transcriptional derepression: growth rates were about 0.57 ± 0.01 to 0.60 ± 0.02 h−1, and glucose consumption rates about 6.84 to 9.7 mmol glucose g cell biomass−1 h−1. del Castillo 2008 Apr 2008 https://doi.org/10.1128/jb.01726-07 (castillo2008asetof pages 5-6, castillo2008asetof pages 4-5)
isozyme context KT2440 has at least two GAPDH isozymes, GAPA and GAPB. Proteomics in 2024 showed GAPA decreased twofold (p < 0.01) in glucose+ferulate versus glucose alone, while GAPB was unchanged (p > 0.392). The authors interpret prior work as functional partitioning in which GAPA supports lower glycolysis and GAPB supports gluconeogenesis. Mendonca 2024 Jun 2024 https://doi.org/10.1021/acs.est.4c01328 (mendonca2024disproportionatecarbondioxide pages 6-8)
pathway/physiology context During glucose or gluconate feeding, gapA grouped with ED/PP/EMP catabolic genes induced by both substrates. During mixed-substrate growth with ferulate, ED pathway usage remained prominent and glycolytic flux downstream of GAP was favored despite lower GAPA abundance. Chen 2024; Mendonca 2024 Nov 2024; Jun 2024 https://doi.org/10.1111/1751-7915.70059 ; https://doi.org/10.1021/acs.est.4c01328 (chen2024gnurrepressesthe pages 4-6, mendonca2024disproportionatecarbondioxide pages 6-8)

Table: This table compiles evidence-based findings for Pseudomonas putida KT2440 gapA/PP_1009, including identity, enzymatic function, regulation, pathway placement, and quantitative observations. It highlights what is directly supported by retrieved sources and where evidence remains inferential or incomplete.

8) Evidence gaps and interpretation notes (important for functional annotation)

  1. KT2440 GapA kinetics and NADP usage: No retrieved KT2440 primary biochemical characterization (Km/kcat, NAD vs NADP preference) was available in the accessed corpus; current statements are therefore limited to the explicit NADH-forming reaction description. (geiger2019investigationofrnabased pages 102-105)
  2. Localization: No direct KT2440 localization experiment was retrieved; cytosolic localization is a reasonable inference for bacterial GAPDH and is supported by predictions in closely related Pseudomonas GapA analyses, but remains inferential for Q88P44 without KT2440-specific experiments. (sun2025theroleof pages 4-6)
  3. Isozyme mapping to PP_1009: The 2024 proteomics/flux paper discusses GAPA/GAPB isozymes; the gathered excerpts do not explicitly map “GAPA” to the locus tag PP_1009, so isozyme nomenclature should be reconciled against genome annotations when producing a definitive functional annotation. (mendonca2024disproportionatecarbondioxide pages 6-8)

9) Annotated primary function statement (for database-style functional annotation)

gapA (PP_1009; UniProt Q88P44) encodes a glyceraldehyde-3-phosphate dehydrogenase that catalyzes the NAD-linked oxidative phosphorylation of glyceraldehyde-3-phosphate to 1,3-bisphosphoglycerate (NADH-producing), operating at the ED-to-lower-EMP interface in KT2440 glucose metabolism and contributing to routing ED-derived triose-phosphate toward downstream metabolism/TCA-cycle intermediates; its expression is negatively regulated by HexR and is induced as part of glucose/gluconate catabolic programs that include ED/PP/EMP genes. (geiger2019investigationofrnabased pages 102-105, chen2024gnurrepressesthe pages 4-6, castillo2008asetof pages 5-6)

References

  1. (castillo2008asetof pages 5-6): Teresa del Castillo, Estrella Duque, and Juan L. Ramos. A set of activators and repressors control peripheral glucose pathways in pseudomonas putida to yield a common central intermediate. Journal of Bacteriology, 190:2331-2339, Apr 2008. URL: https://doi.org/10.1128/jb.01726-07, doi:10.1128/jb.01726-07. This article has 130 citations and is from a peer-reviewed journal.

  2. (castillo2008asetof pages 4-5): Teresa del Castillo, Estrella Duque, and Juan L. Ramos. A set of activators and repressors control peripheral glucose pathways in pseudomonas putida to yield a common central intermediate. Journal of Bacteriology, 190:2331-2339, Apr 2008. URL: https://doi.org/10.1128/jb.01726-07, doi:10.1128/jb.01726-07. This article has 130 citations and is from a peer-reviewed journal.

  3. (castillo2008asetof media 9c82fbd7): Teresa del Castillo, Estrella Duque, and Juan L. Ramos. A set of activators and repressors control peripheral glucose pathways in pseudomonas putida to yield a common central intermediate. Journal of Bacteriology, 190:2331-2339, Apr 2008. URL: https://doi.org/10.1128/jb.01726-07, doi:10.1128/jb.01726-07. This article has 130 citations and is from a peer-reviewed journal.

  4. (castillo2008asetof media 584b5e53): Teresa del Castillo, Estrella Duque, and Juan L. Ramos. A set of activators and repressors control peripheral glucose pathways in pseudomonas putida to yield a common central intermediate. Journal of Bacteriology, 190:2331-2339, Apr 2008. URL: https://doi.org/10.1128/jb.01726-07, doi:10.1128/jb.01726-07. This article has 130 citations and is from a peer-reviewed journal.

  5. (geiger2019investigationofrnabased pages 102-105): S Geiger. Investigation of rna-based regulation of gene expression in proteobacterial energy metabolism. Unknown journal, 2019.

  6. (sun2025theroleof pages 4-6): Lei Sun, Dao-Jiao Tang, Qian-Nan Zhang, Lu-Lu Li, Lei Zhang, Xin-Yi Zan, Feng-Jie Cui, Ling Sun, and Wen-Jing Sun. The role of glyceraldehyde-3-phosphate dehydrogenase in 2-ketogluconic acid industrial production strain pseudomonas plecoglossicida juim01. Foods, 14:3830, Nov 2025. URL: https://doi.org/10.3390/foods14223830, doi:10.3390/foods14223830. This article has 0 citations.

  7. (lorenzo2024pseudomonasputidakt2440 pages 4-7): Victor de Lorenzo, Danilo Pérez-Pantoja, and Pablo I. Nikel. pseudomonas putida kt2440: the long journey of a soil-dweller to become a synthetic biology chassis. Journal of Bacteriology, Jul 2024. URL: https://doi.org/10.1128/jb.00136-24, doi:10.1128/jb.00136-24. This article has 78 citations and is from a peer-reviewed journal.

  8. (chen2024gnurrepressesthe pages 4-6): Wenbo Chen, Rao Ma, Yong Feng, Yunzhu Xiao, Agnieszka Sekowska, Antoine Danchin, and Conghui You. Gnur represses the expression of glucose and gluconate catabolism in pseudomonas putida kt2440. Microbial Biotechnology, Nov 2024. URL: https://doi.org/10.1111/1751-7915.70059, doi:10.1111/1751-7915.70059. This article has 2 citations and is from a peer-reviewed journal.

  9. (castillo2008asetof pages 1-2): Teresa del Castillo, Estrella Duque, and Juan L. Ramos. A set of activators and repressors control peripheral glucose pathways in pseudomonas putida to yield a common central intermediate. Journal of Bacteriology, 190:2331-2339, Apr 2008. URL: https://doi.org/10.1128/jb.01726-07, doi:10.1128/jb.01726-07. This article has 130 citations and is from a peer-reviewed journal.

  10. (chen2024gnurrepressesthe pages 3-4): Wenbo Chen, Rao Ma, Yong Feng, Yunzhu Xiao, Agnieszka Sekowska, Antoine Danchin, and Conghui You. Gnur represses the expression of glucose and gluconate catabolism in pseudomonas putida kt2440. Microbial Biotechnology, Nov 2024. URL: https://doi.org/10.1111/1751-7915.70059, doi:10.1111/1751-7915.70059. This article has 2 citations and is from a peer-reviewed journal.

  11. (mendonca2024disproportionatecarbondioxide pages 6-8): Caroll M. Mendonca, Lichun Zhang, Jacob R. Waldbauer, and Ludmilla Aristilde. Disproportionate carbon dioxide efflux in bacterial metabolic pathways for different organic substrates leads to variable contribution to carbon-use efficiency. Environmental Science & Technology, 58:11041-11052, Jun 2024. URL: https://doi.org/10.1021/acs.est.4c01328, doi:10.1021/acs.est.4c01328. This article has 17 citations and is from a domain leading peer-reviewed journal.

  12. (lorenzo2024pseudomonasputidakt2440 pages 2-4): Victor de Lorenzo, Danilo Pérez-Pantoja, and Pablo I. Nikel. pseudomonas putida kt2440: the long journey of a soil-dweller to become a synthetic biology chassis. Journal of Bacteriology, Jul 2024. URL: https://doi.org/10.1128/jb.00136-24, doi:10.1128/jb.00136-24. This article has 78 citations and is from a peer-reviewed journal.

  13. (chen2024gnurrepressesthe pages 12-13): Wenbo Chen, Rao Ma, Yong Feng, Yunzhu Xiao, Agnieszka Sekowska, Antoine Danchin, and Conghui You. Gnur represses the expression of glucose and gluconate catabolism in pseudomonas putida kt2440. Microbial Biotechnology, Nov 2024. URL: https://doi.org/10.1111/1751-7915.70059, doi:10.1111/1751-7915.70059. This article has 2 citations and is from a peer-reviewed journal.

  14. (lorenzo2024pseudomonasputidakt2440 pages 12-12): Victor de Lorenzo, Danilo Pérez-Pantoja, and Pablo I. Nikel. pseudomonas putida kt2440: the long journey of a soil-dweller to become a synthetic biology chassis. Journal of Bacteriology, Jul 2024. URL: https://doi.org/10.1128/jb.00136-24, doi:10.1128/jb.00136-24. This article has 78 citations and is from a peer-reviewed journal.

Artifacts

Citations

  1. geiger2019investigationofrnabased pages 102-105
  2. sun2025theroleof pages 4-6
  3. chen2024gnurrepressesthe pages 4-6
  4. castillo2008asetof pages 5-6
  5. castillo2008asetof pages 1-2
  6. chen2024gnurrepressesthe pages 3-4
  7. mendonca2024disproportionatecarbondioxide pages 6-8
  8. castillo2008asetof pages 4-5
  9. chen2024gnurrepressesthe pages 12-13
  10. https://doi.org/10.1128/jb.01726-07
  11. https://doi.org/10.1111/1751-7915.70059
  12. https://doi.org/10.1021/acs.est.4c01328
  13. https://doi.org/10.1128/jb.01726-07,
  14. https://doi.org/10.3390/foods14223830,
  15. https://doi.org/10.1128/jb.00136-24,
  16. https://doi.org/10.1111/1751-7915.70059,
  17. https://doi.org/10.1021/acs.est.4c01328,

📄 View Raw YAML

id: Q88P44
gene_symbol: gapA
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: Glyceraldehyde-3-phosphate dehydrogenase (GAPDH; gapA, locus PP_1009) of Pseudomonas putida KT2440. It is a homotetrameric, NAD-dependent oxidoreductase that catalyzes the reversible oxidative phosphorylation of D-glyceraldehyde-3-phosphate with inorganic phosphate to 1,3-bisphosphoglycerate, reducing NAD+ to NADH. This reaction is the central energy-conserving step of the lower glycolytic (Embden-Meyerhof-Parnas) segment. In P. putida, which catabolizes glucose primarily through the Entner-Doudoroff pathway, GAPDH acts at the node where the triose phosphate produced by the ED and EMP routes is channeled toward 3-phosphoglycerate, pyruvate and the TCA cycle. The enzyme is cytoplasmic and is a member of the NAD-dependent glyceraldehyde-3-phosphate dehydrogenase (type I, GAPDH-I) family, with a Rossmann-fold NAD(P)-binding domain and a catalytic cysteine active site.
existing_annotations:
- term:
    id: GO:0006006
    label: glucose metabolic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: involved_in
  review:
    summary: GAPDH catalyzes a core step of glucose catabolism (lower glycolysis), converting glyceraldehyde-3-phosphate to 1,3-bisphosphoglycerate. In P. putida this node integrates Entner-Doudoroff-derived triose phosphate with lower EMP flux. The annotation is biologically correct and represents a core function.
    action: ACCEPT
    reason: Family/domain-based IEA annotation consistent with the well-established role of GAPDH in glucose metabolism; supported by P. putida pathway literature.
    supported_by:
    - reference_id: PMID:18245293
      full_text_unavailable: true
      supporting_text: PP_1009 (gap-1/gapA) encodes glyceraldehyde-3-phosphate dehydrogenase acting in KT2440 glucose catabolism, repressed by the glucose-catabolism regulator HexR (4.89-fold derepression in a hexR mutant).
- term:
    id: GO:0016620
    label: oxidoreductase activity, acting on the aldehyde or oxo group of donors, NAD or NADP as acceptor
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: This is the correct parent molecular function for GAPDH, which oxidizes the aldehyde group of glyceraldehyde-3-phosphate using NAD as acceptor. A more specific child term exists (glyceraldehyde-3-phosphate dehydrogenase (NAD+) (phosphorylating) activity, GO:0004365), which better captures the precise reaction catalyzed by this enzyme.
    action: MODIFY
    reason: The term is correct but too general. The protein is a canonical phosphorylating, NAD-dependent GAPDH (TIGR01534 GAPDH-I, PROSITE PS00071, with catalytic Cys at position 154 and bound NAD+), so the specific child term is warranted.
    proposed_replacement_terms:
    - id: GO:0004365
      label: glyceraldehyde-3-phosphate dehydrogenase (NAD+) (phosphorylating) activity
- term:
    id: GO:0050661
    label: NADP binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: This annotation derives from the broad NAD(P)-binding Rossmann-fold InterPro signature (IPR006424). However, all cofactor-binding residues modeled in the UniProt record bind NAD+ (CHEBI:57540), and the protein is classified as GAPDH-I (TIGR01534), the NAD-specific form of the family. There is no P. putida-specific evidence of NADP usage; NADP-dependent GAPDH is the distinct GapN/GapC/non-phosphorylating class.
    action: REMOVE
    reason: Over-propagated electronic inference from a generic NAD(P)-binding domain signature. The structural evidence (NAD+-only binding sites) and GAPDH-I family membership argue against NADP binding for this specific protein. NAD binding is already captured separately.
- term:
    id: GO:0051287
    label: NAD binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: GAPDH binds NAD+ as its catalytic cofactor. The UniProt record annotates multiple NAD+-binding residues (positions 12-13, 37, 81, 123, 314) via the Rossmann-fold NAD(P)-binding domain. This is correct and a core feature.
    action: ACCEPT
    reason: Strongly supported by domain architecture and conserved NAD+-binding residues; consistent with NAD-dependent GAPDH-I family membership.
core_functions:
- description: NAD-dependent, phosphorylating glyceraldehyde-3-phosphate dehydrogenase catalyzing the reversible oxidation/phosphorylation of glyceraldehyde-3-phosphate to 1,3-bisphosphoglycerate, the energy-conserving step of lower glycolysis.
  molecular_function:
    id: GO:0004365
    label: glyceraldehyde-3-phosphate dehydrogenase (NAD+) (phosphorylating) activity
  supported_by:
  - reference_id: GO_REF:0000002
    supporting_text: InterPro family Glyceraldehyde-3-P_DH_1 (IPR006424) and GAPDH-I signature (TIGR01534); conserved catalytic Cys (ACT_SITE 154) and NAD+-binding residues in the UniProt record.
  - reference_id: PMID:18245293
    full_text_unavailable: true
    supporting_text: In P. putida KT2440 glucose catabolism, glyceraldehyde-3-phosphate dehydrogenase (PP_1009/gap-1) acts on the product of glucose metabolism to channel carbon toward Krebs cycle intermediates.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: PMID:18245293
  title: A set of activators and repressors control peripheral glucose pathways in Pseudomonas putida to yield a common central intermediate
  findings:
  - statement: PP_1009 (gap-1/gapA) encodes glyceraldehyde-3-phosphate dehydrogenase acting in peripheral/central glucose catabolism in P. putida KT2440 and is repressed by HexR (4.89-fold derepression in a hexR mutant).
    reference_section_type: RESULTS
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: del Castillo, Duque & Ramos, J Bacteriol 2008;190:2331-2339 (doi:10.1128/jb.01726-07). Corrected identifier from a hallucinated PMID:18156256 (which resolves to an unrelated lfnA/wbuX O-antigen paper, doi:10.1128/JB.01708-07) to the correct PMID:18245293, confirmed via DOI 10.1128/jb.01726-07 against PubMed (abstract states HexR controls gap-1 encoding glyceraldehyde-3-phosphate dehydrogenase). Establishes PP_1009 identity as gapA-like GAPDH and its regulation in KT2440 glucose metabolism.