aceE

UniProt ID: Q88QZ5
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
Aliases:
PP_0339
📝 Provide Detailed Feedback

Gene Description

aceE (PP_0339) encodes the E1 component of pyruvate dehydrogenase, a thiamine-diphosphate enzyme that decarboxylates pyruvate and transfers the resulting hydroxyethyl/acetyl equivalent to the lipoyl group of the E2 component. Together with AceF and lipoamide dehydrogenase, it supports oxidative decarboxylation of pyruvate to acetyl-CoA in central carbon metabolism.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0004739 pyruvate dehydrogenase (acetyl-transferring) activity
IEA
GO_REF:0000120
ACCEPT
Summary: pyruvate dehydrogenase (acetyl-transferring) activity is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
Reason: This is a specific, biologically appropriate annotation for this gene product.
GO:0016491 oxidoreductase activity
IEA
GO_REF:0000002
KEEP AS NON CORE
Summary: oxidoreductase activity is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
Reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.

Core Functions

pyruvate dehydrogenase (acetyl-transferring) activity supporting the Pyruvate dehydrogenase E1 component (EC 1.2.4.1) role summarized for aceE.

Supporting Evidence:
  • file:PSEPK/aceE/aceE-uniprot.txt
    DR GO; GO:0004739; F:pyruvate dehydrogenase (acetyl-transferring) activity; IEA:UniProtKB-EC.

References

Gene Ontology annotation through association of InterPro records with GO terms
Combined Automated Annotation using Multiple IEA Methods
file:PSEPK/aceE/aceE-uniprot.txt
UniProt record for aceE (Q88QZ5)
  • UniProt identifies aceE as Pyruvate dehydrogenase E1 component (EC 1.2.4.1) and provides the seeded EC/domain/GO evidence reviewed here.
file:PSEPK/aceE/aceE-deep-research-asta.md
Asta deep-research retrieval for aceE
  • Asta retrieval was run for this first-pass pathway curation; direct organism-specific literature was limited for several common enzyme names, so UniProt/family evidence carries the main review weight.

Deep Research

Asta

(aceE-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-06T05:23:59.968115

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.715)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.708)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[3] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.702)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[4] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.694)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[5] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.684)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[6] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.678)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[7] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.676)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[8] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.672)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[9] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.670)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[10] The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize

  • Authors: Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al.
  • Year: 2016
  • Venue: BMC Systems Biology
  • URL: https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  • DOI: 10.1186/s12918-016-0369-x
  • PMID: 27899149
  • PMCID: 5129634
  • Citations: 11
  • Influential citations: 1
  • Summary: CornCyc’s computational predictions are more accurate than those in MaizeCyc when compared to experimentally determined function assignments, demonstrating the relative strength of the enzymatic function assignment pipeline used to generate CornCyc.
  • Evidence snippets:
  • Snippet 1 (score: 0.657)
    > A gold standard set of protein functional annotations was generated by extracting data from UniProt [16] and BRENDA [17]. We extracted all protein sequence and annotation data from UniProt (release 2016_05) for the organism Zea mays, keeping the EC annotations only from the manually reviewed component of UniProt, while removing those annotations that had not undergone manual review. We also extracted experimentally verified protein annotations for Zea mays from BRENDA (release 2016.1). The UniProt and BRENDA annotations were then merged by matching proteins based on the database crosslinks provided by BRENDA, resulting in the union of the reviewed annotations from UniProt and the experimentally verified annotations of BRENDA with duplicates removed. The merged protein annotations were then matched to the B73 RefGen_v2 translated gene models using BLASTP based on a sequence identity cutoff of 96% and an e-value cutoff of 1e-20. We selected the top scoring hit for each protein which resulted in matches to 1,815 unique maize proteins. EC annotations for alternate isoforms were consolidated at the gene level, resulting in 1,475 experimentally verified or manually reviewed protein functional annotations across 1,450 maize genes.

[11] MitoMiner, an Integrated Database for the Storage and Analysis of Mitochondrial Proteomics Data

  • Authors: Anthony C. Smith, A. Robinson
  • Year: 2009
  • Venue: Molecular & Cellular Proteomics : MCP
  • URL: https://www.semanticscholar.org/paper/206a60d6d387688ce7f880e0dfe67af5a723c7b8
  • DOI: 10.1074/mcp.M800373-MCP200
  • PMID: 19208617
  • PMCID: 2690483
  • Citations: 84
  • Influential citations: 7
  • Summary: Analysis indicated that enzymes of some cytosolic metabolic pathways are regularly detected in mitochondrial proteomics experiments, suggesting that they are associated with the outside of the outer mitochondrial membrane.
  • Evidence snippets:
  • Snippet 1 (score: 0.655)
    > Recorded for each protein of the mass spectrometry data sets were, where available, the original protein identifier, subcellular location, sequence of identified peptides, sequence coverage, and the experimental techniques that had been used for the purification, separation, and identification of the protein. If the original protein identifier could not be mapped to a UniProt primary accession number by PIR ID or MGI, then the protein was compared with proteins in UniProt by using BLASTP (14). If there was a significant match, then the UniProt primary accession number was assigned to the protein. Those proteins without a significant match were discarded. By using the PIR ID and the MGI identifier conversion tools, the evidence of mitochondrial localization for a protein was linked to many of the UniProt entries representing it. Identifiers of proteins encoded in the mitochondrial genome of organisms were taken from the Organelle database of the European Molecular Biology Laboratory-European Bioinformatics Institute and used to annotate the appropriate proteins in MitoMiner.
    > The source of protein sequences, related features, and annotation was UniProt (11). All UniProt entries were downloaded for the six species with mitochondrial localization data sets. The literature citations in each UniProt entry were retrieved from PubMed by using an InterMine parser. Additional Gene Ontology annotation on the biological process, metabolic function, and cellular component of each protein was taken from UniProt (15) and individual genome projects of M. musculus (12), Rattus norvegicus (16), Drosophila melanogaster (17), and Saccharomyces cerevisiae (18).
    > Finally lists of human genes and the descriptions of their associated disease phenotypes were taken from OMIM (19), the definitions of groups of homologous proteins were taken from HomoloGene (20), and data on the reactions, enzymes, and compounds of metabolic pathways were taken from KEGG (21). The EC numbers of proteins in UniProt were used to define the cross-reference between proteins and metabolic pathways.

[12] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.640)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[13] A Manual Curation Strategy to Improve Genome Annotation: Application to a Set of Haloarchael Genomes

  • Authors: F. Pfeiffer, D. Oesterhelt
  • Year: 2015
  • Venue: Life
  • URL: https://www.semanticscholar.org/paper/f5983d01e0ac838554f7f5c29481d70a9d728c30
  • DOI: 10.3390/life5021427
  • PMID: 26042526
  • PMCID: 4500146
  • Citations: 38
  • Influential citations: 1
  • Summary: A manual curation effort is described that attempts to produce high-quality genome annotations for a set of haloarchaeal genomes (Halobacterium salinarum and Hbt. hubeiense, Haloferax volcanii and Hfx. mediterranei).
  • Evidence snippets:
  • Snippet 1 (score: 0.640)
    > Labelling of such a gene as "inactivated" seems biologically correct. This is translated to the CDS qualifier /pseudo in EMBL and securely ensures that the protein translation is not present in UniProt (e.g., searching for OE_1059R results in no hit). When, however, an invalid partial translation product is produced but not tagged as disrupted (as is the case for VNG0034H), then this is considered by EMBL as a "regular" gene (CDS). Such a gene fragment is included as a regular protein in UniProt (VNG0034H is Q9HSX6). Upon superficial analysis, this may be taken as evidence for an "improved" (because less incomplete) genome annotation in strain NRC-1 compared to strain R1. In addition, according to EMBL requirements, the "CDS" coordinates of OE_1059R must be given as 29913-31570, thus covering and including the integrated transposon ISH1 (with its transposase gene). Only a "tolerated" misc_feature annotation allows representation of this disrupted gene in a biologically meaningful way, representing the reconstructed ancestral gene.

[14] Text-mining and information-retrieval services for molecular biology

  • Authors: Martin Krallinger, A. Valencia
  • Year: 2005
  • Venue: Genome Biology
  • URL: https://www.semanticscholar.org/paper/558a2745d6e1ac99f77dde88d62566237bd3cfad
  • DOI: 10.1186/gb-2005-6-7-224
  • PMID: 15998455
  • PMCID: 1175978
  • Citations: 237
  • Influential citations: 1
  • Summary: A range of text-mining applications have been developed recently that will improve access to knowledge for biologists and database annotators.
  • Evidence snippets:
  • Snippet 1 (score: 0.635)
    > Biological research is name-centered: proteins are referred to in free text by their names or symbols rather than using the unambiguous identifiers provided by annotation databases (such as SwissProt accession numbers [16]). Identifying mentions of proteins and genes unambiguously within free text is a fundamental step for the later extraction of functional attributes of these entities. Unfortunately this is a difficult process, partly because of the complex nature and usage of gene and protein names. Genes and proteins may be referred to in free text in a range of different ways: as full names (for example, porin), as symbols (the Saccharomyces cerevisiae gene POR1), and also through typographical variants (POR-1). Many genes also have several synonyms (such as OMP2 for POR1), or the gene name may be ambiguous [17] and refer to words that also have a different meanings depending on the context (for example, big brain, the full name for the Drosophila melanogaster gene bib, could also be an anatomical description). Furthermore, it has been suggested that errors in gene names might be introduced automatically by certain applications in bioinformatics [18].
    > In the NLP field, the identification of entities in free text is known as named-entity recognition (NER). To identify biological entities such as genes, proteins and drugs automatically and unambiguously within free text, over 50 information-extraction and text-mining tools have recently been implemented, and two community-wide evaluations have been carried out [19,20]. The top left of Figure 1 shows nine existing NER applications for biology that are provided via an online server or are directly downloadable. Note that the average recovery of biological entities from free text by 15 NER tools was 80%, and the results had an accuracy of 80% [21]; these figures are significantly lower than in the case of entities found in documents from fields such as economics, which demonstrates the complex nature of protein names.
    > Proteins and genes are characterized within biological databases through unique identifiers; each identifier is associated with its corresponding protein or nucleotide sequence and functional descriptions.

[15] RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications

  • Authors: Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al.
  • Year: 2025
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  • DOI: 10.1093/nar/gkaf1206
  • PMID: 41206768
  • PMCID: 12807602
  • Citations: 2
  • Summary: The new RM2Target v2.0 will serve as a foundational resource for exploring RNA epitranscriptomic regulation, enabling investigations into cross-talk among modifications, underlying molecular mechanisms, and disease connections, thereby facilitating both basic research and translational applications in RNA epigenetics.
  • Evidence snippets:
  • Snippet 1 (score: 0.634)
    > To obtain basic information on WERs and their target genes, such as official gene symbols, gene IDs, gene types, and genomic locations, gene annotations were downloaded from the GENCODE project [ 44 ] for human and mouse, and from NCBI [ 45 ] and Ensembl [ 46 ] for the other species. Genomic locations were extracted from the corresponding GTF annotation files. Gene symbols were primarily standardized based on the NCBI Gene database [ 45 ] for mRNAs and lncRNAs, GtR-NAdb [ 47 ] for tRNAs, miRbase [ 48 ] for microRNAs, and cir-cBase [ 49 ] for circRNAs. Deprecated or substituted versions of genes were filtered out. The LiftOver [ 50 ] program was employed to convert and unify genomic coordinates across different genome assembly versions.
    > The functional descriptions of WERs were compiled based on the UniProt database [ 51 ] and further supplemented with evidence from relevant publications, with particular emphasis on their functions as RNA modification regulatory proteins.

[16] Prioritising genetic findings for drug target identification and validation.

  • Authors: N. Hukerikar, A. Hingorani, F. Asselbergs, C. Finan, A. Schmidt
  • Year: 2024
  • Venue: Atherosclerosis
  • URL: https://www.semanticscholar.org/paper/80ee965ca8d81196a8281ab055ff7ff79eda31d9
  • DOI: 10.1016/j.atherosclerosis.2024.117462
  • PMID: 38325120
  • Citations: 11
  • Summary: The current review provides an overview of genetic evidence for drug target identification, and how biomedical databases can be used to provide actionable prioritisation, fully informing downstream experimental validation.
  • Evidence snippets:
  • Snippet 1 (score: 0.633)
    > Other gene identification systems include the Entrez Gene [43] database for gene-specific information, and the HUGO Gene Nomenclature Committee (HGNC) [44] which maintains unique symbols and names for human loci.An analogue for proteins is the UniProt Knowledgebase (UniProtKB) [45], which contains data on protein sequences and function, and each protein in the database is assigned a unique UniProt accession ID.UniProt provides functionality to map between different identifiers, including Ensembl IDs and UniProt accession IDs.
    > A common naming convention is also required to identify the diseases associated with the drug targets.Medical Subject Headings (MeSH) [46] are terms defined by the National Library of Medicine, and act as a standardised thesaurus for diseases and medical conditions which can be used to index PubMed.In some data sources, such as the Chemical Biology Database (ChEMBL) [23], diseases and outcomes will be identified by MeSH terms.However, in other cases, this will not be the case, and a metathesaurus such as the Unified Medical Language System (UMLS) [47] can be used to map synonymous disease terms.

[17] Building a high-quality sense inventory for improved abbreviation disambiguation

  • Authors: Naoaki Okazaki, Sophia Ananiadou, Jun'ichi Tsujii
  • Year: 2010
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/5e28c61947875535bf5edb8960985d3ddc30d716
  • DOI: 10.1093/bioinformatics/btq129
  • PMID: 20360059
  • PMCID: 2859134
  • Citations: 56
  • Influential citations: 3
  • Summary: A supervised approach for clustering expanded forms is presented and the possibility of conflicts of protein and gene names with abbreviations is investigated to investigate the possibility of conflicts of protein and gene names with abbreviations.
  • Evidence snippets:
  • Snippet 1 (score: 0.633)
    > Some researchers have argued that gene symbols are often identical to ambiguous abbreviations (Gaudan et al., 2005;Yu et al., 2006). For example, SCT represents the official gene symbol for the human gene secretin, but it also stands for stem cell transplantation, salmon calcitonin, sacrococcygeal teratoma, etc. (Erhardt et al., 2006). How many protein and gene names actually conflict with abbreviations?
    > To examine the importance of abbreviation disambiguation, we extracted entity names from databases and compared them with the sense inventory. We used entity names in the following resources: description (DE) and gene name (GE) fields in UniProtKB/Swiss-Prot database (as of July 7, 2009); concept names with 'Gene or Genome' type in UMLS (2009AA release as of April 20, 2009); and concept names with 'Amino Acid, Peptide, or Protein' type in UMLS. We assume a database record to have a possible conflict with an abbreviation if the record includes a name that also appears in the abbreviation list. A conflicting name is ambiguous when the sense inventory includes the name as an abbreviation with multiple senses. Table 4 presents the number of database records including abbreviations with at least k senses in the sense inventory. The first row (k ≥ 0) represents the total number of records in each database. Results showed that 149 537 (32.0%) out of 466 739 UniProt records include names that also appear in the abbreviation list (k ≥ 1). Of UniProt records 77 833 (16.7%) have ambiguous abbreviations with multiple senses (k ≥ 2); similarly, 13.2% gene names and 6.4% acid/peptide/protein names in UMLS have possible conflicts with ambiguous abbreviations (k ≥ 2). Moreover, 4 841 (1.0%) of UniProt records are highly ambiguous with at least 30 senses in the abbreviation dictionary. These facts suggest that it is insufficient to identify gene or protein names simply by matching textual expressions with database records.

[18] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.632)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[19] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.632)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[20] Revisiting the Plasmodium falciparum druggable genome using predicted structures and data mining

  • Authors: Karla P. Godinez-Macias, Daisy Chen, J. L. Wallis, Miles G. Siegel, Anna Adam et al.
  • Year: 2024
  • Venue: Research Square
  • URL: https://www.semanticscholar.org/paper/795b2985fdc3cf1cfbd5fda1b3c0502eb4dfe866
  • DOI: 10.21203/rs.3.rs-5412515/v1
  • PMID: 39649165
  • PMCID: 11623766
  • Citations: 3
  • Summary: This study systematically assessed the Plasmodium falciparum genome for proteins amenable to target-based drug discovery, identifying 867 candidate targets with evidence of small molecule binding and blood stage essentiality and implements a generalizable framework for systematically evaluating and prioritizing novel pathogenic disease targets.
  • Evidence snippets:
  • Snippet 1 (score: 0.624)
    > List of genes and genomic features (GFF) for Plasmodium falciparum 3D7 genome (PlasmoDB release 66) was downloaded and protein coding genes were extracted along with their gene annotations and genomic location. Additional genomic annotations were obtained by querying PlasmoDB to extract UniProt and Entrez ID(s), ortholog group (OrthoMCL), protein features (CDS and protein length, molecular weight, isoelectric point), domain annotations (InterPro, PFam, Superfamily), number of transmembrane (TM) domains, and enzyme commission (EC) numbers. Gene function (Gene Ontology; components, functions and processes) was extracted by either PlasmoDB or by querying the InterPro ID under InterPro2GO mapping tool from EMBL-EBI services. Gene essentiality data was obtained for P.
    > falciparum 29 and P. berghei 30,31 parasites that were mapped to their falciparum ortholog using OrthoMCL orthology group IDs. Protein Data Bank (PDB) IDs of crystal structures were obtained by searching either gene symbols, UniProt IDs associated with each gene, or by typing "Plasmodium" in the PDB website search box. A report with gene identi er, organism, accession number, method for structure determination and publication information was extracted for the search hits.
    > Mapping genes to associated literature publications A download from the NCBI FTP site was performed for gene2pubmed.gz (version 2024-02-21) containing taxonomy ID, gene ID (Entrez) and PubMed ID. Gene IDs were mapped to the P. falciparum 3D7 annotation set, and PMIDs matching the criteria were extracted. To include literature references associated with gene symbols, we queried each gene symbol in PubMed using the Eutils 81 efetch function from NCBI; additional information for each publication was obtained pragmatically using the same tool, retrieving title, authors and DOI (digital object identi er).

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  3. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  4. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  5. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  6. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  7. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  8. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  9. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  10. Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al. (2016). The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize. BMC Systems Biology. https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  11. Anthony C. Smith, A. Robinson (2009). MitoMiner, an Integrated Database for the Storage and Analysis of Mitochondrial Proteomics Data. Molecular & Cellular Proteomics : MCP. https://www.semanticscholar.org/paper/206a60d6d387688ce7f880e0dfe67af5a723c7b8
  12. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  13. F. Pfeiffer, D. Oesterhelt (2015). A Manual Curation Strategy to Improve Genome Annotation: Application to a Set of Haloarchael Genomes. Life. https://www.semanticscholar.org/paper/f5983d01e0ac838554f7f5c29481d70a9d728c30
  14. Martin Krallinger, A. Valencia (2005). Text-mining and information-retrieval services for molecular biology. Genome Biology. https://www.semanticscholar.org/paper/558a2745d6e1ac99f77dde88d62566237bd3cfad
  15. Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al. (2025). RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications. Nucleic Acids Research. https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  16. N. Hukerikar, A. Hingorani, F. Asselbergs, C. Finan, A. Schmidt (2024). Prioritising genetic findings for drug target identification and validation.. Atherosclerosis. https://www.semanticscholar.org/paper/80ee965ca8d81196a8281ab055ff7ff79eda31d9
  17. Naoaki Okazaki, Sophia Ananiadou, Jun'ichi Tsujii (2010). Building a high-quality sense inventory for improved abbreviation disambiguation. Bioinformatics. https://www.semanticscholar.org/paper/5e28c61947875535bf5edb8960985d3ddc30d716
  18. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  19. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  20. Karla P. Godinez-Macias, Daisy Chen, J. L. Wallis, Miles G. Siegel, Anna Adam et al. (2024). Revisiting the Plasmodium falciparum druggable genome using predicted structures and data mining. Research Square. https://www.semanticscholar.org/paper/795b2985fdc3cf1cfbd5fda1b3c0502eb4dfe866

OpenScientist

(aceE-deep-research-openscientist.md)
Functional Annotation Report: *aceE* (Pyruvate Dehydrogenase E1 component) in *Pseudomonas putida* KT2440 OpenScientist openscientist-autonomous 2 artifacts 2026-07-11T17:36:43.171185

Functional Annotation Report: aceE (Pyruvate Dehydrogenase E1 component) in Pseudomonas putida KT2440

Target: UniProt Q88QZ5 | Gene aceE | Ordered locus PP_0339
Organism: Pseudomonas putida KT2440 (ATCC 47054 / DSM 6125) — PSEPK
EC: 1.2.4.1 | Cofactor: Thiamine diphosphate (ThDP/TPP)


1. Summary (Answer to the Research Question)

aceE (PP_0339, Q88QZ5) encodes the E1 component (pyruvate dehydrogenase, EC 1.2.4.1) of the pyruvate dehydrogenase multienzyme complex (PDHc). Its primary function is to catalyze the first and rate-limiting step of the complex: the thiamine-diphosphate (ThDP)-dependent oxidative decarboxylation of pyruvate, releasing CO₂ and generating a ThDP-bound C2α-hydroxyethylidene (enamine) intermediate, which E1 then uses to reductively acetylate the lipoyl (lipoamide) prosthetic group of the E2 component. Through the sequential action of E1→E2→E3 the complex converts pyruvate + CoA + NAD⁺ → acetyl-CoA + CO₂ + NADH, the "link reaction" connecting glycolysis to the citric acid cycle. In P. putida, whose glycolysis runs almost exclusively through the Entner–Doudoroff/EDEMP route, this reaction is the principal gateway feeding pyruvate-derived carbon into acetyl-CoA for the TCA cycle, energy metabolism, and biosynthesis. The enzyme functions as a homodimeric peripheral subunit that is non-covalently tethered to the E2 structural core of a large soluble cytoplasmic assembly.


2. Gene / Protein Identity Verification

Attribute Value Consistency check
Gene symbol aceE Matches canonical name for PDH E1 in Gram-negative bacteria (as in E. coli aceEF-lpd operon) ✔
Protein Pyruvate dehydrogenase E1 component Matches UniProt RecName ✔
EC 1.2.4.1 Pyruvate dehydrogenase (acetyl-transferring), ThDP-dependent ✔
Domains PDC_E1_N (IPR035807), PDH_E1 (IPR004660), PDH_E1_M (IPR041621), THDP-binding (IPR029061), PDH/Transketolase (IPR051157) All diagnostic of the ThDP-dependent 2-oxoacid dehydrogenase E1 family ✔
Organism P. putida KT2440

Verdict: The gene symbol, protein description, EC number, and domain architecture are fully mutually consistent. This is an unambiguous, well-characterized enzyme family; annotation is confident. (Note: "aceE" is not ambiguous in bacteria — it is the standard designator for the PDH E1α/E1 subunit. Care is only needed not to conflate the bacterial single-chain E1 with the eukaryotic split E1α/E1β subunits PDHA1/PDHB.)

Sequence-based orthology evidence (this work): The UniProt sequence of Q88QZ5 is an 881-aa single polypeptide. A global Needleman–Wunsch alignment against E. coli K-12 aceE (P0AFG8/ODP1_ECOLI, 887 aa) gives 61.9% amino-acid identity (545/881 identical residues). This is far above the ~30% homology "twilight zone", establishing Q88QZ5 as a confident ortholog of the biochemically characterized E. coli E1p and justifying transfer of the E. coli mechanistic, kinetic, and structural data below. The single ~880-aa chain (vs. the split eukaryotic E1α ~360 aa + E1β ~330 aa) confirms the gammaproteobacterial single-chain E1 architecture that functions as a homodimer. The conserved ThDP-binding GDG motif is present (~residue 224).


3. Primary Molecular Function — the Catalyzed Reaction

Overall complex reaction (link reaction):

pyruvate + CoA-SH + NAD⁺ → acetyl-CoA + CO₂ + NADH + H⁺

Step catalyzed specifically by E1 (aceE):
1. Substrate binding & decarboxylation. Pyruvate binds at the ThDP cofactor. The thiazolium C2-ylide attacks the pyruvate carbonyl to form 2-(2-lactyl)-ThDP (LThDP), which is decarboxylated (loss of CO₂) to yield the resonance-stabilized C2α-carbanion/enamine (2-α-hydroxyethylidene-ThDP) intermediate.
2. Reductive acetylation. The enamine reduces and acetylates the dithiolane of the lipoyl group carried on E2's mobile lipoyl domain, transferring the acetyl (2-carbon) unit and regenerating ThDP.

Substrate specificity: E1 (aceE) is specific for pyruvate as the 2-oxo-acid substrate (2-oxoglutarate is handled by the paralogous OGDC E1o; branched-chain 2-oxoacids by BCKDH). Specificity is imposed by the ThDP-proximal substrate pocket characteristic of the PDH_E1 family. As a ThDP-dependent enzyme, E1 catalysis proceeds through covalent cofactor intermediates common to the ThDP superfamily (transketolase, 2-oxoacid dehydrogenases, decarboxylases).

Ordered reaction sequence (E1's place in it): "The reaction starts with a ThDP-dependent decarboxylation on E1 to an enamine/C2α carbanion, followed by oxidation and acetyl transfer to form S-acetyldihydrolipoamide E2, and then transfer of this acetyl group from the LD [lipoyl domain] to coenzyme A on the [E2 catalytic domain]. The dihydrolipoamide E2 is finally reoxidized by the E3 component" (Song & Jordan, 2012, PMID 22413895). In Gram-negative bacteria — the group that includes P. putida — the complex comprises E1p (pyruvate dehydrogenase/decarboxylase), E2p (dihydrolipoyl acetyltransferase forming a 24-subunit core with multiple E1p/E3 binding sites and mobile lipoyl domains), and E3 (dihydrolipoyl dehydrogenase); the closely related Azotobacter vinelandii γ-proteobacterial complex is the best-characterized structurally (de Kok et al., 1998, PMID 9655933).

Kinetic/mechanistic evidence: In the closely homologous E. coli E1p (aceE), pre-steady-state kinetics show that formation of the LThDP predecarboxylation intermediate is rate-limiting, and that disorder→order transitions of active-site loops upon substrate binding gate covalent catalysis (Balakrishnan et al., 2012, PMID 23088422). E1 is the first and rate-limiting component of the whole complex (Chan et al., 2023, PMID 36723268). Radical/redox mechanisms and the coupling of decarboxylation to reductive acyl transfer in ThDP 2-oxoacid dehydrogenases are reviewed by Tittmann (2009, PMID 19476487).

Cofactor-fold integrity (this work): Sequence analysis of Q88QZ5 confirms an intact ThDP/Mg²⁺-binding signature — a GDG motif at residue 224 followed ~24 residues downstream by the conserved Asn (…MGDGE…IFVINCN…), the diagnostic motif of the transketolase/2-oxoacid-dehydrogenase E1 ThDP-binding fold — indicating a catalytically competent, non-degenerate enzyme.


4. Pathway Context / Biological Process

  • Position in metabolism: PDHc catalyzes the irreversible link reaction between glycolysis and the citric-acid cycle, converting the glycolytic end-product pyruvate into acetyl-CoA — a key substrate for the TCA cycle and fatty-acid synthesis (Bothe & Zdanowicz, 2026, PMID 40808219; Škerlová et al., 2021, PMID 34489474).
  • P. putida-specific context: KT2440 lacks a functional Embden–Meyerhof–Parnas pathway; glucose is catabolized largely via periplasmic oxidation to gluconate and the Entner–Doudoroff pathway, with a cyclic EDEMP architecture that also boosts NADPH supply (Nikel et al., 2015, PMID 26350459). Regardless of upstream route, the pyruvate → acetyl-CoA conversion performed by aceE-containing PDHc is the dominant oxidative decarboxylation node feeding the TCA cycle in this obligate aerobe.
  • Downstream products: Acetyl-CoA feeds the TCA cycle (energy/reducing equivalents), lipid/polyhydroxyalkanoate precursor supply, and biosynthesis; NADH feeds the respiratory chain.
  • Genomic/operon context (this work): In P. putida KT2440, aceE (PP_0339) is immediately adjacent to aceF (PP_0338, Q88QZ6), the E2 acetyltransferase component of PDHc (EC 2.3.1.12, dihydrolipoyllysine-residue acetyltransferase, 546 aa). This aceEF gene cluster mirrors the E. coli aceEF-lpd operon and provides organism-specific genomic evidence that PP_0339 is the E1 of a co-expressed, functional PDH complex whose cognate E2 core sits next to it. Flanking genes (PP_0337 c-di-GMP phosphodiesterase; PP_0340 glnE; PP_0341 waaF) are unrelated, and the shared E3 (dihydrolipoamide dehydrogenase, lpd/lpdG) is encoded elsewhere in the genome, as is typical.
  • Regulation (family-level inference): In E. coli the aceEF-lpd operon is repressed by PdhR, a pyruvate-sensing transcriptional regulator; homologous pyruvate-responsive control is expected for the P. putida locus. (Direct experimental regulation data for PP_0339 were not retrieved here and are flagged as a knowledge gap.)

5. Structural Organization & Subcellular Localization

  • Quaternary structure: PDHc is one of the largest known enzyme assemblies (~5–12 MDa). E2 (dihydrolipoamide acetyltransferase) forms the structural core (octahedral/cubic in most bacteria, icosahedral in others), and E1 and E3 bind the core as peripheral subunits (Bothe & Zdanowicz, 2026, PMID 40808219). Complex integrity is maintained by non-covalent tethering of the peripheral E1 and E3 to E2 via the E2 peripheral-subunit-binding domain (PSBD) (Arjunan et al., 2014, PMID 25210042).
  • Oligomeric state of E1: In gammaproteobacteria (which includes Pseudomonas), E1p is a homodimer (Meinhold et al., 2024, PMID 38324697; Arjunan et al., 2014). Each aceE monomer contributes to a shared active-site environment for ThDP.
  • Substrate channeling: E2's covalently attached, swinging lipoyl domains shuttle reaction intermediates between the E1, E2, and E3 active sites; cryo-EM of the native E. coli E2 core reveals how lipoyl domains dock at active sites (Škerlová et al., 2021, PMID 34489474).
  • Localization: The complex is a soluble cytoplasmic assembly in bacteria (in eukaryotes it is mitochondrial). aceE therefore carries out its function in the cytoplasm/cytosol of P. putida.

6. Supported vs. Refuted Hypotheses

Supported:
- H1 — aceE is a ThDP-dependent pyruvate dehydrogenase E1 (EC 1.2.4.1) catalyzing the first, rate-limiting step of PDHc. Supported (domain architecture + homolog kinetics).
- H2 — Its physiological role is producing acetyl-CoA linking glycolysis (ED/EDEMP in P. putida) to the TCA cycle. Supported.
- H3 — aceE acts as a peripheral homodimer tethered to the cytoplasmic E2 core. Supported (structural literature).

Refuted / excluded:
- aceE is not an isolated soluble monomeric enzyme, and not a membrane transporter or structural protein — it is an enzymatic subunit of a large multienzyme complex.
- The bacterial aceE is a single-chain E1, distinct from the split eukaryotic E1α (PDHA1)/E1β architecture; literature on human PDHA1 describes the orthologous chemistry but a different subunit organization (excluded as a direct structural analogue).


7. Evidence Quality & Limitations

  • Strength: The catalytic chemistry, cofactor, mechanism, and complex architecture are established by decades of precise biochemistry, kinetics, X-ray crystallography, and cryo-EM — largely on the near-identical E. coli E1p and on other bacterial/mammalian PDHc. Given very high sequence/domain conservation, transfer of this mechanism to P. putida PP_0339 is well justified.
  • Limitations / gaps specific to P. putida KT2440:
  • No P. putida-specific crystal/cryo-EM structure or steady-state kinetic constants (Km for pyruvate, kcat) for PP_0339 were retrieved — inference is by orthology.
  • Transcriptional regulation of the P. putida aceE locus (PdhR-like control, growth-condition dependence) was not directly documented here.
  • Quantitative flux through PDH vs. alternative pyruvate-consuming routes (e.g., pyruvate carboxylation, transhydrogenase-linked cycles) in KT2440 warrants organism-specific confirmation.
  • Future directions: Determine PP_0339 kinetic parameters and the P. putida PDHc structure; test PdhR-type regulation; ¹³C-flux quantification of the pyruvate→acetyl-CoA node under industrially relevant carbon sources.

8. Key References

  • Bothe & Zdanowicz (2026) Structural diversity of pyruvate dehydrogenase complexes. PMID 40808219
  • Škerlová et al. (2021) Structure of the native pyruvate dehydrogenase complex reveals the mechanism of substrate insertion. PMID 34489474
  • Arjunan et al. (2014) Novel binding motif… E1p–E2p subcomplex from the E. coli PDH complex. PMID 25210042
  • Balakrishnan et al. (2012) Pre-steady-state rate constants on the E. coli PDH complex… loop movement controls the rate-limiting step. PMID 23088422
  • Chan et al. (2023) Furan-based inhibitors of pyruvate dehydrogenase (PDH E1 as TPP-dependent, rate-limiting). PMID 36723268
  • Tittmann (2009) Reaction mechanisms of thiamin diphosphate enzymes: redox reactions. PMID 19476487
  • Meinhold et al. (2024) Dimerization of a 5-kDa domain defines the architecture of the 5-MDa gammaproteobacterial PDH complex. PMID 38324697
  • Nikel et al. (2015) P. putida KT2440 metabolizes glucose through an ED/EMP/PPP cycle (EDEMP). PMID 26350459
  • Wang et al. (2014) Structure and function of the E2 catalytic domain in E. coli PDHc. PMID 24742683
  • Träger et al. (2026) The Pyruvate Dehydrogenase Complex: A 90-Year-Old Enigma… PMID 42334543
  • Song & Jordan (2012) Interchain acetyl transfer in the E2 component of bacterial pyruvate dehydrogenase… (defines the ordered E1→E2→E3 reaction). PMID 22413895
  • de Kok et al. (1998) The pyruvate dehydrogenase multi-enzyme complex from Gram-negative bacteria. PMID 9655933

Artifacts

📄 View Raw YAML

id: Q88QZ5
gene_symbol: aceE
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: aceE (PP_0339) encodes the E1 component of pyruvate dehydrogenase, a thiamine-diphosphate enzyme that decarboxylates pyruvate and transfers the resulting hydroxyethyl/acetyl equivalent to the lipoyl group of the E2 component. Together with AceF and lipoamide dehydrogenase, it supports oxidative decarboxylation of pyruvate to acetyl-CoA in central carbon metabolism.
existing_annotations:
- term:
    id: GO:0004739
    label: pyruvate dehydrogenase (acetyl-transferring) activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: pyruvate dehydrogenase (acetyl-transferring) activity is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
    action: ACCEPT
    reason: This is a specific, biologically appropriate annotation for this gene product.
- term:
    id: GO:0016491
    label: oxidoreductase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: oxidoreductase activity is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
    action: KEEP_AS_NON_CORE
    reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: file:PSEPK/aceE/aceE-uniprot.txt
  title: UniProt record for aceE (Q88QZ5)
  findings:
  - statement: UniProt identifies aceE as Pyruvate dehydrogenase E1 component (EC 1.2.4.1) and provides the seeded EC/domain/GO evidence reviewed here.
- id: file:PSEPK/aceE/aceE-deep-research-asta.md
  title: Asta deep-research retrieval for aceE
  findings:
  - statement: Asta retrieval was run for this first-pass pathway curation; direct organism-specific literature was limited for several common enzyme names, so UniProt/family evidence carries the main review weight.
aliases:
- PP_0339
core_functions:
- description: pyruvate dehydrogenase (acetyl-transferring) activity supporting the Pyruvate dehydrogenase E1 component (EC 1.2.4.1) role summarized for aceE.
  supported_by:
  - reference_id: file:PSEPK/aceE/aceE-uniprot.txt
    supporting_text: DR   GO; GO:0004739; F:pyruvate dehydrogenase (acetyl-transferring) activity; IEA:UniProtKB-EC.
  molecular_function:
    id: GO:0004739
    label: pyruvate dehydrogenase (acetyl-transferring) activity
  directly_involved_in:
  - id: GO:0006086
    label: pyruvate decarboxylation to acetyl-CoA
proposed_new_terms: []