aceF

UniProt ID: Q88QZ6
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
Aliases:
PP_0338
📝 Provide Detailed Feedback

Gene Description

aceF (PP_0338) encodes the E2 acetyltransferase component of the pyruvate dehydrogenase complex. The protein carries lipoyl domains and a catalytic acyltransferase domain that transfers the acetyl group from S-acetyldihydrolipoyllysine to coenzyme A, coupling the E1 pyruvate-decarboxylating reaction to acetyl-CoA formation. It is a cytoplasmic component of the multienzyme pyruvate dehydrogenase complex, connecting pyruvate produced by lower central carbon metabolism to acetyl-CoA and the TCA cycle.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0004742 dihydrolipoyllysine-residue acetyltransferase activity
IEA
GO_REF:0000120
ACCEPT
Summary: dihydrolipoyllysine-residue acetyltransferase activity is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
Reason: This is a specific, biologically appropriate annotation for this gene product.
GO:0005737 cytoplasm
IEA
GO_REF:0000118
ACCEPT
Summary: cytoplasm is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
Reason: This is a specific, biologically appropriate annotation for this gene product.
GO:0006086 pyruvate decarboxylation to acetyl-CoA
IEA
GO_REF:0000120
ACCEPT
Summary: pyruvate decarboxylation to acetyl-CoA is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
Reason: This is a specific, biologically appropriate annotation for this gene product.
GO:0016407 acetyltransferase activity
IEA
GO_REF:0000118
KEEP AS NON CORE
Summary: acetyltransferase activity is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
Reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
GO:0016746 acyltransferase activity
IEA
GO_REF:0000002
KEEP AS NON CORE
Summary: acyltransferase activity is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
Reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
GO:0031405 lipoic acid binding
IEA
GO_REF:0000118
KEEP AS NON CORE
Summary: lipoic acid binding is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
Reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
GO:0045254 pyruvate dehydrogenase complex
IEA
GO_REF:0000120
ACCEPT
Summary: pyruvate dehydrogenase complex is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
Reason: This is a specific, biologically appropriate annotation for this gene product.

Core Functions

dihydrolipoyllysine-residue acetyltransferase activity supporting the Acetyltransferase component of pyruvate dehydrogenase complex (EC 2.3.1.12) role summarized for aceF.

Supporting Evidence:
  • file:PSEPK/aceF/aceF-uniprot.txt
    DR GO; GO:0004742; F:dihydrolipoyllysine-residue acetyltransferase activity; IEA:UniProtKB-UniRule.

References

Gene Ontology annotation through association of InterPro records with GO terms
TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
file:PSEPK/aceF/aceF-uniprot.txt
UniProt record for aceF (Q88QZ6)
  • UniProt identifies aceF as Acetyltransferase component of pyruvate dehydrogenase complex (EC 2.3.1.12) and provides the seeded EC/domain/GO evidence reviewed here.
file:PSEPK/aceF/aceF-deep-research-asta.md
Asta deep-research retrieval for aceF
  • Asta retrieval was run for this first-pass pathway curation; direct organism-specific literature was limited for several common enzyme names, so UniProt/family evidence carries the main review weight.

Deep Research

Asta

(aceF-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-06T05:23:51.721524

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.750)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[2] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.715)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[3] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.689)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[4] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.688)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[5] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.686)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[6] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[7] The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize

  • Authors: Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al.
  • Year: 2016
  • Venue: BMC Systems Biology
  • URL: https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  • DOI: 10.1186/s12918-016-0369-x
  • PMID: 27899149
  • PMCID: 5129634
  • Citations: 11
  • Influential citations: 1
  • Summary: CornCyc’s computational predictions are more accurate than those in MaizeCyc when compared to experimentally determined function assignments, demonstrating the relative strength of the enzymatic function assignment pipeline used to generate CornCyc.
  • Evidence snippets:
  • Snippet 1 (score: 0.668)
    > A gold standard set of protein functional annotations was generated by extracting data from UniProt [16] and BRENDA [17]. We extracted all protein sequence and annotation data from UniProt (release 2016_05) for the organism Zea mays, keeping the EC annotations only from the manually reviewed component of UniProt, while removing those annotations that had not undergone manual review. We also extracted experimentally verified protein annotations for Zea mays from BRENDA (release 2016.1). The UniProt and BRENDA annotations were then merged by matching proteins based on the database crosslinks provided by BRENDA, resulting in the union of the reviewed annotations from UniProt and the experimentally verified annotations of BRENDA with duplicates removed. The merged protein annotations were then matched to the B73 RefGen_v2 translated gene models using BLASTP based on a sequence identity cutoff of 96% and an e-value cutoff of 1e-20. We selected the top scoring hit for each protein which resulted in matches to 1,815 unique maize proteins. EC annotations for alternate isoforms were consolidated at the gene level, resulting in 1,475 experimentally verified or manually reviewed protein functional annotations across 1,450 maize genes.

[8] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[9] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[10] Ten steps to get started in Genome Assembly and Annotation

  • Authors: Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al.
  • Year: 2018
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  • DOI: 10.12688/f1000research.13598.1
  • PMID: 29568489
  • PMCID: 5850084
  • Citations: 109
  • Influential citations: 1
  • Summary: Ten steps to facilitate researchers getting started in genome assembly and genome annotation are presented and the importance of data management is stressed, and advice on where to submit data and how to make results Findable, Accessible, Interoperable, and Reusable (FAIR).
  • Evidence snippets:
  • Snippet 1 (score: 0.665)
    > The ultimate goal of the functional annotation process (Figure 4) is to assign biologically relevant information to predicted polypeptides, and to the features they derive from (e.g. gene, mRNA). This process is especially relevant nowadays in the context of the NGS era due to the capacity of sequencing, assembling, and annotating full genomes in short periods of time, e.g. less than a month. Functional elements could range from putative name and/or symbols for protein-coding genes, e.g. ADH to its putative biological function, e.g. alcohol dehydrogenase, associated gene ontology terms, e.g. GO:0004022, functional sites, e.g. METAL 47 47 Zinc 1, and domains, e.g. IPR002328, among other features. The function of predicted proteins can be computationally inferred based on the similarity between the sequence of interest and other sequences in different public repositories, e.g. BLASTP against Uniprot. Caution should be taken when assigning results merely based on sequence similarity as two evolutionary independent sequences which share some common domains could be considered homologs 62 . Thus, whenever possible, it is better to use orthologous sequences for annotation purposes rather than simply similar sequences 63 . With the growing number of sequences in those public repositories, it is possible to perform various searches and combine obtained results into a consensus annotation. The accurate assignment of the functional elements is a complex process, and the best annotation will involve manual curation.
    > There are two main outcomes of the functional annotation process. The first is the assignment of functional elements to genes. Downstream analysis of these elements allow further understanding of specific genome properties, e.g. metabolic pathways, and similarities compared with closely related species. The second result of the functional annotation is the additional quality check for the predicted gene set. It is possible to identify problematic and/or suspicious genes by the presence of specific domains, suspicious orthology assignment and/or absence of other functional elements, e.g. functional completeness. These Page 13 of 19

[11] CellPhoneDB: inferring cell–cell communication from combined expression of multi-subunit ligand–receptor complexes

  • Authors: M. Efremova, Miquel Vento-Tormo, S. Teichmann, R. Vento-Tormo
  • Year: 2020
  • Venue: Nature Protocols
  • URL: https://www.semanticscholar.org/paper/6a3b3e4a2eebc3fad3b03ee6d8228264247abbc8
  • DOI: 10.1038/s41596-020-0292-x
  • PMID: 32103204
  • Citations: 2770
  • Influential citations: 298
  • Summary: The structure and content of CellPhoneDB is outlined, procedures for inferring cell–cell communication networks from single-cell RNA sequencing data are provided and a practical step-by-step guide to help implement the protocol is presented.
  • Evidence snippets:
  • Snippet 1 (score: 0.662)
    > CellPhoneDB stores ligand-receptor interactions, as well as other properties of the interacting partners, including their subunit architecture and gene and protein identifiers. To create the content of the database, four main .csv data files are required: gene_input.csv, protein_input.csv, complex_input.csv and interaction_input.csv (Fig. 4).
    > gene_input Mandatory fields are 'gene_name', 'uniprot', 'hgnc_symbol' and 'ensembl'. This file is critical for establishing the link between the scRNA-seq data and the interaction pairs stored at the protein level. It includes the following gene and protein identifiers: (i) gene name ('gene_name'), (ii) UniProt identifier ('uniprot'), (iii) HUGO Nomenclature Committee (HGNC) symbol ('hgnc_symbol') and (iv) gene Ensembl identifier (ENSG) ('ensembl'). To create this file, lists of linked proteins and gene identifiers are downloaded from UniProt and merged using gene names. Several rules need to be considered when merging the files • UniProt annotation prevails over the gene Ensembl annotation when the same gene Ensembl identifier points toward different UniProt identifiers. • UniProt and Ensembl lists are also merged by their UniProt identifier, but this information is used only when the UniProt or Ensembl identifier is missing in the original list merged by gene name. • If the same gene name points toward different HGNC symbols, only the HGNC symbol matching the gene name annotation is considered. • Only one HLA isoform is considered in our interaction analysis, and it is stored in a manually HLA-curated list of genes, named HLA_curated.
    > protein_input Mandatory fields are 'uniprot' and 'protein_name'.

[12] RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications

  • Authors: Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al.
  • Year: 2025
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  • DOI: 10.1093/nar/gkaf1206
  • PMID: 41206768
  • PMCID: 12807602
  • Citations: 2
  • Summary: The new RM2Target v2.0 will serve as a foundational resource for exploring RNA epitranscriptomic regulation, enabling investigations into cross-talk among modifications, underlying molecular mechanisms, and disease connections, thereby facilitating both basic research and translational applications in RNA epigenetics.
  • Evidence snippets:
  • Snippet 1 (score: 0.660)
    > To obtain basic information on WERs and their target genes, such as official gene symbols, gene IDs, gene types, and genomic locations, gene annotations were downloaded from the GENCODE project [ 44 ] for human and mouse, and from NCBI [ 45 ] and Ensembl [ 46 ] for the other species. Genomic locations were extracted from the corresponding GTF annotation files. Gene symbols were primarily standardized based on the NCBI Gene database [ 45 ] for mRNAs and lncRNAs, GtR-NAdb [ 47 ] for tRNAs, miRbase [ 48 ] for microRNAs, and cir-cBase [ 49 ] for circRNAs. Deprecated or substituted versions of genes were filtered out. The LiftOver [ 50 ] program was employed to convert and unify genomic coordinates across different genome assembly versions.
    > The functional descriptions of WERs were compiled based on the UniProt database [ 51 ] and further supplemented with evidence from relevant publications, with particular emphasis on their functions as RNA modification regulatory proteins.

[13] Into the metabolic wild: Unveiling hidden pathways of microbial metabolism

  • Authors: Özge Ata, D. Mattanovich
  • Year: 2024
  • Venue: Microbial Biotechnology
  • URL: https://www.semanticscholar.org/paper/81ac8bfd05a0c52f8a5b6d058c5414b45bb5c1f9
  • DOI: 10.1111/1751-7915.14548
  • PMID: 39126421
  • PMCID: 11316390
  • Citations: 5
  • Summary: The predictive power of metabolic modelling, well‐founded on biochemical knowledge and genomic information is discussed in the light of both discovery of yet unknown existing metabolic routes and the prediction of others, new to Nature.
  • Evidence snippets:
  • Snippet 1 (score: 0.659)
    > One major concern regarding genome data is the potential spread of false or inaccurate functional annotations iterating by sequence similarity-based gene annotations where functions are ascribed by analogy to other organisms' genes rather than being experimentally verified. As the source of information on a putative gene function is not always clearly indicated in databases, we advise researchers to always trace the original literature related to a gene function to verify its plausibility in a given non-model organism.
    > A special case is the genes with no known functional annotation in model organisms, like the y-genes in E. coli (so named because their gene names begin with the letter y). Ghatak et al. (2019) used multiple databases to define the y-ome of E. coli as the 35% genes that lack experimental evidence of their function, and highlight the differences between genes that lack any known function and those with attributed localization or computationally annotated functional domains. A similar situation is observed with eukaryotic genomes, where a consistent fraction of about 20% of the genes remains with unknown (or unstudied) gene products. Wood et al. (2019) compared unknown gene products of Schizosaccharomyces pombe, S. cerevisiae, and Homo sapiens, and concluded that these genes remained understudied even since genome sequences were published. They disprove the common assumption that unknown gene products are mostly orphaned, as the majority are conserved from yeasts to humans, and they conclude that the reasons for the lack of research on these proteins is two-fold. For many of these genes, no apparent knockout phenotype was observed in standard conditions. On the other hand, researchers still focus on a few well-studied proteins rather than exploring unknown territories, which the authors partly ascribe to risk-adverse strategies of funders and reviewers.
    > Although the primary use of metabolic models is the analysis of already known metabolic pathways, the potential of metabolic models extends far beyond these analyses. They can be used to uncover hidden pathways, unknown reactions, and to construct de novo pathways.

[14] Next Generation Sequencing and Transcriptome Analysis Predicts Biosynthetic Pathway of Sennosides from Senna (Cassia angustifolia Vahl.), a Non-Model Plant with Potent Laxative Properties

  • Authors: Nagaraja Reddy Rama Reddy, Rucha Harishbhai Mehta, Palak Soni, Jayanti Makasana, N. Gajbhiye et al.
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/fec6263b90f9cd765752bdb1be3d162872f2e64e
  • DOI: 10.1371/journal.pone.0129422
  • PMID: 26098898
  • PMCID: 4476680
  • Citations: 81
  • Influential citations: 3
  • Summary: A set of putative genes involved in various secondary metabolite pathways, especially those related to the synthesis of sennosides are identified which will serve as an important platform for public information about gene expression, genomics, and functional genomics in senna.
  • Evidence snippets:
  • Snippet 1 (score: 0.659)
    > Further the assembled transcript contigs were validated using CLC Genomics workbench (CLC Bio, Boston, MA 02108 USA) by mapping high quality reads back to the assembled transcript contigs. ORF-Predictor [57], an online tool, was used on default parameters to identify the coding DNA sequences (CDS) from assembled transcript contigs. GC counts of transcripts was determined using a custom-made perl script.
    > Functional annotations. The functional annotation was performed by aligning coding DNA sequence (CDS) to NCBI 'green plant database (txid 33090)' database using basic local alignment search tool (BLASTX) [58] with an E-value threshold of 1e -06 and GO assignments were used to classify the functions of the predicted CDS. The GO mapping also provided ontology of defined terms representing gene product properties which were grouped into three main domains: biological process (BP), molecular function (MF) and cellular component (CC). GO mapping was carried out in order to retrieve GO terms for all the BLASTX functionally annotated CDS. The GO mapping used defined criteria to retrieve GO terms for annotated CDS which included use of BLASTX result accession IDs to retrieve gene names or symbols, UniProt IDs and direct search in the dbxref table of GO database. Identified gene names or symbols were then searched in the species specific entries of the gene-product tables of GO database. UniProt IDs made use of protein information resource (PIR) which includes protein sequence database (PSD), UniProt, SwissProt, TrEMBL, RefSeq, GenPept, and PDB databases. Gene Ontology analysis helps in specifying all the annotated nodes comprising of GO functional groups. CDS were compared against the COG (Clusters of Orthologous Groups) database for the analysis of phylogenetically widespread domain families. CDS were compared against Pfam database for higher-level groupings of related protein families, known as clans and the identification of domains that occurs within proteins. BLASTX was used against uniprot-swissprot database with cut-off e-value 1e-6 to annotate predicted CDS against protein.

[15] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[16] The Protein Naming Utility: a rules database for protein nomenclature

  • Authors: Johannes Goll, R. Montgomery, L. Brinkac, S. Schobel, D. Harkins et al.
  • Year: 2009
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/dad2c69d04a5a3abdeda129612d52a7849531c09
  • DOI: 10.1093/nar/gkp958
  • PMID: 20007151
  • PMCID: 2808875
  • Citations: 8
  • Summary: The Protein Naming Utility (PNU) is a web-based database for storing and applying naming rules to identify and correct syntactically incorrect protein names, or to replace synonyms with their preferred name.
  • Evidence snippets:
  • Snippet 1 (score: 0.647)
    > During the annotation phase of a typical modern genomics project, functional names are assigned to identified genes and proteins in an automated or semiautomated fashion. Ideally, before such names are submitted to public sequence databases, they should be manually reviewed by experts to ensure that they are consistent, syntactically correct and unambiguous. However, with the scale of genomic data produced by next-generation sequencing technology and with increasingly automated functional annotation processes, the manual correction of names is no longer feasible. This issue is further complicated by the prevalence of ambiguous names resulting from the lack of interspecies naming conventions (1). New proteins are often named based on homology to existing proteins and many existing proteins have syntactically incorrect or ambiguous names, producing transitive annotation errors. Consequently, poor-quality names have proliferated in both public databases and the scientific literature.
    > The need for consistent and unambiguous names has led to the development of a number of conventions for naming genes and proteins [UniProt protein nomenclature (2), HUGO human gene name nomenclature (3) and various other model organism databases (4)(5)(6)(7)]. In addition, the biological text mining community has created dictionaries to resolve gene/protein synonyms to improve the identification of genes and proteins in scientific articles (1,8).
    > The Broad Institute has developed BioNames, a tool to resolve these difficulties using collections of hard-coded regular expressions (https://sourceforge.net/projects/ microbiomeutil). Here, we present our solution to this problem in the form of the Protein Naming Utility (PNU), a web-based database to store and apply customizable sets of naming rules to correct and standardize gene and protein names within an annotated genome or metagenome. The database provides an intuitive web interface that allows users to create and maintain their own naming rules and organize these rules in projects that can be shared with the community.

[17] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.647)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[18] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.647)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[19] Tandem mass tag labeled quantitative proteomic analysis of differential protein expression on total alkaloid of Aconitum flavum Hand.-Mazz. against melophagus ovinus

  • Authors: Xinjian Wang, Zhen Yang, Yujun Zhang, Feng Cheng, Xiaoyong Xing et al.
  • Year: 2022
  • Venue: Frontiers in Veterinary Science
  • URL: https://www.semanticscholar.org/paper/1bbb99fff242bf51635626ad3578683e39182d20
  • DOI: 10.3389/fvets.2022.951058
  • PMID: 35968012
  • PMCID: 9365070
  • Citations: 2
  • Summary: The mechanism of Aconitum flavum Hand.-Mazz.
  • Evidence snippets:
  • Snippet 1 (score: 0.644)
    > The Gene Ontology, or GO, is a major bioinformatics initiative to unify the representation of genes and gene product attributes across all species. More specifically, the project aims to maintain and develop its controlled vocabulary of genes and gene product attributes, annotate genes and gene products, assimilate and disseminate annotation data, and provide tools to facilitate access to all aspects of the project data. The gene ontology covers three domains including cellular component, molecular function, and biological process. The cellular component is a component of a cell, but with the proviso that it is part of some larger object. This may be an anatomical structure (e.g., rough endoplasmic reticulum or nucleus) or a gene product group (e.g., ribosome, proteasome, or a protein dimer). Molecular function describes activities, such as catalytic or binding activities, that occur at the molecular level. GO molecular function terms represent activities rather than the entities (molecules or complexes) that perform the actions and do not specify where, when, or in which context the action takes place. The biological process is a series of events accomplished by an ordered assembly of one or more molecular functions. Biological processes and molecular functions are difficult to distinguish, but the general rule is that a process must have multiple distinct steps. Gene ontology (GO) annotation proteome was derived from the UniProt-GOA database (http:// www.ebi.ac.uk/GOA/). First, the identified protein IDs were converted to UniProt IDs and then mapped to GO IDs by protein IDs. If some identified proteins were not annotated by the UniProt-GOA database, the InterProScan software would be used to the annotated protein's GO functional based on the protein sequence alignment method. Then, proteins were classified by Gene Ontology annotation based on three categories: biological process, cellular component, and molecular function.

[20] Protocol for gene annotation, prediction, and validation of genomic gene expansion

  • Authors: Quanwei Zhang, Zhengdong D. Zhang
  • Year: 2022
  • Venue: STAR Protocols
  • URL: https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3
  • DOI: 10.1016/j.xpro.2022.101692
  • PMID: 36125934
  • PMCID: 9494284
  • Citations: 1
  • Summary: A detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation is described and steps to discover functionality of each copy of replicated genes are detailed.
  • Evidence snippets:
  • Snippet 1 (score: 0.642)
    > 3. Gene annotation and functional annotation. a. Gene structure annotation.
    > In addition to gene prediction models, evidence from orthologous protein sequences and transcriptome assembly could be used to improve annotation quality. Protein sequences of orthologous genes can be obtained from UniProt (The UniProt, 2017). Ones from Swiss-Port have been reviewed and thus are of higher quality. Transcriptome assembly may be available from previous studies or can be assembled de novo from RNA-seq reads by Trinity (Haas et al., 2013). High quality transcriptome assembly can be selected as described in (Zhang et al., 2021). Note: Details about gene structure annotation (Holt and Yandell, 2011) can be found at http:// gmod.org/wiki/MAKER_Tutorial, https://darencard.net/blog/2017-05-16-maker-genomeannotation/, and the protocol (Campbell et al., 2014).
    > b. Quality measurement and functional annotation.
    > For each predicted gene, Maker2 provides the annotation edit distance (AED) score, which measures the goodness of fit between its predicted gene structure and its evidence support. The lower the score, the more accurate the prediction. If more than 90% genes with AED scores lower than 0.5, the genome can be considered well annotated. In addition to the AED score, a high proportion of recognizable domains contained in predicted protein -e.g., higher than 50% -also indicates a good annotation. Recognizable protein domains can by scanned by InterProScan (Jones et al., 2014), assigning potential function to predicted genes.
    > Note: Besides the aforementioned quality measurement, we strongly recommend measuring the completeness of the genome assembly and annotation by checking the existence of a set of Benchmarking Universal Single-Copy Orthologs (BUSCO) (Simao et al., 2015). A high-level completeness of genome assembly and annotation is imperative for a better identification of gene expansion. Based on the result of this analysis, researchers can decide whether they need to further improve the genome assembly before predicting gene expansion. A detailed protocol of BUSCO is available at

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  2. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  3. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  4. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  5. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  6. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  7. Jesse R. Walsh, M. Schaeffer, Peifen Zhang, S. Rhee, J. Dickerson et al. (2016). The quality of metabolic pathway resources depends on initial enzymatic function assignments: a case for maize. BMC Systems Biology. https://www.semanticscholar.org/paper/c41be7766c80fddb3f81c57ced799b8562370cc8
  8. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  9. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  10. Victoria Dominguez Del Angel, Erik Hjerde, L. Sterck, S. Capella-Gutiérrez, C. Notredame et al. (2018). Ten steps to get started in Genome Assembly and Annotation. F1000Research. https://www.semanticscholar.org/paper/1b1090dcbd0f6a609f0448957b7e464997879ea8
  11. M. Efremova, Miquel Vento-Tormo, S. Teichmann, R. Vento-Tormo (2020). CellPhoneDB: inferring cell–cell communication from combined expression of multi-subunit ligand–receptor complexes. Nature Protocols. https://www.semanticscholar.org/paper/6a3b3e4a2eebc3fad3b03ee6d8228264247abbc8
  12. Xiaoqiong Bao, Qi Jiang, Weixuan Chen, Huiqin Li, Xuan Li et al. (2025). RM2Target v2.0: an updated database for the target genes of writers, erasers, and readers of RNA modifications. Nucleic Acids Research. https://www.semanticscholar.org/paper/15dbf5509222f17a624912633aa05cc9183fcc70
  13. Özge Ata, D. Mattanovich (2024). Into the metabolic wild: Unveiling hidden pathways of microbial metabolism. Microbial Biotechnology. https://www.semanticscholar.org/paper/81ac8bfd05a0c52f8a5b6d058c5414b45bb5c1f9
  14. Nagaraja Reddy Rama Reddy, Rucha Harishbhai Mehta, Palak Soni, Jayanti Makasana, N. Gajbhiye et al. (2015). Next Generation Sequencing and Transcriptome Analysis Predicts Biosynthetic Pathway of Sennosides from Senna (Cassia angustifolia Vahl.), a Non-Model Plant with Potent Laxative Properties. PLoS ONE. https://www.semanticscholar.org/paper/fec6263b90f9cd765752bdb1be3d162872f2e64e
  15. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  16. Johannes Goll, R. Montgomery, L. Brinkac, S. Schobel, D. Harkins et al. (2009). The Protein Naming Utility: a rules database for protein nomenclature. Nucleic Acids Research. https://www.semanticscholar.org/paper/dad2c69d04a5a3abdeda129612d52a7849531c09
  17. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  18. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  19. Xinjian Wang, Zhen Yang, Yujun Zhang, Feng Cheng, Xiaoyong Xing et al. (2022). Tandem mass tag labeled quantitative proteomic analysis of differential protein expression on total alkaloid of Aconitum flavum Hand.-Mazz. against melophagus ovinus. Frontiers in Veterinary Science. https://www.semanticscholar.org/paper/1bbb99fff242bf51635626ad3578683e39182d20
  20. Quanwei Zhang, Zhengdong D. Zhang (2022). Protocol for gene annotation, prediction, and validation of genomic gene expansion. STAR Protocols. https://www.semanticscholar.org/paper/af8e3a73daaa04214d43f4ec1d9b1c0fcd42b8e3

OpenScientist

(aceF-deep-research-openscientist.md)
Functional Annotation Report: *aceF* (PP_0338, UniProt Q88QZ6) OpenScientist openscientist-autonomous 2 artifacts 2026-07-11T17:37:13.955097

Functional Annotation Report: aceF (PP_0338, UniProt Q88QZ6)

Acetyltransferase (E2) component of the pyruvate dehydrogenase complex

Organism: Pseudomonas putida KT2440 (strain ATCC 47054 / DSM 6125 / NCIMB 11950), a Gammaproteobacterium of the order Pseudomonadales.


1. Summary (Answer to the Research Question)

aceF encodes the dihydrolipoyllysine-residue acetyltransferase (E2) component of the pyruvate dehydrogenase multienzyme complex (PDHc)EC 2.3.1.12. Its primary enzymatic function is to catalyze the transfer of an acetyl group from the reductively-acetylated lipoyl arm (S8-acetyldihydrolipoyl-lysine) to coenzyme A, producing acetyl-CoA:

N6-[(R)-S8-acetyldihydrolipoyl]-L-lysyl-[protein] + CoA ⇌ N6-[(R)-dihydrolipoyl]-L-lysyl-[protein] + acetyl-CoA

Beyond catalysis, E2 is the structural and organizational heart of PDHc: 24 copies of its catalytic domain self-assemble into a hollow cube (octahedral 432 symmetry) that forms the core to which the peripheral E1 (pyruvate dehydrogenase, aceE/PP_0339) and E3 (dihydrolipoamide dehydrogenase, lpd) subunits attach. The enzyme functions in the cytoplasm and sits at the pivotal metabolic node linking sugar catabolism (in P. putida, principally the Entner–Doudoroff/EDEMP route converging on pyruvate) to the TCA cycle and acetyl-CoA-dependent biosynthesis.

Gene-identity verification: the symbol aceF, the protein description, the 2-oxoacid-dehydrogenase family assignment, and the InterPro domains (lipoyl/biotinyl-binding, 2-oxoacid DH acyltransferase, PSBD) all mutually agree. This is an unambiguous, well-characterized enzyme; no symbol-collision problem exists.


2. Molecular Identity and Domain Architecture

The 546-residue protein (UniProt Q88QZ6) shows the canonical modular architecture of a Gram-negative PDHc E2, confirmed from the UniProt feature table:

Region Residues Module Role
Lipoyl domain 1 2–75 Biotinyl/lipoyl-binding (β-barrel) Carries covalent (R)-lipoate on a conserved Lys (in the …LESDKASMEIP… motif); "swinging arm"
Lipoyl domain 2 117–191 Biotinyl/lipoyl-binding (β-barrel) Second lipoyl-lysine (second …LESDKASMEIP… motif)
Linkers Ala/Pro-rich Flexible hinges Allow the lipoyl arms to visit E1, E2 and E3 active sites
PSBD 245–282 Peripheral subunit-binding domain Docks E1 and E3 onto the E2 core
Catalytic domain ~290–546 Acetyltransferase (chloramphenicol-acetyltransferase fold) Acetyl transfer to CoA; core assembly

Notable points:
- Two lipoyl domains — intermediate between E. coli (three lipoyl domains) and mammalian/Bacillus E2 (typically one to two). Multiple lipoyl domains increase the effective local concentration and reach of the acetyl-carrying arm.
- The catalytic domain contains a conserved His-Ser-Asn catalytic set, all present in Q88QZ6 (see §3): catalytic Ser468 (in the SSLGH motif), catalytic His519 (in the DHR motif), and Asn523 just downstream. Cofactor: covalently bound (R)-lipoate (UniProt COFACTOR) on Lys41 and Lys157.

Cofactor loading (activation): the lipoyl domains are catalytically inert until a lipoyl group is attached to their conserved lysines. This is done either de novoLipB octanoylates the lysine and the radical-SAM enzyme LipA inserts two sulfur atoms to form lipoate — or by LplA-mediated salvage of exogenous lipoate; in the absence of this modification the dehydrogenase is inactive and aerobic metabolism is blocked (PMID 21209092). P. putida KT2440 encodes the orthologous lipoylation machinery.


3. Primary Catalytic Function and Mechanism

Reaction (EC 2.3.1.12): E2 catalyzes reversible transacetylation between the protein-bound dihydrolipoyl-lysine and CoA. In the physiological (PDHc) direction, E1 first decarboxylates pyruvate and reductively acetylates the E2 lipoyl-lysine; E2 then transfers that acetyl group to CoA to yield acetyl-CoA, leaving a reduced (dihydro)lipoyl arm.

Substrate/acyl specificity. aceF is an acetyl-specific transferase acting on the acetyl group derived from pyruvate. P. putida KT2440 encodes two other E2 acyltransferase paralogs with distinct acyl specificities — sucB (PP_4188, Q88FB0, EC 2.3.1.61, succinyltransferase of the 2-oxoglutarate dehydrogenase complex) and bkdB (PP_4403, Q88EQ0, branched-chain 2-oxoacid dehydrogenase E2). aceF is only ~33% identical to each of these paralogs, yet 69% identical to a true PDH-E2 ortholog (A. vinelandii E2p). This roughly two-fold difference confirms that aceF's acetyl specificity is orthology-defined and encoded in its divergent catalytic domain, functionally separating the pyruvate→acetyl-CoA node from the TCA-cycle 2-oxoglutarate step (sucB) and branched-chain amino-acid catabolism (bkdB).

Mechanistic role of the modular design (substrate channelling):
1. A lipoyl domain presents its lipoyl-lysine to the E1 active site, where it is reductively acetylated (S8-acetyldihydrolipoamide).
2. The flexible Ala/Pro linkers swing the acetylated arm into the E2 catalytic channel. Structural work on the near-identical Azotobacter vinelandii E2p core shows a ~29 Å active-site channel in which CoA enters from the inside of the cube and the lipoamide arm enters from the outside, so the two substrates meet buried within the trimer interface (PMID 1549782).
3. Acetyl transfer produces acetyl-CoA; the resulting dihydrolipoyl arm is then presented to E3, which reoxidizes it (regenerating oxidized lipoamide and reducing NAD+ via FAD).

This "swinging-arm" coupling channels reactive intermediates between three spatially separated active sites without releasing them to bulk solvent. Cryo-EM of the human complex shows that CoA binding modulates the conformational landscape of the lipoyl domains, indicating the arm dynamics are actively coupled to substrate occupancy (PMID 31130485).

Catalytic residues (experimentally defined in the homolog; conserved in aceF): Site-directed mutagenesis with crystallography of A. vinelandii E2p (PMID 7703242) established the active-site chemistry: His610 is the general base for proton transfer (His610→Cys reduced activity ~500-fold), Ser558 provides transition-state stabilization (Ser558→Ala ~200-fold reduction), and Asn614 activates proton transfer. All three are conserved in P. putida aceF at His519 (DHR motif), Ser468 (SSLGH motif), and Asn523. Notably, aceF retains the rare Asn at the 614-equivalent position — a feature described as "exceptional" in A. vinelandii (most E2 homologs have Asp there) — reflecting their close Pseudomonadales kinship and validating A. vinelandii E2p as the structural surrogate for aceF.


4. Structural / Scaffolding Role

E2 is described structurally and functionally as the central enzyme of PDHc. Key evidence, from high-resolution structures of the close Pseudomonadales relative A. vinelandii E2p:

  • The catalytic domain forms a 24-subunit oligomer with octahedral 432 symmetry (PMID 8487300); UniProt independently annotates Q88QZ6 as forming "a 24-polypeptide structural core with octahedral symmetry."
  • Eight tightly-associated trimers assemble into a hollow truncated cube (~120–125 Å edge) with pores on each face, "forming the core of the multienzyme complex" (PMID 1549782). The trimer is the true building block; two levels of contacts (3-fold trimer + 2-fold trimer–trimer) build the cube.
  • Each catalytic subunit has a topology identical to chloramphenicol acetyltransferase (CAT), the structural basis for its acyl-transfer chemistry (PMID 1549782).

The PSBD provides the attachment platform: in the E. coli PDHc (the best-studied Gram-negative model), point substitutions in the PSBD (R129E, R150E) severely reduce complex activity and disrupt binding of both E1 and E3 as well as reductive acetylation of E2 (PMID 23580650). Thus E2 both builds the core and recruits the peripheral catalytic subunits, giving PDHc its megadalton multienzyme architecture.

Quantitative homology (validation of the structural inference): a full-length global (Needleman–Wunsch) alignment of aceF (546 aa) gives 69.2% identity (444/642) to A. vinelandii E2p (P10802, the crystallized/mutationally-dissected model) and 49.6% (316/637) to E. coli AceF (P06959). This exceptionally high identity to a same-order (Pseudomonadales) enzyme means its solved structures and mechanism transfer to aceF with high confidence; length differences arise mainly from lipoyl-domain copy number and Ala/Pro linker length rather than the conserved catalytic core.

Complex partners in KT2440: aceF (E2, PP_0338, Q88QZ6) assembles with E1 = aceE (PP_0339, Q88QZ5, 881 aa) and the shared E3 = lpdG (PP_4187, Q88FB1, dihydrolipoyl dehydrogenase, 478 aa). The adjacent loci PP_0338/PP_0339 are consistent with a co-transcribed aceEF operon, while the distal E3 (lpdG) is typical of a dihydrolipoyl dehydrogenase shared among the 2-oxoacid dehydrogenases and glycine-cleavage system.


5. Localization

  • Cytoplasm (GO:0005737, and as a component of the cytoplasmic PDH complex, GO:0045254). As a bacterial enzyme, it operates in the cytosol — there is no mitochondrion; this is the prokaryotic counterpart of the mitochondrial-matrix mammalian PDHc-E2.

6. Pathway Context and Physiological Role

  • Metabolic node: PDHc performs the irreversible oxidative decarboxylation of pyruvate → acetyl-CoA + CO2 + NADH, the principal gateway from central sugar catabolism into the TCA cycle and into acetyl-CoA-dependent pathways (fatty-acid synthesis, PHA/polyhydroxyalkanoate biosynthesis, acetylation).
  • In P. putida KT2440 specifically: glucose is catabolized predominantly through the Entner–Doudoroff pathway (KT2440 lacks a functional Embden–Meyerhof–Parnas glycolysis for net catabolism), with the cyclic EDEMP arrangement providing NADPH; carbon converges on pyruvate, and PDHc (E1 aceE/PP_0339 – E2 aceF/PP_0338 – E3 lpd) feeds acetyl-CoA to the TCA cycle. The adjacent PP_0338/PP_0339 loci are consistent with an ace gene cluster encoding the E1 and E2 components.
  • Family: 2-oxoacid dehydrogenase family. The E2 module is the paradigm shared with the 2-oxoglutarate (E2o, succinyltransferase EC 2.3.1.61) and branched-chain 2-oxoacid dehydrogenase complexes; aceF is the pyruvate-specific (acetyltransferase) member.

7. Evidence Summary

Claim Evidence type Source
EC 2.3.1.12; acetyl-transfer reaction; (R)-lipoate cofactor Curated annotation (RuleBase/UniRule) UniProt Q88QZ6
Two lipoyl domains + PSBD + catalytic domain Sequence/domain features UniProt Q88QZ6; InterPro IPR000089, IPR003016, IPR001078, IPR006256
24-mer octahedral cubic core; CAT fold; 29 Å active-site channel X-ray crystallography of Pseudomonadales homolog A. vinelandii E2p (2.6 Å) PMID 1549782; PMID 8487300
Catalytic His610/Ser558/Asn614 (→ aceF His519/Ser468/Asn523) Site-directed mutagenesis + crystallography (A. vinelandii) PMID 7703242
Lipoyl domains require LipB/LipA (de novo) or LplA (salvage) lipoylation Biochemistry / proteomics (E. coli) PMID 21209092
Lipoyl "swinging arm"; covalent lipoyl-lysine for active-site coupling Biochemistry / MS mapping PMID 21798751
PSBD tethers E1/E3 and enables reductive acetylation Mutagenesis + structural MS (E. coli) PMID 23580650
CoA-modulated lipoyl-domain dynamics / channelling Cryo-EM + native MS (human PDHc) PMID 31130485
Cytoplasmic localization; PDHc membership Curated GO UniProt Q88QZ6 (GO:0005737, GO:0045254, GO:0004742, GO:0006086)
69.2% identity to A. vinelandii E2p; 49.6% to E. coli AceF Global sequence alignment (this work) UniProt P10802, P06959
Partners aceE (PP_0339/Q88QZ5) & lpdG (PP_4187/Q88FB1); aceEF operon Genomic loci / UniProt UniProt Q88QZ5, Q88FB1
Acetyl-specific: only ~33% identity to paralogs sucB (succinyl) & bkdB Comparative alignment (this work) UniProt Q88FB0, Q88EQ0

Strength of inference: No P. putida-specific enzymological study of aceF was located; however, the function is established at high confidence by (i) unambiguous, mutually-consistent UniProt/InterPro annotation, and (ii) direct structural/biochemical characterization of very close bacterial homologs (A. vinelandii, same order; E. coli, same class), which are the standard models for this enzyme family.


8. Supported vs. Refuted Hypotheses

Supported:
- H1 — aceF is the E2 acetyltransferase (EC 2.3.1.12) of PDHc catalyzing acetyl-CoA formation. ✔
- H2 — aceF forms the 24-mer cubic core and scaffolds the complex via its PSBD. ✔
- H3 — Function depends on covalent lipoyl "swinging arms" enabling substrate channelling. ✔
- H4 — The enzyme acts in the cytoplasm at the glycolysis/ED–TCA junction. ✔
- H5 — aceF is acetyl-specific, distinct from the succinyl (sucB) and branched-chain (bkdB) E2 paralogs (~33% identity vs 69% to a PDH-E2 ortholog). ✔
- H6 — Catalytic His519/Ser468/Asn523 are conserved from the mutationally-validated A. vinelandii active site. ✔ (inferred)

Refuted / not applicable:
- The gene-symbol-ambiguity contingency was ruled out: aceF unambiguously matches the annotated PDHc-E2 identity.


9. Limitations and Future Directions

  • Direct experimental characterization is from homologs, not P. putida KT2440 itself; a KT2440-specific structure or kinetic study would confirm exact catalytic-residue geometry and the functional consequence of the two-lipoyl-domain arrangement.
  • The precise lipoyl-lysine positions and the catalytic His were inferred from sequence motifs and homology; site-directed mutagenesis in KT2440 would confirm them.
  • Whether both lipoyl domains are equally lipoylated/functionally redundant in KT2440 is untested.
  • Regulatory features (e.g., any PDH-kinase control) and operon structure in KT2440 warrant genomic/proteomic confirmation.

Report generated during autonomous functional-annotation investigation (Iterations 1–5). Conclusions rest on curated UniProt/InterPro annotation, high-resolution structures and site-directed mutagenesis of close Pseudomonadales/Gammaproteobacteria homologs (A. vinelandii E2p, E. coli AceF), conserved-residue and quantitative sequence-identity analysis, and KT2440 genomic context.

Artifacts

📄 View Raw YAML

id: Q88QZ6
gene_symbol: aceF
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: aceF (PP_0338) encodes the E2 acetyltransferase component of the pyruvate dehydrogenase complex. The protein carries lipoyl domains and a catalytic acyltransferase domain that transfers the acetyl group from S-acetyldihydrolipoyllysine to coenzyme A, coupling the E1 pyruvate-decarboxylating reaction to acetyl-CoA formation. It is a cytoplasmic component of the multienzyme pyruvate dehydrogenase complex, connecting pyruvate produced by lower central carbon metabolism to acetyl-CoA and the TCA cycle.
existing_annotations:
- term:
    id: GO:0004742
    label: dihydrolipoyllysine-residue acetyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: dihydrolipoyllysine-residue acetyltransferase activity is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
    action: ACCEPT
    reason: This is a specific, biologically appropriate annotation for this gene product.
- term:
    id: GO:0005737
    label: cytoplasm
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: cytoplasm is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
    action: ACCEPT
    reason: This is a specific, biologically appropriate annotation for this gene product.
- term:
    id: GO:0006086
    label: pyruvate decarboxylation to acetyl-CoA
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: pyruvate decarboxylation to acetyl-CoA is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
    action: ACCEPT
    reason: This is a specific, biologically appropriate annotation for this gene product.
- term:
    id: GO:0016407
    label: acetyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: enables
  review:
    summary: acetyltransferase activity is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
    action: KEEP_AS_NON_CORE
    reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
- term:
    id: GO:0016746
    label: acyltransferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: acyltransferase activity is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
    action: KEEP_AS_NON_CORE
    reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
- term:
    id: GO:0031405
    label: lipoic acid binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: enables
  review:
    summary: lipoic acid binding is biologically plausible for this enzyme but is ancillary to the more specific catalytic function.
    action: KEEP_AS_NON_CORE
    reason: Retain as a supporting/non-core annotation rather than using it as the main functional summary.
- term:
    id: GO:0045254
    label: pyruvate dehydrogenase complex
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: part_of
  review:
    summary: pyruvate dehydrogenase complex is consistent with the curated UniProt name, EC/family evidence, and the gene product role summarized here.
    action: ACCEPT
    reason: This is a specific, biologically appropriate annotation for this gene product.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: file:PSEPK/aceF/aceF-uniprot.txt
  title: UniProt record for aceF (Q88QZ6)
  findings:
  - statement: UniProt identifies aceF as Acetyltransferase component of pyruvate dehydrogenase complex (EC 2.3.1.12) and provides the seeded EC/domain/GO evidence reviewed here.
- id: file:PSEPK/aceF/aceF-deep-research-asta.md
  title: Asta deep-research retrieval for aceF
  findings:
  - statement: Asta retrieval was run for this first-pass pathway curation; direct organism-specific literature was limited for several common enzyme names, so UniProt/family evidence carries the main review weight.
aliases:
- PP_0338
core_functions:
- description: dihydrolipoyllysine-residue acetyltransferase activity supporting the Acetyltransferase component of pyruvate dehydrogenase complex (EC 2.3.1.12) role summarized for aceF.
  supported_by:
  - reference_id: file:PSEPK/aceF/aceF-uniprot.txt
    supporting_text: DR   GO; GO:0004742; F:dihydrolipoyllysine-residue acetyltransferase activity; IEA:UniProtKB-UniRule.
  molecular_function:
    id: GO:0004742
    label: dihydrolipoyllysine-residue acetyltransferase activity
  directly_involved_in:
  - id: GO:0006086
    label: pyruvate decarboxylation to acetyl-CoA
proposed_new_terms: []