hisC

UniProt ID: Q88P86
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

Histidinol-phosphate aminotransferase (HisC; EC 2.6.1.9), a cytoplasmic pyridoxal-5'-phosphate (PLP)-dependent class-II aminotransferase that catalyzes the seventh step of L-histidine biosynthesis. It transfers an amino group from L-glutamate to imidazole-acetol phosphate (3-(imidazol-4-yl)-2-oxopropyl phosphate), producing L-histidinol phosphate and 2-oxoglutarate. The enzyme functions as a homodimer with active sites at the dimer interface; PLP is covalently bound as an internal aldimine to an active-site lysine (Lys210 in this protein) and catalysis proceeds via a ping-pong mechanism through a pyridoxamine-5'-phosphate intermediate. In Pseudomonas putida KT2440 the gene (PP_0967) lies within a histidine-biosynthesis gene cluster. Aromatic-amino-acid transamination is documented for some HisC orthologs but has not been demonstrated for the KT2440 protein.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0000105 L-histidine biosynthetic process
IEA
GO_REF:0000120
ACCEPT
Summary: HisC catalyzes the seventh step of histidine biosynthesis; this BP term is well supported by family/HAMAP-rule assignment, the UniProt pathway annotation, and operon context in KT2440.
Reason: Core biological process for this enzyme. The histidinol-phosphate aminotransferase function places it squarely in the L-histidine biosynthetic pathway (UniPathway UPA00031; HAMAP-Rule MF_01023).
Supporting Evidence:
file:PSEPK/hisC/hisC-deep-research-openscientist.md
catalyzes the **seventh step of de novo L-histidine biosynthesis**
GO:0004400 L-histidinol-phosphate:2-oxoglutarate transaminase activity
IEA
GO_REF:0000120
ACCEPT
Summary: This is the specific molecular function of HisC (EC 2.6.1.9), transaminating L-histidinol phosphate with 2-oxoglutarate/L-glutamate. The UniProt CATALYTIC ACTIVITY block and HAMAP rule directly support this.
Reason: Represents the core molecular function. Strongly supported by family assignment (HisP_aminotrans subfamily, TIGR01141 hisC), Rhea:23744, and EC 2.6.1.9.
Supporting Evidence:
file:PSEPK/hisC/hisC-deep-research-openscientist.md
*hisC* encodes **histidinol-phosphate aminotransferase (HisC, EC 2.6.1.9)**
GO:0016740 transferase activity
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: A high-level parent of the specific aminotransferase activity already annotated (GO:0004400). It is correct but uninformative given the more precise term.
Reason: Redundant generic ancestor of GO:0004400; adds no information beyond the specific transaminase MF term.
GO:0030170 pyridoxal phosphate binding
IEA
GO_REF:0000002
KEEP AS NON CORE
Summary: HisC is a PLP-dependent enzyme that covalently binds pyridoxal 5'-phosphate as an internal aldimine at the active-site lysine (MOD_RES 210 in this entry). Well supported by the COFACTOR annotation and conserved PLP-lysine motif.
Reason: PLP binding is an essential, well-supported cofactor interaction, but the substrate-specific transaminase activity is the defining molecular function. The PLP-lysine internal aldimine and ping-pong mechanism are documented for HisC homologs (see hisC-deep-research-falcon.md).
GO:0140385 amino acid transaminase activity
IEA
GO_REF:0000117
MARK AS OVER ANNOTATED
Summary: A broad parent term covering aminotransferase activity on amino acid substrates. Correct but less specific than GO:0004400, which is already annotated.
Reason: Generic ancestor of the specific histidinol-phosphate transaminase activity; the more precise term GO:0004400 already captures this function.

Core Functions

Catalyzes the PLP-dependent transamination of imidazole-acetol phosphate to L-histidinol phosphate (using L-glutamate as amino donor), the seventh step of L-histidine biosynthesis.

Supporting Evidence:
  • file:PSEPK/hisC/hisC-deep-research-falcon.md
    KT2440 histidine-biosynthesis genes PP0965-PP0967 are annotated as the hisGDC cluster, placing PP_0967 as hisC within histidine biosynthesis.
  • file:PSEPK/hisC/hisC-deep-research-openscientist.md
    transfers the α-amino group of **L-glutamate onto imidazole-acetol phosphate (3-(imidazol-4-yl)-2-oxopropyl phosphate)** to yield **L-histidinol phosphate + 2-oxoglutarate**

References

Gene Ontology annotation through association of InterPro records with GO terms
Electronic Gene Ontology annotations created by ARBA machine learning models
Combined Automated Annotation using Multiple IEA Methods
Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium through the screening of a genome-wide mutant library
  • In P. putida KT2440 the histidine-biosynthesis genes PP0965-PP0967 are annotated as the hisGDC cluster, with PP_0967 corresponding to hisC, and RT-PCR co-transcription assays support operon organization of these clusters.
Crystal structure of histidinol phosphate aminotransferase (HisC) from Escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and L-histidinol phosphate
  • E. coli HisC is a PLP-dependent homodimeric aminotransferase with active sites at the dimer interface; the active-site lysine (Lys214) forms an internal aldimine with PLP and catalysis proceeds through PMP via a ping-pong mechanism.
file:PSEPK/hisC/hisC-deep-research-openscientist.md
OpenScientist functional report for PSEPK HisC
  • Supports the exact HisC reaction and histidine-pathway assignment for Q88P86 while distinguishing family-level aromatic-amino-acid activity from a demonstrated physiological role in KT2440.

Deep Research

Asta

(hisC-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:19:18.686850

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.705)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[2] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.703)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[3] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.700)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[4] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.689)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[5] Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana

  • Authors: K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al.
  • Year: 2020
  • Venue: Journal of Food Science and Engineering
  • URL: https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  • DOI: 10.17265/2159-5828/2020.02.001
  • Summary: Baccaurea motleyana (rambai) is underutilized fruits that are native to Malaysia, Indonesia and Thailand and used for simple sequence repeat (SSR) analysis by MIcroSAtellite (MISA).
  • Evidence snippets:
  • Snippet 1 (score: 0.686)
    > To get comprehensive gene function of rambai genes, gene annotation to seven databases, namely National Center for Biotechnology Information (NCBI) non-redundant protein sequences (NR), NCBI nucleotide sequences (NT), Kyoto Encyclopedia of Genes and Genome Ortholog (KO), SwissProt, Protein family (Pfam), Gene Ontology (GO) and Cluster of Orthologous Groups (KOG), was used as reference.
    > The NCBI non-redundant protein sequences (NR), include protein sequence information from GenBank, Protein Data Bank (PDB), SwissProt, Protein Information Resource (PIR) and Protein Research Foundation (PRF). The NCBI nucleotide sequences (NT) are the nucleotide sequence database that includes nucleotide sequence from GenBank of the European Bioinformatics Institute (EMBL) and DNA Data Bank of Japan (DDBJ). KEGG is a database resource for understanding high-level functions and utilities of the biological system, such as cell, organism and ecosystem, from molecular-level information, especially for large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies. KEGG is an established Cluster of Orthologous (KO) annotation system that can accomplish the function annotation of the genome/transcriptome of a newly sequenced species. SwissProt is a manual annotated and reviewed protein sequence database that has a high-quality protein sequence database from experimental results, computed features and scientific conclusions. Pfam is comprehensive collection of protein domains and families, represented as multiple sequence alignments and as profile of hidden Markov models. Many proteins are composed of structural domains, and the protein sequence of a specific structural domain possesses a certain degree of conservative property. GO is the established standard for the functional annotation of gene products and controlled vocabulary used to classify the functional attributes of gene products of a biological process, a molecular function and a cellular component.

[6] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.684)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[7] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.678)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[8] Quantitative proteomic dataset of whole protein in three melanoma samples of 92.1, 92.1-A and 92.1-B

  • Authors: Xi-feng Fei, Xiangtong Xie, X. Ji, Haiyan Tian, F. Sun et al.
  • Year: 2022
  • Venue: Data in Brief
  • URL: https://www.semanticscholar.org/paper/33312bb6cc2cf985d32f3a31cf4fdee6b4e17385
  • DOI: 10.1016/j.dib.2022.108592
  • PMID: 36164296
  • PMCID: 9508510
  • Citations: 2
  • Influential citations: 1
  • Summary: Covering differential proteomes of three cell lines in a pairwise model, the data could be used to further screen the kinesins that play a vital role in regulating the growth of UM.
  • Evidence snippets:
  • Snippet 1 (score: 0.676)
    > 2.8.1. Annotation methods 2.8.1.1. Functional annotation. UniProt-GOA database was utilized to retrieve Gene Ontology (GO) annotation proteome. First, the identified protein identity was converted to UniProt identity and then mapped to GO identity based on the protein identity. When the identified protein was not annotated by UniProt-GOA, the functional annotation of that protein was conducted using the InterProScan software according to the amino acid sequence alignment approach. All proteins were then classified into 3 groups: molecular function, cellular component and biological process.

[9] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.671)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[10] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.667)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[11] Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells

  • Authors: Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al.
  • Year: 2011
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  • DOI: 10.1371/journal.pone.0020489
  • PMID: 21647379
  • PMCID: 3103581
  • Citations: 51
  • Influential citations: 3
  • Summary: This work reports the first proteomic study of the mitotic spindle from Chinese Hamster Ovary (CHO) cells and identifies proteins that are unique to the CHO spindle.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > The lists of proteins used for the comparison contain more items than listed previously due to expansion out of gene clusters, for example, to allow the updating and comparison of current HGNC gene symbols. These lists of proteins were compared in Microsoft Excel 2011 using PivotTable.
    > The protein set for the CHO midbody was derived from the accession numbers in Table S1 and Table S2 from Skop et al. [9]. Original accession numbers were updated to more recent UniProt accessions (2/2010), and duplicates from different species or different protein isoforms were removed. The unique UniProt accessions were mapped to gene names using UniProt KB Unimart, UniProt dataset [82], and these gene names were confirmed manually as HGNC symbols using HGNC, with ambiguities checked using BLASTP of the sequence corresponding to the original accession number. Accessions that didn't map successfully in Unimart were manually analyzed using BLASTP against the human RefSeq protein set using sequences from the original accessions, combined with TreeFam.org data for the non-human UniProt accessions. HGNC symbols were updated again on 12/14/2010 before comparison with this paper's protein set.
    > The protein set for the HeLa spindle proteome is derived from the 1121 accession numbers in Sauer et al. supplementary table 1 column 2 [17]. Updating the 1116 UniProt accession numbers and 5 IPI accession numbers from 795 rows required several steps. Most proteins were updated to current UniProt accessions using UniProt retrieve. Duplicates were removed. Sequences for the IPI accession numbers and 16 defunct UniProt accession numbers were recovered from other sources on the web, and BLASTP against the human RefSeq protein set with a cutoff of at least 90% identity was used to update some of these accessions. The unique current UniProt accessions were mapped to HGNC symbols using UniProt ID mapping to HGNC IDs. Biomart, database Ensembl Genes 60, dataset GRCh37.p2 [http://uswest.ensembl.org/biomart/martview/]

[12] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.665)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[13] Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates

  • Authors: Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa
  • Year: 2023
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  • DOI: 10.1101/2023.09.20.558425
  • Citations: 1
  • Summary: It is found that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration, and protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within OFPs, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies.
  • Evidence snippets:
  • Snippet 1 (score: 0.665)
    > We defined the subcellular localization for each protein in the human proteome by integrating data from Gene Ontology annotations in UniProt (GOA), UniProt annotations, the Human Transmembrane Proteome (HTP) 121 , MatrixDB 122 , and MatrisomeDB 123 . We divided the UniProt and the Gene Ontology annotations (GOA) into tier 1 (more reliable) and tier 2 (less reliable) annotations, depending on the attached evidence codes. For UniProt, annotations with the evidence codes ECO:0000269 or ECO:0000305 are considered as tier 1, while annotations with evidence codes ECO:0000250, ECO:0000255, or ECO:0000303 are tier 2. For Gene Ontology, annotations with evidence codes IDA, IMP, IPI, IGI, EXP, IBA, IKR, TAS, NAS, IC, or ND are tier 1, while annotations with evidence codes HDA, ISS, ISA, RCA, ISO, ISM, IGC, or IEA are tier 2. Based on these, each protein was assigned exactly one broad localization. It was considered to be a transmembrane protein (TMP), if it is assigned the 'integral component of membrane (GO:0016021)' GO term in tier 1 GOA annotations, or it is annotated as a TMP in HTP with a confidence score over 85, or is annotated in HTP as a TMP with a confidence score above 50 and is also annotated as a TMP in GOA (either tier).

[14] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.658)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[15] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.656)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[16] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.649)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[17] Functional annotation of parasitic worm genomes, by assigning protein names and GO terms

  • Authors: Avril Coghlan, M. Berriman
  • Year: 2018
  • Venue: Unknown venue
  • URL: https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  • DOI: 10.1038/protex.2018.055
  • Citations: 1
  • Summary: A computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and Go terms from orthologues in other species.
  • Evidence snippets:
  • Snippet 1 (score: 0.643)
    > Given a set of predicted protein-coding genes for a newly sequenced genome, functional annotation involves assigning putative functions to the predicted genes. Two ways in which this can be done are assigning protein names and Gene Ontology (GO;Gene Ontology Consortium, 2010) terms to the predicted proteins. Here we describe a computational pipeline for assigning protein names and GO terms to predicted proteins in parasitic worm (nematode and platyhelminth) genomes, which transfers names and GO terms from orthologues in other species.
    > When assigning protein names, UniProt protein naming rules (www.uniprot.org/docs/nameprot) are followed where possible. This recommends that a good and stable name for a protein is "as neutral as possible"; that a protein name "should be, as far as possible, unique and attributed to all orthologs"; and a protein name "should not contain a specific characteristic of the protein, and in particular it should not reflect the function or role of the protein, nor its subcellular location, its domain structure, its tissue specificity, its molecular weight or its species of origin".
    > In our protocol, a protein name is assigned to each predicted protein based on curated names in UniProt (Bairoch & Apweiler, 2000) for human, zebrafish, Drosophila melanogaster, Caenorhabditis elegans, and Schistosoma mansoni orthologues identified from a database of gene families (e.g. built using Ensembl Compara; Vilella et al. 2009), or (if no information is found from orthologues) based on InterPro (Hunter et al. 2012) domains. Figure 1 shows an example of using our protein naming pipeline for four Strongyloides ratti genes that belong to the tubulin polyglutamylase family (underlined in pink), where four different protein names were assigned to them (in pink), based on names of their C. elegans or human orthologues.
    > Since each of the S. ratti genes belonged to a different subfamily of the tubulin polyglutamylase family, they were assigned different names.

[18] GOnet: a tool for interactive Gene Ontology analysis

  • Authors: M. Pomaznoy, Brendan Ha, Bjoern Peters
  • Year: 2018
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  • DOI: 10.1186/s12859-018-2533-3
  • PMID: 30526489
  • PMCID: 6286514
  • Citations: 247
  • Influential citations: 17
  • Summary: The open-source GOnet web-application is created, which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis and provides insight into the functional interconnection of the submitted entries.
  • Evidence snippets:
  • Snippet 1 (score: 0.642)
    > In a basic workflow, the GOnet application receives a list of gene symbols, protein symbols, or protein IDs (UniProt IDs) as an input, and outputs a graph (an example given in Fig. 1). There are various input parameters which will affect the actual structure of the graph visualized and its appearance. The first main user choice is which GO terms the genes are annotated against:
    > 1. GO terms statistically significantly over-represented in the gene list submitted. 2. A predefined subset (also known as 'GO slim'), or a user-supplied list of terms.
    > In the first case the analysis will be referred to as an 'enrichment' analysis, in the second as an 'annotation' analysis.
    > Input parameters 1) Gene list. A mandatory input parameter containing the genes/proteins of interest. Currently human and mouse data is supported. An example of a human gene list might look like this:
    > Fig. 1 Sample network output generated by GOnet application. Gene differentially expressed in CD4 Bulk Memory T cells in Latent TB patients compared to healthy controls were used as an example [22] The gene list can also be accompanied with a contrast value. For example, This contrast value can be any decimal number, such as the log-fold change of gene expression between two conditions. This is merely a visualization enhancement. If the value is supplied it can be used later to differentially color specific genes in the graph (note different colors of gene nodes in Fig. 1), and visually indicate up-or down-regulation of specific genes and gene clusters.
    > The application can process common gene symbols (like in the example above), UniProt IDs, and MGI Accession IDs (mouse only). The former type of ID (gene symbols), although is the most human friendly, can unfortunately be ambiguous. For example, AIM1 can mean 'absent in melanoma' (also called CRYBG1) or 'Aurora and Ipl1-like midbody-associated protein' (also known as AURKB). Due to this ambiguity UniProt IDs or MGI accession IDs (for mouse) are preferred.
    > 2) GO namespace. Can be any of 'biological process', 'molecular function' or 'cellular component'.

[19] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.634)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[20] The y-ome defines the thirty-four percent of Escherichia coli genes that lack experimental evidence of function

  • Authors: S. Ghatak, Zachary A. King, Anand V. Sastry, B. Palsson
  • Year: 2018
  • Venue: bioRxiv
  • URL: https://www.semanticscholar.org/paper/b55c92ec2394bb9ebdff9d019f9afb7238e5c8ba
  • DOI: 10.1101/328591
  • Citations: 4
  • Summary: This work identified the genes that lack direct experimental evidence of function (the “y-ome”) and discusses the value of the y-ome for systematic improvement of E. coli knowledge bases and its extension to other organisms.
  • Evidence snippets:
  • Snippet 1 (score: 0.634)
    > Any attempt to systematically assess the function of unannotated genes must therefore draw from multiple knowledge bases and resolve these conflicts.
    > Many research groups have categorized E. coli genes and proteins by annotation quality as a part of their studies. In 2009, Hu et al. constructed a global functional atlas of E. coli proteins (18) . First, they identified all unannotated proteins in the K-12 W3110 and MG1655 genomes. In order for a protein-encoding gene to be considered functionally uncharacterized in their analysis, it had to meet the following criteria: (i) The gene name begins with "y", (ii) the gene does not have a known pathway within EcoCyc, and (iii) the gene does not have a functional description in GenProtEC (19) (any gene with a description containing the words "predicted", "hypothetical", or "conserved"). Based on these criteria, it was determined that 1431 of 4225 protein coding sequences were functionally unannotated. In 2015, Kim et al. published a database called EcoliNet that curated and predicted cofunctional gene networks for every protein coding gene in the E. coli genome (20) . This study also quantified the number of uncharacterized protein coding genes in E. coli . To assess functional annotation, they used the presence of experimentally supported "biological process" annotations in the Gene Ontology database (21) . They concluded that ~2000 protein coding genes in E. coli were functionally unannotated. The most comprehensive effort to assess the level of annotation in bacterial genomes has been Computational Bridges to Experiments (COMBREX) (22,23) . The COMBREX knowledge base currently contains information about 4182 protein coding genes in E. coli K-12 MG1655, of which 2378 (57%) have experimentally verified function, 1741 (42%) have predicted but not experimentally verified function, and 63 (2%) have no predicted function. These studies of unannotated genes in E. coli K-12 MG1655 provided inspiration for this work.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  2. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  3. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  4. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  5. K. Nasir, Muhammad Fairuz Mohd Yusof, M. S. F. A. Razak, Siti Norsaidah Ibrahim, Mira Farzana Mohamad Moktar et al. (2020). Discovery of Simple Sequence Repeat Markers through Transcriptome Analysis of Baccaurea motleyana. Journal of Food Science and Engineering. https://www.semanticscholar.org/paper/f99fe2940881ec45ecbd8ba3da7f10b4fb22fc3b
  6. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  7. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  8. Xi-feng Fei, Xiangtong Xie, X. Ji, Haiyan Tian, F. Sun et al. (2022). Quantitative proteomic dataset of whole protein in three melanoma samples of 92.1, 92.1-A and 92.1-B. Data in Brief. https://www.semanticscholar.org/paper/33312bb6cc2cf985d32f3a31cf4fdee6b4e17385
  9. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  10. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  11. Mary Kate Bonner, D. Poole, Tao Xu, Ali Sarkeshik, J. Yates et al. (2011). Mitotic Spindle Proteomics in Chinese Hamster Ovary Cells. PLoS ONE. https://www.semanticscholar.org/paper/8a46e242e657489c1933c76e06a37618f7d1901f
  12. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  13. Nazanin Farahi, Tamas Lazar, P. Tompa, Bálint Mészáros, Rita Pancsa (2023). Phase-separating fusion proteins drive cancer by dysregulating transcription through ectopic condensates. bioRxiv. https://www.semanticscholar.org/paper/57a63e3228a18d5d68d54eb8303eeb7c0ae29da6
  14. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  15. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  16. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  17. Avril Coghlan, M. Berriman (2018). Functional annotation of parasitic worm genomes, by assigning protein names and GO terms. https://www.semanticscholar.org/paper/583c74a2e225dbff5fca04d36298e5b690491e82
  18. M. Pomaznoy, Brendan Ha, Bjoern Peters (2018). GOnet: a tool for interactive Gene Ontology analysis. BMC Bioinformatics. https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  19. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  20. S. Ghatak, Zachary A. King, Anand V. Sastry, B. Palsson (2018). The y-ome defines the thirty-four percent of Escherichia coli genes that lack experimental evidence of function. bioRxiv. https://www.semanticscholar.org/paper/b55c92ec2394bb9ebdff9d019f9afb7238e5c8ba

Falcon

(hisC-deep-research-falcon.md)
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate. Falcon Edison Scientific Literature 27 citations 2 artifacts 2026-06-11T22:01:27.871031

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: hisC (PP_0967; UniProt Q88P86) in Pseudomonas putida KT2440

Executive summary

The Pseudomonas putida KT2440 gene hisC (annotated as PP_0967) encodes histidinol-phosphate aminotransferase (also called imidazole acetol-phosphate transaminase), an enzyme of histidine biosynthesis that catalyzes a pyridoxal-5′-phosphate (PLP)-dependent transamination step (EC 2.6.1.9). In KT2440, hisC is part of a histidine-biosynthesis gene cluster annotated hisGDC (PP_0965–PP_0967) and is supported as an operon by RT-PCR co-transcription assays (data referenced but not shown). Although KT2440 gene-level quantitative fitness for hisC was not found in the retrieved texts, histidine biosynthesis is conditionally essential for minimal-medium growth, and multiple other his genes yield histidine auxotrophy in a genome-wide mutant screen. The enzyme is expected to be cytosolic, consistent with its role in core amino-acid biosynthesis.

Mandatory target verification (avoid symbol ambiguity)

Target identity required by the prompt: UniProt Q88P86, gene hisC, locus PP_0967, organism P. putida KT2440.

Strain-specific verification from KT2440 literature: A genome-wide KT2440 study explicitly lists histidine-biosynthesis genes as organized into four clusters and identifies PP0965–PP0967 as the “hisGDC” cluster, placing PP_0967 as hisC within that cluster. The same source reports co-transcription assays indicating the clusters form independent operons, supporting operon organization for this region (molina‐henares2010identificationofconditionally pages 7-9). In parallel, an authoritative histidine-biosynthesis review defines HisC as histidinol aminotransferase with EC 2.6.1.9 (winkler2009biosynthesisofhistidine pages 46-47).

Limitations of verification: The tools in this run did not directly retrieve the UniProt record for Q88P86, so the UniProt accession-to-locus mapping is indirect (PP_0967 ↔ hisC) rather than confirmed by UniProt text in context.

1) Key concepts, definitions, and current understanding

1.1 Histidine biosynthesis context

Histidine is synthesized in bacteria via a conserved multi-step pathway; HisC performs a late aminotransferase step (often described as the 7th step in bacteria) in which an amino group is installed on the histidine precursor (sivaraman2001crystalstructureof pages 1-2, winkler2009biosynthesisofhistidine pages 12-13).

1.2 Enzymatic function (reaction, EC number, substrates/products)

Primary biochemical role of HisC (EC 2.6.1.9):
- Amino donor: typically L-glutamate
- Amino acceptor: imidazole acetol-phosphate (also described as 3-(imidazol-4-yl)-2-oxo-propyl phosphate, or imidazoleacetol-phosphate)
- Products: L-histidinol phosphate + 2-oxoglutarate (α-ketoglutarate)

This reaction is explicitly described in structural/enzymology studies and reviews (sivaraman2001crystalstructureof pages 1-2, matte2003contributionofstructural pages 2-3, fernandez2004structuralstudiesof pages 1-2).

1.3 Cofactor dependence and enzyme class

HisC is a PLP-dependent aminotransferase. Structural work shows the canonical PLP chemistry: PLP is covalently linked to an active-site lysine as an internal aldimine, cycles through pyridoxamine-5′-phosphate (PMP), and catalysis proceeds via a ping-pong (double-displacement) mechanism characteristic of aminotransferases (sivaraman2001crystalstructureof pages 1-2, sivaraman2001crystalstructureof pages 7-9, winkler2009biosynthesisofhistidine pages 12-13).

1.4 Mechanism and structural determinants (expert-level structural biology)

High-resolution crystallography on bacterial HisC (not KT2440-specific) provides mechanistic anchors useful for functional annotation:
- Oligomerization: HisC is dimeric, and the active sites lie at the dimer interface (sivaraman2001crystalstructureof pages 1-2, sivaraman2001crystalstructureof pages 7-9, sivaraman2001crystalstructureof pages 2-4).
- Active-site lysine: in E. coli HisC, Lys214 forms the internal aldimine with PLP (sivaraman2001crystalstructureof pages 13-14).
- Captured intermediates: structures include PLP internal aldimine, PMP state, and an unusual covalent tetrahedral/gem-diamine–like intermediate involving PLP + L-histidinol phosphate + active-site Lys, supporting the transimination mechanism (sivaraman2001crystalstructureof pages 1-2, matte2003contributionofstructural pages 2-3, winkler2009biosynthesisofhistidine pages 12-13).
- Conserved PLP-contact residues: residues such as Tyr55, Asn157, Asp184, Tyr187, Ser213, Lys214, Arg222 (numbering from E. coli) are described as conserved PLP-interacting positions (sivaraman2001crystalstructureof pages 1-2, matte2003contributionofstructural pages 2-3).

Image evidence: Sivaraman et al. provide figures schematizing (i) the covalent PLP–L-histidinol phosphate complex interactions and (ii) the transimination mechanism states (internal aldimine → gem-diamine intermediates → external aldimine), which directly support the mechanistic model (sivaraman2001crystalstructureof media 1b7db48c, sivaraman2001crystalstructureof media 8b8302fc).

2) KT2440-specific biology: gene context, pathway integration, phenotypes, localization

2.1 Genomic organization and operon context (KT2440)

In P. putida KT2440, histidine-biosynthesis genes are described as distributed in four genomic clusters, including PP0965–PP0967 (“hisGDC”). Co-transcription assays by RT-PCR are reported to show that the clusters form independent operons, supporting that PP_0965–PP_0967 are co-transcribed (molina‐henares2010identificationofconditionally pages 7-9).

2.2 Functional genetics: minimal-medium conditional essentiality and auxotrophy

A genome-wide KT2440 mutant-library screen on glucose minimal medium provides quantitative and phenotype-level evidence that histidine biosynthesis is crucial under nutrient limitation:
- Library size: 7,760 independent clones screened.
- Minimal-medium growth defects: 79 mutants unable to grow on glucose minimal medium.
- Unique genes implicated: 47 independent knockout genes mapped from those mutants.
- Histidine auxotroph-associated hits recovered include hisB (PP0289; 1 hit), hisF (PP0293; 2 hits), hisH (PP0290; 2 hits), and hisZ (PP4890; 1 hit) (molina‐henares2010identificationofconditionally pages 2-3).

Notably, hisC was not recovered as a mutant hit in that screen, despite being in a cluster predicted by in silico models to yield histidine auxotrophy (molina‐henares2010identificationofconditionally pages 7-9, molina‐henares2010identificationofconditionally pages 2-3). This is consistent with the broader point made by the authors that transposon mutagenesis screens can miss some predicted conditionally essential loci due to library coverage and gene organization effects (molina‐henares2010identificationofconditionally pages 2-3).

2.3 Cellular localization (KT2440)

No retrieved KT2440 paper provided an explicit subcellular localization statement for HisC. Given HisC’s role in core amino-acid biosynthesis and the absence of any membrane/periplasmic context in the KT2440 evidence presented here, the most defensible statement from the present evidence base is that HisC functions in the intracellular (cytosolic) metabolic network that supplies histidine for translation and metabolism (sivaraman2001crystalstructureof pages 1-2, molina‐henares2010identificationofconditionally pages 2-3).

3) Quantitative biochemical data relevant to functional annotation

Direct biochemical kinetics for KT2440 HisC were not retrieved in this run. However, quantitative parameters from well-studied bacterial homologs provide a calibrated expectation for activity and specificity (with appropriate caution about species differences).

3.1 Kinetic constants and specificity (Thermotoga maritima HisC homolog)

A hyperthermophilic HisC (tmHspAT) study reported kinetic constants (measured at 20°C) for multiple substrates:
- Histidinol phosphate (Hsp): Km 0.8 mM; kcat 2.8 min⁻¹; kcat/Km 3.5 min⁻¹·mM⁻¹
- Tyrosine: Km 2.3 mM; kcat 2.6 min⁻¹; kcat/Km 1.13
- Tryptophan: Km 3.4 mM; kcat 0.85 min⁻¹; kcat/Km 0.25
- Phenylalanine: Km 38.0 mM; kcat 0.52 min⁻¹; kcat/Km 0.014

The same study reports no measurable activity with L-histidine and notes temperature dependence with maximal activity above 60°C (fernandez2004structuralstudiesof pages 9-10).

These data illustrate a key annotation nuance: some HisC homologs can display broadened substrate ranges (e.g., aromatic amino acids), while still functioning in histidine biosynthesis (fernandez2004structuralstudiesof pages 1-2, fernandez2004structuralstudiesof pages 9-10).

3.2 Spectral signatures of PLP/PMP states (E. coli HisC)

UV–visible spectroscopy provides quantitative cofactor-state signatures:
- A peak around 327 nm (assigned to PMP form)
- Conversion to PLP internal aldimine yields peaks at 338 nm and 427 nm; addition of α-ketoglutarate drives this conversion, with an observed shift above 15 mM α-ketoglutarate (sivaraman2001crystalstructureof pages 7-9).

3.3 Structural/biophysical quantitative descriptors (E. coli HisC)

Reported measurements include:
- Dimer in solution, with dynamic light scattering Mr ≈ 60 kDa and monomer ≈ 40 kDa
- Dimer dimensions ≈ 94 × 55 × 54 Å
- PLP–PLP phosphate distance across the dimer: 22.6 Å
- Soaking concentration for L-histidinol phosphate in crystallography: 4 mM
- A PLP ring rotation ~20° and Lys movement ~1 Å upon covalent complex formation (sivaraman2001crystalstructureof pages 2-4, sivaraman2001crystalstructureof pages 9-10).

4) Recent developments (prioritizing 2023–2024) and current applications

4.1 Recent advances in KT2440 functional annotation workflows (2024)

A 2024 mSystems paper demonstrates use of independent component analysis (ICA) on a large RB-TnSeq fitness compendium to identify “functional modules” (fModules) in P. putida KT2440 and links these to regulatory iModulons. The retrieved excerpt specifically notes histidine-related signals (e.g., hisA in a histidine/purine-related module and a “His metabolism” connection via HutC) as part of this modern data-driven annotation strategy (borchert2024machinelearninganalysis pages 11-13). While this excerpt does not provide hisC-specific values, it reflects a current trend: integrating high-throughput fitness and transcriptomics with machine learning to refine gene-function relationships.

4.2 Real-world implementation: KT2440 as a biotechnological chassis (2023)

A 2023 Science Advances study reports metabolic engineering and bioprocess development of P. putida KT2440 for lignin-related aromatic conversion to β-ketoadipic acid, achieving titers of 44.5 g/L (model LRCs) and 25 g/L (corn stover-derived LRCs), and predicted a minimum selling price of $2.01/kg (Werner et al., 2023; URL in retrieved metadata). This positions KT2440 as an industrially relevant chassis; although the retrieved text segments did not connect this directly to histidine biosynthesis, such chassis optimization depends on robust central metabolism including amino acid supply and PLP-dependent enzyme networks (paper metadata retrieved; no direct in-text hisC evidence found here).

4.3 Broader (non-KT2440) translational relevance of HisC

While outside the KT2440 scope, recent microbiology frequently treats HisC and histidine biosynthesis as potential antimicrobial or host-adaptation nodes because humans lack de novo histidine biosynthesis. This supports the general relevance of accurate HisC functional annotation, but pathogen-specific claims should not be transferred to KT2440 without direct evidence.

5) Expert interpretation and annotation confidence

5.1 Primary function and substrate specificity (best-supported statements)

Across authoritative reviews and structural enzymology, HisC is best described as a PLP-dependent aminotransferase that transfers the amino group from glutamate to imidazole acetol-phosphate, producing L-histidinol phosphate and α-ketoglutarate (sivaraman2001crystalstructureof pages 1-2, matte2003contributionofstructural pages 2-3, fernandez2004structuralstudiesof pages 1-2). This is the strongest functional basis for annotating PP_0967/Q88P86 as histidinol-phosphate aminotransferase.

5.2 KT2440 context strengthens pathway assignment but lacks direct hisC phenotyping

KT2440 operon context (PP_0965–PP_0967 annotated as hisGDC; RT-PCR co-transcription) places PP_0967 within histidine biosynthesis at the genomic level (molina‐henares2010identificationofconditionally pages 7-9). However, currently retrieved KT2440 genetic screens did not directly yield a PP_0967/hisC mutant phenotype, so essentiality/auxotrophy for hisC remains an inference from pathway logic plus operon annotation rather than directly demonstrated in these sources (molina‐henares2010identificationofconditionally pages 2-3).

Consolidated evidence table

Category Key facts Organism/Scope Evidence source (with DOI URL and year)
Verified identity User-specified target is hisC / PP_0967 / UniProt Q88P86 in Pseudomonas putida KT2440. KT2440 histidine-biosynthesis genes are organized in four clusters, and PP0965–PP0967 is annotated as the hisGDC cluster, placing PP_0967 as hisC in this strain-specific genomic context; this matches the expected role of histidinol-phosphate aminotransferase in histidine biosynthesis. Generic histidine-pathway references also identify HisC = histidinol aminotransferase, EC 2.6.1.9. (molina‐henares2010identificationofconditionally pages 7-9, winkler2009biosynthesisofhistidine pages 46-47) P. putida KT2440 for locus/operon context; broad bacterial annotation for enzyme name/EC Molina-Henares et al., 2010, Environmental Microbiology, DOI: https://doi.org/10.1111/j.1462-2920.2010.02166.x; Winkler & Ramos-Montañez, 2009, EcoSal Plus, DOI: https://doi.org/10.1128/ecosalplus.3.6.1.9
Catalyzed reaction and pathway step HisC (EC 2.6.1.9) catalyzes the 7th step of histidine biosynthesis: amino-group transfer from L-glutamate to imidazole acetol-phosphate / 3-(imidazol-4-yl)-2-oxo-propyl phosphate, producing L-histidinol phosphate and 2-oxoglutarate (α-ketoglutarate). The transferred amino group becomes the product’s α-amino group. (sivaraman2001crystalstructureof pages 1-2, matte2003contributionofstructural pages 2-3, winkler2009biosynthesisofhistidine pages 12-13, fernandez2004structuralstudiesof pages 1-2) Broad bacterial HisC biochemistry and structural enzymology Sivaraman et al., 2001, J. Mol. Biol., DOI: https://doi.org/10.1006/jmbi.2001.4882; Matte et al., 2003, J. Bacteriol., DOI: https://doi.org/10.1128/jb.185.14.3994-4002.2003; Winkler & Ramos-Montañez, 2009, EcoSal Plus, DOI: https://doi.org/10.1128/ecosalplus.3.6.1.9; Fernandez et al., 2004, J. Biol. Chem., DOI: https://doi.org/10.1074/jbc.m400291200
Mechanistic/structural features HisC is a PLP-dependent aminotransferase that follows a ping-pong (double-displacement) mechanism. Structural work shows a dimeric enzyme (~80 kDa total in E. coli), with each monomer containing a large PLP-binding domain, a smaller domain, and an N-terminal arm involved in dimerization/active-site shielding. The catalytic Lys214 (numbering from E. coli HisC) forms the internal aldimine with PLP; crystallography captured PMP, internal aldimine, and a covalent tetrahedral/gem-diamine-like intermediate with PLP and L-histidinol phosphate. Conserved PLP-interacting residues include Tyr55, Asn157, Asp184, Tyr187, Ser213, Lys214, Arg222. (sivaraman2001crystalstructureof pages 1-2, sivaraman2001crystalstructureof pages 7-9, matte2003contributionofstructural pages 2-3, winkler2009biosynthesisofhistidine pages 12-13, sivaraman2001crystalstructureof pages 13-14, fernandez2004structuralstudiesof pages 5-7, sivaraman2001crystalstructureof media 1b7db48c) Broad bacterial HisC structural mechanism; residue numbering from E. coli and Thermotoga maritima homologs used for functional inference Sivaraman et al., 2001, J. Mol. Biol., DOI: https://doi.org/10.1006/jmbi.2001.4882; Matte et al., 2003, J. Bacteriol., DOI: https://doi.org/10.1128/jb.185.14.3994-4002.2003; Fernandez et al., 2004, J. Biol. Chem., DOI: https://doi.org/10.1074/jbc.m400291200
KT2440 genomic context / operon evidence In P. putida KT2440, histidine biosynthesis genes occur in four genomic clusters. One cluster is PP0965–PP0967 (hisGDC), and RT-PCR evidence indicated these histidine clusters form independent operons, supporting that hisC/PP_0967 is cotranscribed with neighboring histidine-biosynthesis genes in this region. A separate monocistronic hisZ locus is PP4890. (molina‐henares2010identificationofconditionally pages 7-9) P. putida KT2440 Molina-Henares et al., 2010, Environmental Microbiology, DOI: https://doi.org/10.1111/j.1462-2920.2010.02166.x
KT2440 functional genomics / essentiality In a genome-wide miniTn5 screen of 7,760 KT2440 mutants, 79 mutants failed to grow on glucose minimal medium, mapping to 47–48 conditionally essential genes; histidine auxotrophs were recovered, including hisB (PP0289), hisF (PP0293), hisH (PP0290), and hisZ (PP4890), but no hisC mutant was recovered, so this study supports histidine-pathway importance in minimal medium without direct knockout evidence for PP_0967. A 2024 RB-TnSeq/ICA reanalysis identified histidine-related functional modules (e.g., hisA in fModule_71, “His metabolism” connection to HutC iModulon), but the cited text provides no quantitative hisC-specific fitness value or essentiality call. (molina‐henares2010identificationofconditionally pages 11-12, molina‐henares2010identificationofconditionally pages 2-3, borchert2024machinelearninganalysis pages 11-13) P. putida KT2440 functional genomics Molina-Henares et al., 2010, Environmental Microbiology, DOI: https://doi.org/10.1111/j.1462-2920.2010.02166.x; Borchert et al., 2024, mSystems, DOI: https://doi.org/10.1128/msystems.00942-23

Table: This table consolidates strain-specific identity and operon evidence for PP_0967/hisC in Pseudomonas putida KT2440 with core biochemical and structural knowledge for HisC enzymes. It also distinguishes direct KT2440 evidence from broader homolog-based inference and notes current limits of hisC-specific functional-genomics data.

Key primary sources (publication dates and URLs)

  • Molina-Henares et al. “Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium…” Environmental Microbiology (Jun 2010). https://doi.org/10.1111/j.1462-2920.2010.02166.x (molina‐henares2010identificationofconditionally pages 2-3)
  • Borchert et al. “Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440.” mSystems (Mar 2024). https://doi.org/10.1128/msystems.00942-23 (borchert2024machinelearninganalysis pages 11-13)
  • Winkler & Ramos-Montañez “Biosynthesis of Histidine.” EcoSal Plus (Dec 2009). https://doi.org/10.1128/ecosalplus.3.6.1.9 (winkler2009biosynthesisofhistidine pages 12-13)
  • Sivaraman et al. “Crystal structure of histidinol phosphate aminotransferase (HisC) from Escherichia coli…” Journal of Molecular Biology (Aug 2001). https://doi.org/10.1006/jmbi.2001.4882 (sivaraman2001crystalstructureof pages 1-2)
  • Fernandez et al. “Structural Studies of the Catalytic Reaction Pathway of a Hyperthermophilic Histidinol-phosphate Aminotransferase.” Journal of Biological Chemistry (May 2004). https://doi.org/10.1074/jbc.m400291200 (fernandez2004structuralstudiesof pages 1-2)

Notes on evidence gaps and recommendations

  1. UniProt Q88P86 record not retrieved in-tool: to fully satisfy accession-level verification, the UniProt entry should be programmatically retrieved and cross-checked for sequence features (PLP-lysine motif, domain architecture) and any curated experimental evidence.
  2. KT2440 hisC quantitative fitness/essentiality: Borchert et al. point to supplemental files for fitness metrics; extracting PP_0967 values would require retrieval of those supplements.
  3. Subcellular localization: direct localization experiments for HisC in KT2440 were not found here; targeted proteomics or fractionation references would improve this section.

References

  1. (molina‐henares2010identificationofconditionally pages 7-9): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

  2. (winkler2009biosynthesisofhistidine pages 46-47): Malcolm E. Winkler and Smirla Ramos-Montañez. Biosynthesis of histidine. Dec 2009. URL: https://doi.org/10.1128/ecosalplus.3.6.1.9, doi:10.1128/ecosalplus.3.6.1.9. This article has 247 citations.

  3. (sivaraman2001crystalstructureof pages 1-2): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  4. (winkler2009biosynthesisofhistidine pages 12-13): Malcolm E. Winkler and Smirla Ramos-Montañez. Biosynthesis of histidine. Dec 2009. URL: https://doi.org/10.1128/ecosalplus.3.6.1.9, doi:10.1128/ecosalplus.3.6.1.9. This article has 247 citations.

  5. (matte2003contributionofstructural pages 2-3): Allan Matte, J. Sivaraman, Irena Ekiel, Kalle Gehring, Zongchao Jia, and Miroslaw Cygler. Contribution of structural genomics to understanding the biology of escherichia coli. Journal of Bacteriology, 185:3994-4002, Jul 2003. URL: https://doi.org/10.1128/jb.185.14.3994-4002.2003, doi:10.1128/jb.185.14.3994-4002.2003. This article has 24 citations and is from a peer-reviewed journal.

  6. (fernandez2004structuralstudiesof pages 1-2): Francisco J. Fernandez, M. Cristina Vega, Frank Lehmann, Erika Sandmeier, Heinz Gehring, Philipp Christen, and Matthias Wilmanns. Structural studies of the catalytic reaction pathway of a hyperthermophilic histidinol-phosphate aminotransferase*. Journal of Biological Chemistry, 279:21478-21488, May 2004. URL: https://doi.org/10.1074/jbc.m400291200, doi:10.1074/jbc.m400291200. This article has 54 citations and is from a domain leading peer-reviewed journal.

  7. (sivaraman2001crystalstructureof pages 7-9): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  8. (sivaraman2001crystalstructureof pages 2-4): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  9. (sivaraman2001crystalstructureof pages 13-14): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  10. (sivaraman2001crystalstructureof media 1b7db48c): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  11. (sivaraman2001crystalstructureof media 8b8302fc): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  12. (molina‐henares2010identificationofconditionally pages 2-3): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

  13. (fernandez2004structuralstudiesof pages 9-10): Francisco J. Fernandez, M. Cristina Vega, Frank Lehmann, Erika Sandmeier, Heinz Gehring, Philipp Christen, and Matthias Wilmanns. Structural studies of the catalytic reaction pathway of a hyperthermophilic histidinol-phosphate aminotransferase*. Journal of Biological Chemistry, 279:21478-21488, May 2004. URL: https://doi.org/10.1074/jbc.m400291200, doi:10.1074/jbc.m400291200. This article has 54 citations and is from a domain leading peer-reviewed journal.

  14. (sivaraman2001crystalstructureof pages 9-10): J Sivaraman, Yunge Li, Robert Larocque, Joseph D Schrag, Miroslaw Cygler, and Allan Matte. Crystal structure of histidinol phosphate aminotransferase (hisc) from escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and l-histidinol phosphate. Journal of molecular biology, 311 4:761-76, Aug 2001. URL: https://doi.org/10.1006/jmbi.2001.4882, doi:10.1006/jmbi.2001.4882. This article has 90 citations and is from a domain leading peer-reviewed journal.

  15. (borchert2024machinelearninganalysis pages 11-13): Andrew J. Borchert, Alissa C. Bleem, Hyun Gyu Lim, Kevin Rychel, Keven D. Dooley, Zoe A. Kellermyer, Tracy L. Hodges, Bernhard O. Palsson, and Gregg T. Beckham. Machine learning analysis of rb-tnseq fitness data predicts functional gene modules in pseudomonas putida kt2440. Mar 2024. URL: https://doi.org/10.1128/msystems.00942-23, doi:10.1128/msystems.00942-23. This article has 13 citations and is from a peer-reviewed journal.

  16. (fernandez2004structuralstudiesof pages 5-7): Francisco J. Fernandez, M. Cristina Vega, Frank Lehmann, Erika Sandmeier, Heinz Gehring, Philipp Christen, and Matthias Wilmanns. Structural studies of the catalytic reaction pathway of a hyperthermophilic histidinol-phosphate aminotransferase*. Journal of Biological Chemistry, 279:21478-21488, May 2004. URL: https://doi.org/10.1074/jbc.m400291200, doi:10.1074/jbc.m400291200. This article has 54 citations and is from a domain leading peer-reviewed journal.

  17. (molina‐henares2010identificationofconditionally pages 11-12): M. Antonia Molina‐Henares, Jesús De La Torre, Adela García‐Salamanca, A. Jesús Molina‐Henares, M. Carmen Herrera, Juan L. Ramos, and Estrella Duque. Identification of conditionally essential genes for growth of pseudomonas putida kt2440 on minimal medium through the screening of a genome‐wide mutant library. Environmental Microbiology, 12:1468-1485, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2010.02166.x, doi:10.1111/j.1462-2920.2010.02166.x. This article has 89 citations and is from a domain leading peer-reviewed journal.

Artifacts

Citations

  1. winkler2009biosynthesisofhistidine pages 46-47
  2. sivaraman2001crystalstructureof pages 13-14
  3. fernandez2004structuralstudiesof pages 9-10
  4. sivaraman2001crystalstructureof pages 7-9
  5. borchert2024machinelearninganalysis pages 11-13
  6. winkler2009biosynthesisofhistidine pages 12-13
  7. sivaraman2001crystalstructureof pages 1-2
  8. fernandez2004structuralstudiesof pages 1-2
  9. matte2003contributionofstructural pages 2-3
  10. sivaraman2001crystalstructureof pages 2-4
  11. sivaraman2001crystalstructureof pages 9-10
  12. fernandez2004structuralstudiesof pages 5-7
  13. https://doi.org/10.1111/j.1462-2920.2010.02166.x;
  14. https://doi.org/10.1128/ecosalplus.3.6.1.9
  15. https://doi.org/10.1006/jmbi.2001.4882;
  16. https://doi.org/10.1128/jb.185.14.3994-4002.2003;
  17. https://doi.org/10.1128/ecosalplus.3.6.1.9;
  18. https://doi.org/10.1074/jbc.m400291200
  19. https://doi.org/10.1111/j.1462-2920.2010.02166.x
  20. https://doi.org/10.1128/msystems.00942-23
  21. https://doi.org/10.1006/jmbi.2001.4882
  22. https://doi.org/10.1111/j.1462-2920.2010.02166.x,
  23. https://doi.org/10.1128/ecosalplus.3.6.1.9,
  24. https://doi.org/10.1006/jmbi.2001.4882,
  25. https://doi.org/10.1128/jb.185.14.3994-4002.2003,
  26. https://doi.org/10.1074/jbc.m400291200,
  27. https://doi.org/10.1128/msystems.00942-23,

OpenScientist

(hisC-deep-research-openscientist.md)
Functional Annotation of *hisC* (Q88P86, PP_0967) in *Pseudomonas putida* KT2440 OpenScientist openscientist-autonomous 2 artifacts 2026-07-17T05:51:42.205579

Functional Annotation of hisC (Q88P86, PP_0967) in Pseudomonas putida KT2440

Gene: hisC (OrderedLocusName PP_0967)
Protein: Histidinol-phosphate aminotransferase (HisC); AltName: imidazole-acetol-phosphate transaminase
UniProt: Q88P86 · EC: 2.6.1.9 · KEGG Ortholog: K00817
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / NCIMB 11950 / KT2440)


1. Summary (Answer to the Research Question)

hisC encodes histidinol-phosphate aminotransferase (HisC, EC 2.6.1.9), a soluble, cytoplasmic, pyridoxal-5′-phosphate (PLP)-dependent class-II aminotransferase that catalyzes the seventh step of de novo L-histidine biosynthesis. Working as a homodimer, it performs a reversible transamination that, in the biosynthetic direction, transfers the α-amino group of L-glutamate onto imidazole-acetol phosphate (3-(imidazol-4-yl)-2-oxopropyl phosphate) to yield L-histidinol phosphate + 2-oxoglutarate. Its substrate specificity is dominated by recognition of the substrate phosphate group; members of this subfamily can additionally act as aromatic-amino-acid aminotransferases. In P. putida the gene sits in a compact hisG–hisD–hisC cluster and is conditionally essential — its loss causes histidine auxotrophy on minimal medium.

The identity of the target was rigorously verified: gene symbol, organism, EC number, protein family, and catalytic residues are all mutually consistent across UniProt, KEGG, and the primary structural literature on close orthologs. No gene-symbol ambiguity was encountered.


2. Identity Verification

Attribute Provided target Confirmed by this study
Gene symbol hisC UniProt Q88P86; KEGG ppu:PP_0967 (SYMBOL hisC)
Enzyme Histidinol-phosphate aminotransferase, EC 2.6.1.9 UniProt catalytic activity; KEGG KO K00817; EC 2.6.1.9
Organism P. putida KT2440 KEGG ORGANISM ppu; UniProt organism
Family Class-II PLP-dependent aminotransferase UniProt SIMILARITY; KEGG BRITE "Aminotransferase Class II"; Pfam Aminotran_1_2
Locus/position PP_0967 KEGG POSITION 1,106,849–1,107,895

All identifiers converge on a single, well-characterized enzyme family. The verification requirement is satisfied.


3. Primary Function: Reaction Catalyzed and Substrate Specificity

3.1 The reaction

HisC catalyzes the PLP-dependent, reversible transamination (UniProt Q88P86 catalytic activity):

L-histidinol phosphate + 2-oxoglutarate ⇌ 3-(imidazol-4-yl)-2-oxopropyl phosphate (imidazole-acetol phosphate) + L-glutamate

Physiologically, the enzyme operates in the biosynthetic (amination) direction: "histidinol-phosphate aminotransferase catalyzes the transfer of the amino group from glutamate to imidazole acetol-phosphate producing 2-oxoglutarate and histidinol phosphate" (Fernández et al., 2004, PMID 15007066). This is the seventh step in the synthesis of histidine within eubacteria (Sivaraman et al., 2001, PMID 11518529), corresponding to step 7 of 9 from 5-phospho-α-D-ribose-1-diphosphate (PRPP) in the UniProt/KEGG pathway map (KEGG module M00026).

3.2 Substrate specificity

  • The natural substrate pair is L-histidinol phosphate / 2-oxoglutarate (amino donor L-glutamate in the biosynthetic direction).
  • The substrate phosphate group is the principal specificity determinant. In Corynebacterium glutamicum HisC, "the hydrogen bond between the side chain of this residue [Tyr21] and the phosphate group of His-P is important for recognition of the natural substrate and discrimination against other potential amino donors such as phenylalanine and leucine" (Marienhagen et al., 2008, PMID 18560156).
  • Family-level substrate promiscuity / moonlighting: In organisms lacking dedicated aromatic aminotransferases, HisC also transaminates aromatic amino acids. The Thermotoga maritima enzyme "accepts histidinol phosphate, tyrosine, tryptophan, and phenylalanine, but not histidine, as substrates" (Fernández et al., 2004, PMID 15007066). Consistent with this, KEGG maps P. putida PP_0967 not only to histidine metabolism (ppu00340) but also to tyrosine (ppu00350), phenylalanine (ppu00360), and aromatic amino acid biosynthesis (ppu00400) pathways. However, this mapping reflects family-level catalytic capacity rather than a demonstrated dedicated role: P. putida KT2440 encodes two dedicated aromatic-amino-acid aminotransferases (PP_1972 and PP_3590; KEGG K00832, EC 2.6.1.57), whereas T. maritima—where HisC broadens to aromatic substrates—lacks such enzymes. Any aromatic-amino-acid transamination by P. putida HisC is therefore likely biochemically possible but physiologically redundant/minor, and has not been directly demonstrated.
  • Genetic corroboration of the pathway step: hisC mutants (in Micrococcus luteus) "accumulated imidazoleacetol" (Kane-Falce & Kloos, 1975, PMID 1126626), confirming HisC acts at the imidazole-acetol(-phosphate) node.

4. Mechanism, Cofactor, and Quaternary Structure

  • Cofactor: pyridoxal 5′-phosphate (PLP), covalently bound as an internal aldimine (Schiff base) at Lys210 of P. putida HisC (UniProt "N6-(pyridoxal phosphate)lysine" at residue 210).
  • Catalytic cycle: classic aminotransferase ping-pong (two half-reaction) mechanism cycling between the PLP and pyridoxamine-5′-phosphate (PMP) forms. In the E. coli ortholog, a covalent tetrahedral complex "consisting of PLP and l-histidinol phosphate attached to Lys214" was captured crystallographically, resembling the transient gem-diamine intermediate (Sivaraman et al., 2001, PMID 11518529). PMP-enzyme complexes have been trapped in both E. coli (PMID 11518529) and C. glutamicum (PMID 18560156).
  • Quaternary structure: homodimer. "HisC is a dimeric enzyme with a mass of approximately 80 kDa" (Sivaraman et al., 2001, PMID 11518529); UniProt lists Q88P86 as a homodimer. Each monomer comprises a large α/β/α PLP-binding domain plus a small domain, with an N-terminal arm contributing to dimerization; the active site lies at the dimer interface.
  • Conserved catalytic apparatus in Q88P86 (direct sequence evidence): A global alignment of Q88P86 against the crystallographically characterized E. coli HisC (P06986) shows the active-site residues are conserved — Tyr55→Tyr57, Asp184→Asp181, Tyr187→Tyr184, Ser213→Ser209, Lys214→Lys210, Arg222→Arg218. The reference set is described as: "Residues that interact with the PLP cofactor, including Tyr55, Asn157, Asp184, Tyr187, Ser213, Lys214 and Arg222, are conserved in the family of aspartate, tyrosine and histidinol phosphate aminotransferases" (Sivaraman et al., 2001, PMID 11518529). The mapped catalytic lysine (Lys214→Lys210) matches UniProt's independent PLP-attachment annotation exactly, validating the alignment. Asp181 stabilizes the protonated PLP pyridinium nitrogen (a fold-type-I/class-II hallmark), Arg218 binds the substrate α-carboxylate/2-oxoglutarate, and Tyr57 contributes to substrate/phosphate binding. This upgrades the annotation from database transfer to residue-level evidence that Q88P86 is a catalytically competent HisC.

5. Subcellular Localization

HisC acts in the cytoplasm as a soluble enzyme. UniProt Q88P86 shows no signal peptide, transmembrane segment, or lipidation/anchor; all characterized bacterial orthologs (E. coli, Salmonella typhimurium, C. glutamicum, T. maritima) are soluble proteins purified from soluble extracts and crystallized as such (e.g., "Crystalline L-histidinol phosphate aminotransferase from Salmonella typhimurium", Henderson & Snell, 1973, PMID 4632247). Its substrates are cytosolic phosphorylated intermediates and glutamate/2-oxoglutarate. The entire de novo histidine biosynthetic pathway is cytoplasmic, so HisC exerts its function there.


6. Pathway Context and Biological Role

  • Pathway: de novo L-histidine biosynthesis (PRPP → histidine), an unbranched, ancient pathway. HisC provides the transamination that installs the α-amino group of the histidine backbone, converting imidazole-acetol phosphate to L-histidinol phosphate. The immediately downstream enzyme HisD (histidinol dehydrogenase) then oxidizes L-histidinol phosphate/L-histidinol to L-histidine. (Pathway architecture reviewed in Alifano et al., 1996, Microbiol. Rev., PMID 8852895.)
  • Amino-donor coupling: By consuming L-glutamate and releasing 2-oxoglutarate, HisC links histidine biosynthesis to the cell's central glutamate/2-oxoglutarate nitrogen pool.
  • Genomic organization / co-regulation: In P. putida KT2440, hisC (PP_0967) lies immediately downstream of hisG (PP_0965, ATP phosphoribosyltransferase — the first, feedback-regulated step) and hisD (PP_0966, histidinol dehydrogenase — the terminal steps), all on the same strand. hisD and hisC are separated by only 2 bp, indicating translational coupling and operon-like co-transcription of a hisG–hisD–hisC cluster. This physically and transcriptionally embeds HisC within the histidine biosynthetic program.
  • Physiological requirement (organism-specific): A genome-wide mini-Tn5 transposon screen of P. putida KT2440 identified de novo amino-acid biosynthesis genes as conditionally essential on glucose minimal medium — "Auxotrophs for all amino acids predicted by the in silico models were found" (Molina-Henares et al., 2010, PMID 20158506). Because HisC catalyzes an obligatory, non-bypassable step of the single linear histidine pathway (no isozyme), hisC loss yields histidine auxotrophy: required for prototrophic growth but dispensable when histidine is supplied.

7. Evidence Summary

Claim Evidence type Source
EC 2.6.1.9; His-P aminotransferase; step 7 of His biosynthesis Database annotation + primary structural lit. UniProt Q88P86; KEGG K00817; PMID 11518529, 15007066
PLP cofactor at Lys210; ping-pong (PLP↔PMP) mechanism UniProt residue annotation + ortholog crystal structures UniProt Q88P86; PMID 11518529, 18560156
Homodimer, ~80 kDa Ortholog biochemistry/crystallography PMID 11518529; UniProt
Catalytic residues conserved in Q88P86 Bioinformatic alignment (this study) vs P06986; PMID 11518529
Phosphate-group specificity; aromatic-AA moonlighting Site-directed mutagenesis + substrate assays PMID 18560156, 15007066; KEGG pathway mapping
Cytoplasmic, soluble Sequence features + ortholog purification UniProt Q88P86; PMID 4632247
hisGDC cluster / co-regulation Genome coordinates KEGG ppu genome
Conditionally essential (His auxotrophy) Genome-wide transposon screen PMID 20158506

8. Supported and Refuted Hypotheses

Supported:
- H1 — hisC encodes a functional PLP-dependent histidinol-phosphate aminotransferase (EC 2.6.1.9). Strongly supported (database + conserved catalytic residues incl. Lys210-PLP).
- H2 — HisC operates in the cytoplasm as a soluble homodimer. Supported.
- H3 — HisC is embedded in a co-regulated histidine operon (hisG–hisD–hisC) and is required for de novo His synthesis. Supported (genomic + auxotrophy evidence).

Partially supported / open:
- H4 — P. putida HisC physiologically moonlights as an aromatic-amino-acid aminotransferase. Plausible but unproven, and likely minor: supported by family behavior (T. maritima) and KEGG pathway mapping, but no direct P. putida enzymology exists. P. putida encodes two dedicated aromatic-amino-acid aminotransferases (PP_1972, PP_3590; K00832), making any HisC moonlighting role physiologically redundant/minor in this organism.

Refuted / ruled out:
- The protein is not a membrane transporter, structural protein, or signaling molecule; it is a soluble metabolic enzyme (no TM/signal features).
- HisC does not transaminate free histidine (orthologs do not accept histidine as substrate).


9. Limitations and Future Directions

  • No direct enzymology exists for the P. putida KT2440 protein itself. All mechanistic and kinetic detail is transferred from orthologs (E. coli, C. glutamicum, T. maritima, Salmonella) plus bioinformatic residue mapping specific to Q88P86. This is standard and well-justified for a HAMAP-ruled housekeeping enzyme, but P. putida-specific kinetics (Km for His-P, imidazole-acetol phosphate, glutamate) and any aromatic-AA side activity remain to be measured.
  • The conditional-essentiality inference for hisC rests on pathway logic plus a genome-wide screen that reports amino-acid auxotrophs collectively; a targeted hisC knockout with His-supplementation rescue would confirm it directly.
  • Future work: solve/model the P. putida HisC structure (AlphaFold model is available), test aromatic-amino-acid transaminase activity, and define operon boundaries and HisG feedback regulation experimentally.

References (PMIDs)

  • Sivaraman J. et al. (2001) Crystal structure of histidinol phosphate aminotransferase (HisC) from E. coli. PMID 11518529
  • Fernández F.J. et al. (2004) Structural studies of the catalytic reaction pathway of a hyperthermophilic HisC (T. maritima). PMID 15007066
  • Marienhagen J. et al. (2008) Substrate recognition by HisC from C. glutamicum. PMID 18560156
  • Kane-Falce & Kloos (1975) Histidine biosynthesis in Micrococcus luteus. PMID 1126626
  • Henderson & Snell (1973) Crystalline His-P aminotransferase from S. typhimurium. PMID 4632247
  • Alifano P. et al. (1996) Histidine biosynthetic pathway and genes. PMID 8852895
  • Molina-Henares M.A. et al. (2010) Conditionally essential genes for growth of P. putida KT2440. PMID 20158506
  • Rutkiewicz M. et al. (2023) Plant HISN6 substrate specificity/structure. PMID 36842242

Database sources: UniProt Q88P86; KEGG ppu:PP_0967 (K00817); Pfam Aminotran_1_2; InterPro IPR005861 (HisP_aminotrans), IPR050106.

Artifacts

📄 View Raw YAML

id: Q88P86
gene_symbol: hisC
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: Histidinol-phosphate aminotransferase (HisC; EC 2.6.1.9), a cytoplasmic pyridoxal-5'-phosphate (PLP)-dependent class-II aminotransferase that catalyzes the seventh step of L-histidine biosynthesis. It transfers an amino group from L-glutamate to imidazole-acetol phosphate (3-(imidazol-4-yl)-2-oxopropyl phosphate), producing L-histidinol phosphate and 2-oxoglutarate. The enzyme functions as a homodimer with active sites at the dimer interface; PLP is covalently bound as an internal aldimine to an active-site lysine (Lys210 in this protein) and catalysis proceeds via a ping-pong mechanism through a pyridoxamine-5'-phosphate intermediate. In Pseudomonas putida KT2440 the gene (PP_0967) lies within a histidine-biosynthesis gene cluster. Aromatic-amino-acid transamination is documented for some HisC orthologs but has not been demonstrated for the KT2440 protein.
existing_annotations:
- term:
    id: GO:0000105
    label: L-histidine biosynthetic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: involved_in
  review:
    summary: HisC catalyzes the seventh step of histidine biosynthesis; this BP term is well supported by family/HAMAP-rule assignment, the UniProt pathway annotation, and operon context in KT2440.
    action: ACCEPT
    reason: Core biological process for this enzyme. The histidinol-phosphate aminotransferase function places it squarely in the L-histidine biosynthetic pathway (UniPathway UPA00031; HAMAP-Rule MF_01023).
    supported_by:
    - reference_id: file:PSEPK/hisC/hisC-deep-research-openscientist.md
      supporting_text: >-
        catalyzes the **seventh step of de novo L-histidine biosynthesis**
- term:
    id: GO:0004400
    label: L-histidinol-phosphate:2-oxoglutarate transaminase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: This is the specific molecular function of HisC (EC 2.6.1.9), transaminating L-histidinol phosphate with 2-oxoglutarate/L-glutamate. The UniProt CATALYTIC ACTIVITY block and HAMAP rule directly support this.
    action: ACCEPT
    reason: Represents the core molecular function. Strongly supported by family assignment (HisP_aminotrans subfamily, TIGR01141 hisC), Rhea:23744, and EC 2.6.1.9.
    supported_by:
    - reference_id: file:PSEPK/hisC/hisC-deep-research-openscientist.md
      supporting_text: >-
        *hisC* encodes **histidinol-phosphate aminotransferase (HisC, EC
        2.6.1.9)**
- term:
    id: GO:0016740
    label: transferase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: A high-level parent of the specific aminotransferase activity already annotated (GO:0004400). It is correct but uninformative given the more precise term.
    action: MARK_AS_OVER_ANNOTATED
    reason: Redundant generic ancestor of GO:0004400; adds no information beyond the specific transaminase MF term.
- term:
    id: GO:0030170
    label: pyridoxal phosphate binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: HisC is a PLP-dependent enzyme that covalently binds pyridoxal 5'-phosphate as an internal aldimine at the active-site lysine (MOD_RES 210 in this entry). Well supported by the COFACTOR annotation and conserved PLP-lysine motif.
    action: KEEP_AS_NON_CORE
    reason: PLP binding is an essential, well-supported cofactor interaction, but the substrate-specific transaminase activity is the defining molecular function. The PLP-lysine internal aldimine and ping-pong mechanism are documented for HisC homologs (see hisC-deep-research-falcon.md).
- term:
    id: GO:0140385
    label: amino acid transaminase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000117
  qualifier: enables
  review:
    summary: A broad parent term covering aminotransferase activity on amino acid substrates. Correct but less specific than GO:0004400, which is already annotated.
    action: MARK_AS_OVER_ANNOTATED
    reason: Generic ancestor of the specific histidinol-phosphate transaminase activity; the more precise term GO:0004400 already captures this function.
core_functions:
- description: Catalyzes the PLP-dependent transamination of imidazole-acetol phosphate to L-histidinol phosphate (using L-glutamate as amino donor), the seventh step of L-histidine biosynthesis.
  molecular_function:
    id: GO:0004400
    label: L-histidinol-phosphate:2-oxoglutarate transaminase activity
  supported_by:
  - reference_id: file:PSEPK/hisC/hisC-deep-research-falcon.md
    supporting_text: KT2440 histidine-biosynthesis genes PP0965-PP0967 are annotated as the hisGDC cluster, placing PP_0967 as hisC within histidine biosynthesis.
  - reference_id: file:PSEPK/hisC/hisC-deep-research-openscientist.md
    supporting_text: >-
      transfers the α-amino group of **L-glutamate onto imidazole-acetol
      phosphate (3-(imidazol-4-yl)-2-oxopropyl phosphate)** to yield
      **L-histidinol phosphate + 2-oxoglutarate**
  directly_involved_in:
  - id: GO:0000105
    label: L-histidine biosynthetic process
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000117
  title: Electronic Gene Ontology annotations created by ARBA machine learning models
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:20158506
  title: Identification of conditionally essential genes for growth of Pseudomonas putida KT2440 on minimal medium through the screening of a genome-wide mutant library
  full_text_unavailable: true
  findings:
  - statement: In P. putida KT2440 the histidine-biosynthesis genes PP0965-PP0967 are annotated as the hisGDC cluster, with PP_0967 corresponding to hisC, and RT-PCR co-transcription assays support operon organization of these clusters.
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: 'Citation-integrity fix: the original identifier PMID:19838707 was a hallucinated/wrong identifier (resolves to an unrelated knee-arthroplasty article). Replaced with PMID:20158506, the correct Molina-Henares et al. 2010 Environ Microbiol paper (DOI 10.1111/j.1462-2920.2010.02166.x), recovered via DOI lookup and PubMed-verified to match the supporting text (KT2440 hisGDC cluster, PP_0967 = hisC). Supporting snippet is paraphrased from the abstract-only source.'
- id: PMID:11518529
  title: Crystal structure of histidinol phosphate aminotransferase (HisC) from Escherichia coli, and its covalent complex with pyridoxal-5'-phosphate and L-histidinol phosphate
  findings:
  - statement: E. coli HisC is a PLP-dependent homodimeric aminotransferase with active sites at the dimer interface; the active-site lysine (Lys214) forms an internal aldimine with PLP and catalysis proceeds through PMP via a ping-pong mechanism.
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: >-
      Citation-integrity fix. The original identifier PMID:11470432 was a wrong
      identifier (resolves to an unrelated article, "Crystal structures of the
      MJ1267 ATP binding cassette reveal an induced-fit effect at the ATPase active
      site of an ABC transporter"). Replaced with PMID:11518529, the correct
      Sivaraman et al. 2001 J Mol Biol 311:761-776 paper
      (DOI 10.1006/jmbi.2001.4882), recovered via DOI lookup and PubMed-verified to
      match the title and supporting text (E. coli HisC crystal structure, PLP
      internal aldimine at Lys214, dimer interface). Structural/mechanistic support
      for the HisC family (E. coli homolog).
- id: file:PSEPK/hisC/hisC-deep-research-openscientist.md
  title: OpenScientist functional report for PSEPK HisC
  findings:
  - statement: >-
      Supports the exact HisC reaction and histidine-pathway assignment for
      Q88P86 while distinguishing family-level aromatic-amino-acid activity from
      a demonstrated physiological role in KT2440.