tyrB

UniProt ID: Q88LG1
Organism: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
Review Status: DRAFT
📝 Provide Detailed Feedback

Gene Description

tyrB (PP_1972; also referred to as tyrB-1) is a cytoplasmic, pyridoxal 5'-phosphate (PLP)-dependent aminotransferase of the class-I (fold-type I) aspartate aminotransferase superfamily. It catalyzes reversible transamination in which an amino group is transferred from an amino acid donor to a 2-oxoacid acceptor (commonly 2-oxoglutarate, yielding L-glutamate). It is annotated as an aromatic-amino-acid aminotransferase, interconverting aromatic amino acids (L-tyrosine, L-phenylalanine) and their cognate aromatic 2-oxoacids (4-hydroxyphenylpyruvate, phenylpyruvate), and contributes to aromatic amino acid biosynthesis and catabolism. Like many class-I PLP aminotransferases it is a homodimer with active sites formed at the subunit interface. In P. putida KT2440 it is one of several aminotransferase isozymes acting on aromatic amino acids; genetic studies show that loss of tyrB alone causes only mild aromatic-amino-acid utilization phenotypes because of redundancy with paralogous aminotransferases (notably PP_3590/AmaC and tyrB-2/phhC), so its physiological role overlaps with those enzymes.

Existing Annotations Review

GO Term Evidence Action Reason
GO:0003824 catalytic activity
IEA
GO_REF:0000002
MARK AS OVER ANNOTATED
Summary: Generic root-level molecular function term. tyrB is an enzyme, so this is not wrong, but it is uninformatively general and is fully subsumed by the more specific transaminase/aminotransferase activity terms.
Reason: "catalytic activity" is the MF root and conveys no specific information. The more precise terms GO:0008483 (transaminase activity) and GO:0004838 (L-tyrosine:2-oxoglutarate transaminase activity) capture the actual function.
GO:0004838 L-tyrosine:2-oxoglutarate transaminase activity
IEA
GO_REF:0000118
ACCEPT
Summary: Specific aromatic (tyrosine) aminotransferase activity assigned by TreeGrafter from the PTHR11879:SF37 "aromatic-amino-acid aminotransferase" subfamily. This is consistent with the protein family, the COG1448 assignment, and with biochemical/genetic characterization of P. putida aromatic aminotransferases. This is the best representation of the gene's core molecular function.
Reason: Domain/family evidence (class-I PLP aminotransferase, PANTHER ArAT subfamily SF37) plus genetic evidence in KT2440 (tyrB mutants show a tyrosine-utilization phenotype; PMID:20050871) support tyrosine aminotransferase activity. The enzyme is likely promiscuous across aromatic amino acids (also acting on phenylalanine), but this term well captures the central characterized activity.
GO:0005829 cytosol
IEA
GO_REF:0000118
ACCEPT
Summary: Cytosolic localization predicted by TreeGrafter. Soluble class-I PLP aminotransferases acting in amino acid metabolism are cytoplasmic enzymes; the sequence has no signal peptide or transmembrane region. Consistent with the expected localization.
Reason: Aromatic aminotransferases in this family are soluble cytoplasmic enzymes. No experimental KT2440 localization data exist, but the prediction is biologically appropriate and there is no evidence for periplasmic/membrane localization.
GO:0006520 amino acid metabolic process
IEA
GO_REF:0000002
MODIFY
Summary: Broad biological process term. tyrB participates in (aromatic) amino acid metabolism, so this is correct but general. A more specific process such as aromatic amino acid family metabolism / phenylalanine or tyrosine biosynthesis or catabolism would be more informative.
Reason: The annotation is correct in essence but too high-level. tyrB acts specifically on aromatic amino acids (Tyr/Phe), so the more specific "aromatic amino acid family metabolic process" better reflects the characterized role while remaining defensible from family + genetic evidence. (Chorismate metabolic process is not appropriate: tyrB acts downstream of chorismate on the aromatic amino acids/2-oxoacids, not on chorismate itself.)
GO:0008483 transaminase activity
IEA
GO_REF:0000120
KEEP AS NON CORE
Summary: General transaminase (aminotransferase) activity. Correct and well supported by the class-I PLP-dependent aminotransferase family assignment, but less specific than GO:0004838. Useful as a parent term.
Reason: Accurately describes the enzymatic class but is a parent of the more specific aromatic aminotransferase term that represents the core function. Retain as supporting/non-core rather than as the primary MF.
GO:0030170 pyridoxal phosphate binding
IEA
GO_REF:0000120
ACCEPT
Summary: PLP cofactor binding. tyrB is a PLP-dependent enzyme (UniProt COFACTOR: pyridoxal 5'-phosphate; conserved PROSITE PS00105 class-I aminotransferase PLP-binding motif). This is a well-supported and informative molecular function annotation.
Reason: Strong family/motif evidence (IPR004838/IPR004839, PROSITE AA_TRANSFER_ CLASS_1, UniProt cofactor annotation) for PLP binding, which is essential for the transamination mechanism.
GO:0042802 identical protein binding
IEA
GO_REF:0000118
MARK AS OVER ANNOTATED
Summary: Self-association annotation reflecting the homodimeric quaternary structure typical of class-I aminotransferases (UniProt SUBUNIT: Homodimer). While the homodimer assignment is reasonable, "identical protein binding" is an uninformative interaction term that does not convey biological function and is a frequent TreeGrafter over-propagation.
Reason: Homodimerization is a structural property rather than a distinct molecular function; the term adds little and is propagated electronically without direct evidence for this protein. Per curation guidance, generic "protein binding"-type terms are discouraged.

Core Functions

PLP-dependent aromatic-amino-acid aminotransferase catalyzing reversible transamination between aromatic amino acids (L-tyrosine, L-phenylalanine) and their 2-oxoacids using 2-oxoglutarate/L-glutamate as the amino acceptor/donor pair, functioning in aromatic amino acid biosynthesis and catabolism.

Supporting Evidence:

References

Gene Ontology annotation through association of InterPro records with GO terms
TreeGrafter-generated GO annotations
Combined Automated Annotation using Multiple IEA Methods
tyrB-2 and phhC genes of Pseudomonas putida encode aromatic amino acid aminotransferase isozymes: evidence at the protein level
  • P. putida aromatic aminotransferase isozymes preferentially transaminate aromatic amino acids and aromatic 2-oxoacids (best substrates L-phenylalanine and phenylpyruvate), using PLP cofactor and 2-oxoglutarate as amino acceptor.
Identification and characterization of the PhhR regulon in Pseudomonas putida
  • Genetic study of aromatic amino acid catabolism in P. putida KT2440; tyrB-1 (PP_1972) and tyrB-2 mutants show altered doubling times on tyrosine and phenylalanine as nitrogen sources, implicating tyrB-family aminotransferases in aromatic amino acid utilization.
Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440
  • Disruption of tyrB (PP_1972) did not inhibit growth on L-phenylalanine or L-tyrosine as sole nitrogen sources, whereas disruption of AmaC (PP_3590) abolished growth; the authors propose PP_3590 as the dominant L-tyrosine aminotransferase, indicating PP_1972 is functionally redundant under those conditions.
Nitrogen metabolism in Pseudomonas putida: functional analysis using random barcode transposon sequencing
  • RB-TnSeq fitness data show only weak single-gene fitness effects for PP_1972 on aromatic nitrogen sources, and a PP_3590/PP_1972 double knockout did not cause phenylalanine auxotrophy, consistent with redundancy among aromatic aminotransferases.

Suggested Questions for Experts

Q: What is the in vitro substrate range and kinetic preference of purified PP_1972 (Tyr vs Phe vs Trp; 2-oxoglutarate vs pyruvate as acceptor), given that direct enzymology exists for the paralogs but not for PP_1972 itself?

Q: What is the division of labor among the P. putida KT2440 aromatic aminotransferase isozymes (PP_1972/tyrB-1, PP_3590/AmaC, tyrB-2/phhC) in aromatic amino acid biosynthesis versus catabolism, and under what conditions is PP_1972 non-redundant?

Suggested Experiments

Experiment: Express and purify recombinant PP_1972 and determine kinetic constants (Km/kcat) against a panel of amino donors (Tyr, Phe, Trp, Asp) and 2-oxoacid acceptors to define substrate specificity directly.

Experiment: Construct single and combinatorial in-frame deletions of PP_1972, PP_3590, and tyrB-2/phhC and assay growth on aromatic amino acids as sole nitrogen and carbon sources to resolve the redundancy network and any condition-specific, non-redundant role of PP_1972.

Deep Research

Asta

(tyrB-deep-research-asta.md)
Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y... Asta Asta Scientific Corpus Retrieval 20 citations 2026-07-05T20:21:40.925977

Asta Literature Retrieval: Gene Research for Functional Annotation ⚠️ CRITICAL: Gene/Protein Identification Context BEFORE YOU BEGIN RESEARCH: Y...

This report is retrieval-only and is generated directly from Asta results.

  • Papers retrieved: 20
  • Snippets retrieved: 20

Relevant Papers

[1] LMPD: LIPID MAPS proteome database

  • Authors: Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam
  • Year: 2005
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  • DOI: 10.1093/nar/gkj122
  • PMID: 16381922
  • PMCID: 1347484
  • Citations: 92
  • Influential citations: 2
  • Summary: The initial release of the LIPID MAPS Proteome Database contains 2959 records, representing human and mouse proteins involved in lipid metabolism, and this LMPD protein list was enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources.
  • Evidence snippets:
  • Snippet 1 (score: 0.790)
    > For each record selected from the results summary, all LMPD data relevant to that protein are displayed, with external database IDs linked to their respective resources.
    > Annotations are organized by category: Record Overview, Gene/GO/KEGG Information, UniProt Annotations, and Related Proteins. The record overview contains LMPD_ID, species, description, gene symbols, lipid categories, EC number, molecular weight, sequence length and protein sequence. Gene information includes Entrez Gene ID, chromosome, map location, primary name, primary symbol and alternate names and symbols; Gene Ontology (GO) IDs and descriptions, and KEGG pathway IDs and descriptions. UniProt annotations include primary accession number, entry name and comments such as catalytic activity, enzyme regulation, function and similarity. For related proteins and splice variants, we display source database, database ID, sequence length, and title.

[2] Avian Immunome DB: an example of a user-friendly interface for extracting genetic information

  • Authors: Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al.
  • Year: 2020
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  • DOI: 10.1186/s12859-020-03764-3
  • PMID: 33176685
  • PMCID: 7661159
  • Citations: 6
  • Summary: The Avian Immunome DB (Avimm) for easy gene property extraction as exemplified by avian immune genes is presented and described, which contains 1170 distinct avian immune genes with canonical gene symbols and 612 synonyms across 363 bird species.
  • Evidence snippets:
  • Snippet 1 (score: 0.744)
    > Ever since the advent of commercial next-generation sequencing platforms in the early 2000s with its associated decrease in sequencing costs [1], the number of DNA sequences increased considerably [2]. Generally, these data become publicly accessible in databases provided by projects focussing on different aspects of biological sequence information [3,4]. Ensembl [5] and NCBI [6] for instance, have a strong focus on genome annotation with the help of RNA transcript information while UniProt has a pronounced emphasis on protein-coding genes and biological function of proteins. UniProt's records are either based on manually annotated, non-redundant protein sequences (SwissProt) or on highquality computationally analysed records, which are enriched with automatic annotation (TrEMBL) [7]. Relying on accurate genome annotations and protein descriptions, Gene Ontology (GO) [8,9] categorises gene products and fits them into a computational model of biological systems. Their assignment deploys a controlled vocabulary, so-called GO terms, to link genes and gene products to biological processes, cellular components, or molecular functions.
    > However, genome annotation is not standardised, and each service provider uses their own custom-built annotation pipelines. As a consequence, this often leads to ambiguity in gene names during genome annotation with different gene symbols being given to the same gene or the same gene symbol being given to different, but similar genes. Additionally, since the pre-existing wealth of sequencing information relies on model organisms like human and mouse, there is a strong bias in gene symbols towards those chosen for these species. Particularly for model species, this issue has been partially addressed, for example by the Human Genome Organisation (HUGO) Gene Nomenclature Committee (HGNC) [10], the Vertebrate Gene Nomenclature Committee (VGNC) [11], or the Chicken Gene Nomenclature Consortium [12]. However, this neither guarantees that gene names are harmonised among these consortia, nor does it keep researchers from assigning alternative gene symbols in their annotations, especially when working with non-model species.

[3] Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes

  • Authors: Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al.
  • Year: 2021
  • Venue: mSystems
  • URL: https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  • DOI: 10.1128/mSystems.00673-21
  • PMID: 34726489
  • PMCID: 8562490
  • Citations: 15
  • Summary: This work systematically updates the functional genome annotation of Mycobacterium tuberculosis virulent type strain H37Rv and identifies hundreds of high-confidence candidates for mechanisms of antibiotic resistance, virulence factors, and basic metabolism and other functions key in clinical and basic tuberculosis research.
  • Evidence snippets:
  • Snippet 1 (score: 0.744)
    > 3. Fig. S2B -match/mismatch colours mixed up? (I think match should be teal and mismatch -red?) 4. Line 162-163: Rv1430 is in UniProt (EC 3.1.1.-) and has been present in Uniprot since version 45 of the gene record: https://www.uniprot.org/uniprot/L7N697. I presume you had conducted your literature analysis before the UniProt entry was updated to include the EC code, so maybe you can add the dates when the data was retrieved from UniProt and other databases you used in the Materials and Methods section? 5. Supplementary text, p. 9, first paragraph. I believe that an unrelated fragment of text was copy-pasted into the second sentence of the paragraph ("Many mutations that altered bacterial clearance...") 6. Supplementary text, p. 12, final paragraph. It should be Rv1191, not Rv1191c. Could you also add a short explanation why you believe it should be classified as a cathepsin (what protein did you transfer this annotation from)?
    > Reviewer #3 (Comments for the Author):
    > In this manuscript, Modlin et al., attempt to tackle the problem of assigning functions to ~1700 hypothetical and/or underannotated genes in the Mycobacterium tuberculosis H37Rv (Mtb) genome. Rapid and accurate annotation of microbial genomes is indeed a very critical and under appreciated part of microbial ecophysiology. This step is especially crucial for pathogenic organisms such as Mtb where accurate functional annotation of these hypothetical proteins could unravel mechanisms which could act as drug targets. The authors employed a two-pronged strategy to define a set of these unannotated or under-annotated genes and to then provide possible functions for many of these genes. First, they undertook a large-scale manual curation of literature to assign functions (including EC numbers for enzymatic functions) to ~575 genes.

[4] Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome

  • Authors: Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani
  • Year: 2023
  • Venue: Medicine
  • URL: https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  • DOI: 10.1097/MD.0000000000034514
  • PMID: 37565925
  • PMCID: 10419424
  • Citations: 2
  • Summary: A network pharmacology-based method to determine the potential therapeutic pathways of phytoconstituents of UHPLC-PDA standardized stigma-based Crocus sativus extract for the management of PCOS revealed that the apocarotenoids and carotenoidal could act on various targets to regulate multiple pathways related to PCOS.
  • Evidence snippets:
  • Snippet 1 (score: 0.731)
    > The target protein name of the active ingredient was converted to the standard target gene name using the UniProt Knowledge Base (UniProtKB). UniProt KB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. The target protein names were uploaded into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol. The potential targets obtained from the UniproKB are depicted in Figures 3 and 4.

[5] Identifying orthologs with OMA: A primer

  • Authors: Monique Zahn-Zabal, C. Dessimoz, Natasha M. Glover
  • Year: 2020
  • Venue: F1000Research
  • URL: https://www.semanticscholar.org/paper/3b77eadcdd6979352c81d0876b0ed3a3ef4215d6
  • DOI: 10.12688/f1000research.21508.1
  • PMID: 32089838
  • PMCID: 7014581
  • Citations: 40
  • Summary: This Primer is organized in two parts and provides all the necessary background information to understand the concepts of orthology, how to infer them and the different subtypes of Orthology in OMA, as well as what types of analyses they should be used for.
  • Evidence snippets:
  • Snippet 1 (score: 0.727)
    > Get more information about your gene. After searching for your gene, you will be taken to the gene's page, which provides some external information. You can also find this by clicking on the Information tab. The information for our example gene, which corresponds to the human protein S100 calcium binding protein P, is shown in Figure 5. The information page includes the OMA ID, description, organism, locus, other IDs and cross-reference, domain architectures, and Gene Ontology annotations.

[6] Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir

  • Authors: Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al.
  • Year: 2025
  • Venue: Molecular Ecology
  • URL: https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  • DOI: 10.1111/mec.70012
  • PMID: 40613337
  • PMCID: 12288799
  • Citations: 3
  • Summary: It is hypothesized that the combination of down‐ and up‐regulated immune gene expression may prevent overstimulation of the immune response, acting as an adaptation in grey seals to resist IAV‐associated mortality.
  • Evidence snippets:
  • Snippet 1 (score: 0.721)
    > Top hits were required to have a percent query coverage (QC) ≥ 80 to be used for annotating transcripts. A subsequent blastx search against the Swiss-Prot database (downloaded from NCBI 07/02/2021) for transcripts without a sufficient hit was conducted using the parameters max_target_seqs 2, max_hsps 1, e-value 0.001 and qcov_hsp_perc 80. Genes without a published gene symbol (named 'LOC' + Gene ID in NCBI's database) were assigned a UniProt gene symbol based on the protein annotation listed in the RefSeq entries, when possible. Ultimately, a list of transcript identifiers and gene symbols were compiled into a transcript-to-gene map for subsequent analyses.
    > To further facilitate analyses of gene functions, additional steps were taken to identify the putative function of genes annotated with a symbol that began with 'LOC'. First 'LOC' genes with a protein-coding gene description in NCBI's database were manually assigned the appropriate gene symbol. Transcript sequences of remaining 'LOC' genes were queried through a blastn search against NCBI's nucleotide database (parameters: max target seq = 2, max hsps = 1, evalue = 0.01, perc identity = 90). Results from the blastn search were filtered to exclude hits with query coverage < 90 and hits that included vague terms (e.g., 'uncharacterized', 'genome assembly' and 'chromosome'). Gene symbols were extracted from the first hit for each transcript and assigned as the identity of that gene for genes with only a single transcript with hits or with consistent hits across transcripts. For genes with multiple transcripts that matched different gene symbols, the annotation was manually determined based on the number of transcripts for each gene symbol hit and query coverage/% identity values. For any 'LOC' genes that were identified as a gene already present in the dataset, gene counts were concatenated for further analysis.

[7] Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data

  • Authors: H. Chiba, Hiroyo Nishide, I. Uchiyama
  • Year: 2015
  • Venue: PLoS ONE
  • URL: https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  • DOI: 10.1371/journal.pone.0122802
  • PMID: 25875762
  • PMCID: 4395280
  • Citations: 14
  • Summary: The ortholog database using the Semantic Web technology can contribute to biological knowledge discovery through integrative data analysis and examples demonstrate that the ortholog information described in RDF can be used to link various biological data such as taxonomy information and Gene Ontology.
  • Evidence snippets:
  • Snippet 1 (score: 0.718)
    > A typical use of an ortholog database is transferring functional annotations from known genes in model organisms to genes of unknown function in other organisms, on the basis of the conjecture that orthologs are usually functionally conserved. To demonstrate such an application in our database, we showed a query to retrieve ortholog information of a specified protein.
    > Here, we specified a UniProt ID to obtain ortholog information. For describing functional categories of genes, we used Gene Ontology (GO) [24]. The UniProt GO Annotation (UniProt-GOA) database [25] (http://www.ebi.ac.uk/GOA) provides GO term assignment to proteins with evidence codes (http://www.geneontology.org/GO.evidence.shtml). We created an ontology for GO annotation (GOA-O, Table 1, http://purl.jp/bio/11/goa) and described UniProt-GOA data in RDF using it (Table 2). If some model organisms have experimentally verified GO annotations, we can transfer such a validated annotation to orthologs of other organisms.

[8] PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse

  • Authors: P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al.
  • Year: 2011
  • Venue: Nucleic Acids Research
  • URL: https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  • DOI: 10.1093/nar/gkr1122
  • PMID: 22135298
  • PMCID: 3245126
  • Citations: 1605
  • Influential citations: 155
  • Summary: PhosphoSitePlus (http://www.phosphosite.org) is an open, comprehensive, manually curated and interactive resource for studying experimentally observed post-translational modifications, primarily of human and mouse proteins. It encompasses 1 30 000 non-redundant modification sites, primarily phosphorylation, ubiquitinylation and acetylation. The interface is designed for clarity and ease of navigation. From the home page, users can launch simple or complex searches and browse high-throughput d...
  • Evidence snippets:
  • Snippet 1 (score: 0.712)
    > Accession numbers from UniPROT KB, NCBI and Ensembl (24)(25)(26), as well as gene symbols from HGNC (27), are curated for all proteins when possible. Basic protein descriptions include information parsed from UniPROT KB (24), and may include additional information from the literature. Descriptions are updated in bulk occasionally. Gene Ontology (28) annotations are parsed from NCBI (25). Editors assign protein types.

[9] CRONOS: the cross-reference navigation server

  • Authors: Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al.
  • Year: 2008
  • Venue: Bioinformatics
  • URL: https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  • DOI: 10.1093/bioinformatics/btn590
  • PMID: 19010804
  • PMCID: 2638938
  • Citations: 20
  • Summary: CRONOS, a cross-reference server that contains entries from five mammalian organisms presented by major gene and protein information resources, is developed, which shows that the cross-references are highly accurate.
  • Evidence snippets:
  • Snippet 1 (score: 0.704)
    > In order to detect gene and protein names which are assigned to products of different genes and thus result in erroneous cross-references, dedicated lists are created for each organism separately. Organism-specific lists are necessary, since terms that are ambiguous in one organism might be explicit in another. For example, ADORA2 is an ambiguous gene name in Homo sapiens but not in mouse, and GALT in mouse but not in H.sapiens.
    > In a first step, ambiguous names within the databases were extracted. If a name occurs in at least two entries describing different genes or proteins (splice variants count as one gene/protein), this particular name is marked as ambiguous and is excluded from the mapping process. In a second step, corresponding gene names occurring in the manually annotated sections of RefSeq as well as in UniProt were analyzed. Entries containing the same gene product name and having a one-to-many or many-to-many relation (e.g. one Swiss-Prot entry maps to many RefSeq entries) were scrutinized for misleading annotation. This process is done manually by inspecting additional information like sequence similarity or functional information about the involved entries. In most of the cases, the exclusion of the ambiguous gene names resulted in correct one-to-one relations.
    > As statistical analysis revealed (Supplementary Material S2) that gene names with less than four letters are exceptionally error-prone, only gene names with at least four letters are kept for mapping purposes. However, gene names with less than four letters can be queried, e.g. a search for the tumor suppressor 'p53' reveals the respective entries with the official gene name 'TP53'. Organism-specific lists of ambiguous gene and protein names are available for download on the CRONOS home page.

[10] The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa

  • Authors: P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al.
  • Year: 2025
  • Venue: Animals : an Open Access Journal from MDPI
  • URL: https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  • DOI: 10.3390/ani15040484
  • PMID: 40002966
  • PMCID: 11852025
  • Citations: 2
  • Summary: Differences in surface proteins between X- and Y-chromosome-bearing bovine spermatozoa are explored to identify potential targets for sperm sexing by LC-MS/MS analysis, with 5 transmembrane proteins showing promise as markers for X-sperm.
  • Evidence snippets:
  • Snippet 1 (score: 0.691)
    > The protein sequences were functionally annotated by combining information retrieved from the UniProt database ( [25], accessed on 9 June 2022) and one-to-one fast orthology assignments using the eggNOG-mapper v.2.1.7 tool ( [26], accessed on 9 June 2022), as described in [27]. Briefly, the gene name, protein name, length, Gene Ontology (GO) IDs, and chromosome associated with each protein entry were obtained from UniProt using the retrieve/ID mapping tool. Additionally, gene names, descriptions, and experimentally validated GO IDs were obtained from eggNOG, with consideration given to a taxonomic scope auto-adjusted per query, a minimum hit bit-score of 60, and thresholds of 80% for identity, minimum query coverage, and minimum subject coverage.
    > Out of the 130 detected proteins, 71 (54.6%) had information manually verified by UniProt curators (Supplementary Spreadsheet S1.5). Utilizing the eggNOG-mapper v2.1.7 tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a descrip-tion, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.
    > tool, a total of 122 entries were scanned (Supplementary Spreadsheet S1.6). By combining data from both tools, a total of 123 proteins were characterized with a gene name, and 127 had GO information. However, 2 proteins still lacked information on a description, protein, gene, and preferred names. Figure 1 provides a summary of the functional annotation results.

[11] Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging

  • Authors: Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al.
  • Year: 2020
  • Venue: Evidence-based Complementary and Alternative Medicine : eCAM
  • URL: https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  • DOI: 10.1155/2020/8508491
  • PMID: 32802136
  • PMCID: 7403930
  • Citations: 8
  • Summary: The study found that flavonoids (quercetin, luteolin, and kaempferol) and beta-sitosterol and the top eight candidate targets, namely, PTGS2, PPARG, DPP4, GSK3B, CCNA2, AR, MAPK14, and ESR1, were selected as the main therapeutic targets of EGb.
  • Evidence snippets:
  • Snippet 1 (score: 0.688)
    > e validated target proteins of the active components were obtained from the TCMSP database. e target protein name of the active ingredient was converted to the standard target gene name through the UniProt Knowledge Base (UniProtKB, http://www. uniprot.org/). e UniProtKB is the central hub for the collection of functional information on proteins, with accurate, consistent, and rich annotation. e target protein names were inputted into UniProtKB, with the organism restricted to "Homo sapiens," eventually gaining the official symbol.

[12] The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research

  • Authors: Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al.
  • Year: 2021
  • Venue: Insects
  • URL: https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  • DOI: 10.3390/insects12070626
  • PMID: 34357286
  • PMCID: 8307976
  • Citations: 41
  • Influential citations: 2
  • Summary: It is shown that the Ag100Pest Initiative will greatly expand the diversity of publicly available arthropod genome assemblies and demonstrate the high quality of preliminary contig assemblies, which should help other researchers attain similarly high-quality assemblies.
  • Evidence snippets:
  • Snippet 1 (score: 0.687)
    > Structural annotation refers to the prediction of gene structures on a genome assembly, including the positions of transcripts, exons, introns, coding sequences, and other features [49]. Functional annotation provides information about the gene's biological role(s), for example, gene ontologies [50], pathways, functional domains, and names. Model organism databases can manually assign biological function to genes by accumulating evidence from the scientific literature and structuring it in human and machine-readable formats. In contrast, for non-model organisms such as those in the Ag100Pest Initiative, most, if not all, functional annotation is performed computationally, as (1) gene function in very few genes have been established experimentally for these non-model species, and (2) the capacity for literature-based curation of gene function does not yet exist for these species.
    > Most of the genome assemblies generated by the Ag100Pest project are being annotated using the NCBI eukaryotic annotation pipeline [51]. This pipeline relies on Gnomon [52] for gene prediction and uses genome assembly, RNA sequencing (RNA-Seq) alignments, transcripts, and protein alignments as inputs. The resulting gene predictions are given an accession number and made publicly available. Gene names are assigned based on homology to proteins in SwissProt [53,54]. The NCBI eukaryotic annotation pipeline requires both the genome assembly and associated RNA-Seq evidence to be publicly available in the NCBI's GenBank and Sequence Read Archive, respectively (SRA; see [55]). In the event that an Ag100Pest species lacks sufficient RNA-Seq evidence in SRA, additional data will be generated, as appropriate, and submitted to aid with NCBI gene structure prediction and annotation.
    > NCBI does not currently generate additional functional annotations. Proteins deposited in GenBank or generated by RefSeq should eventually be functionally annotated by UniProt [53].

[13] GeneTools – application for functional annotation and statistical hypothesis testing

  • Authors: V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al.
  • Year: 2006
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  • DOI: 10.1186/1471-2105-7-470
  • PMID: 17062145
  • PMCID: 1630634
  • Citations: 105
  • Influential citations: 11
  • Summary: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once.
  • Evidence snippets:
  • Snippet 1 (score: 0.687)
    > The database enables searching by gene symbols/names, GenBank accession numbers, UniGene cluster IDs, Swiss-Prot entry names and several unique clone IDs (IMAGE clone IDs, University of Iowa clone IDs, Operon oligo IDs, TAIR IDs and a subset of selected Affymetrix and Agilent IDs).
    > The names and symbols of genes/proteins may be highly ambiguous [20]. We therefore recommend using primary gene IDs, like GeneBank accession numbers or specific probe IDs when querying the database. However, if gene names or symbols are used, caution is advised because only official names/symbols associated with UniProt knowledgebase will be recognized. The underlying database is updated on a weekly basis with annotation information from several external databases including UniGene, Swiss-Prot, Entrez Gene and GO. User data are submitted to the database as text files of gene reporters and analysis of the annotation data can be performed through three user interfaces: the NMC Annotation Tool, the GO Annotator Tool and eGOn. Analysis results and annotation data can be exported in various formats.

[14] Quantitative proteomic dataset of whole protein in three melanoma samples of 92.1, 92.1-A and 92.1-B

  • Authors: Xi-feng Fei, Xiangtong Xie, X. Ji, Haiyan Tian, F. Sun et al.
  • Year: 2022
  • Venue: Data in Brief
  • URL: https://www.semanticscholar.org/paper/33312bb6cc2cf985d32f3a31cf4fdee6b4e17385
  • DOI: 10.1016/j.dib.2022.108592
  • PMID: 36164296
  • PMCID: 9508510
  • Citations: 2
  • Influential citations: 1
  • Summary: Covering differential proteomes of three cell lines in a pairwise model, the data could be used to further screen the kinesins that play a vital role in regulating the growth of UM.
  • Evidence snippets:
  • Snippet 1 (score: 0.681)
    > 2.8.1. Annotation methods 2.8.1.1. Functional annotation. UniProt-GOA database was utilized to retrieve Gene Ontology (GO) annotation proteome. First, the identified protein identity was converted to UniProt identity and then mapped to GO identity based on the protein identity. When the identified protein was not annotated by UniProt-GOA, the functional annotation of that protein was conducted using the InterProScan software according to the amino acid sequence alignment approach. All proteins were then classified into 3 groups: molecular function, cellular component and biological process.

[15] Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem

  • Authors: L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al.
  • Year: 2021
  • Venue: Frontiers in Research Metrics and Analytics
  • URL: https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  • DOI: 10.3389/frma.2021.689059
  • PMID: 34322655
  • PMCID: 8311438
  • Citations: 23
  • Influential citations: 1
  • Summary: The literature knowledge panels developed and implemented in PubChem help to uncover and summarize important relationships between chemicals, genes, proteins, and diseases by analyzing co-occurrences of terms in biomedical literature abstracts.
  • Evidence snippets:
  • Snippet 1 (score: 0.680)
    > We decided to prioritize human genes and proteins. The following strategy has been implemented to resolve gene and protein text entities to the most reasonable gene, protein, or enzyme symbol (corresponding to human, when possible):
    > -Try to find a match among Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) names (Braschi et al., 2019;HUGO, 2021);
    > -Try to find a match among names in The IUPHAR/BPS Guide to Pharmacology (Armstrong et al., 2020; IUPHAR/BPS, 2021); -Try to find matches among names in UniProt (Bateman et al., 2017); -Otherwise, try to match to an enzyme name and resolve to an EC number (Bairoch, 2000;Expassy, 2021).
    > In general, it is very difficult and often impossible to distinguish the name of a gene from the name of the protein encoded by that gene. Therefore, gene and protein names are not strictly distinguished from each other but considered as one category. Therefore, the annotations considered in this study can be grouped into three categories: chemicals, genes/proteins, and diseases.

[16] GOnet: a tool for interactive Gene Ontology analysis

  • Authors: M. Pomaznoy, Brendan Ha, Bjoern Peters
  • Year: 2018
  • Venue: BMC Bioinformatics
  • URL: https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  • DOI: 10.1186/s12859-018-2533-3
  • PMID: 30526489
  • PMCID: 6286514
  • Citations: 247
  • Influential citations: 17
  • Summary: The open-source GOnet web-application is created, which takes a list of gene or protein entries from human or mouse data and performs GO term annotation analysis and provides insight into the functional interconnection of the submitted entries.
  • Evidence snippets:
  • Snippet 1 (score: 0.678)
    > In a basic workflow, the GOnet application receives a list of gene symbols, protein symbols, or protein IDs (UniProt IDs) as an input, and outputs a graph (an example given in Fig. 1). There are various input parameters which will affect the actual structure of the graph visualized and its appearance. The first main user choice is which GO terms the genes are annotated against:
    > 1. GO terms statistically significantly over-represented in the gene list submitted. 2. A predefined subset (also known as 'GO slim'), or a user-supplied list of terms.
    > In the first case the analysis will be referred to as an 'enrichment' analysis, in the second as an 'annotation' analysis.
    > Input parameters 1) Gene list. A mandatory input parameter containing the genes/proteins of interest. Currently human and mouse data is supported. An example of a human gene list might look like this:
    > Fig. 1 Sample network output generated by GOnet application. Gene differentially expressed in CD4 Bulk Memory T cells in Latent TB patients compared to healthy controls were used as an example [22] The gene list can also be accompanied with a contrast value. For example, This contrast value can be any decimal number, such as the log-fold change of gene expression between two conditions. This is merely a visualization enhancement. If the value is supplied it can be used later to differentially color specific genes in the graph (note different colors of gene nodes in Fig. 1), and visually indicate up-or down-regulation of specific genes and gene clusters.
    > The application can process common gene symbols (like in the example above), UniProt IDs, and MGI Accession IDs (mouse only). The former type of ID (gene symbols), although is the most human friendly, can unfortunately be ambiguous. For example, AIM1 can mean 'absent in melanoma' (also called CRYBG1) or 'Aurora and Ipl1-like midbody-associated protein' (also known as AURKB). Due to this ambiguity UniProt IDs or MGI accession IDs (for mouse) are preferred.
    > 2) GO namespace. Can be any of 'biological process', 'molecular function' or 'cellular component'.

[17] CellPhoneDB: inferring cell–cell communication from combined expression of multi-subunit ligand–receptor complexes

  • Authors: M. Efremova, Miquel Vento-Tormo, S. Teichmann, R. Vento-Tormo
  • Year: 2020
  • Venue: Nature Protocols
  • URL: https://www.semanticscholar.org/paper/6a3b3e4a2eebc3fad3b03ee6d8228264247abbc8
  • DOI: 10.1038/s41596-020-0292-x
  • PMID: 32103204
  • Citations: 2770
  • Influential citations: 298
  • Summary: The structure and content of CellPhoneDB is outlined, procedures for inferring cell–cell communication networks from single-cell RNA sequencing data are provided and a practical step-by-step guide to help implement the protocol is presented.
  • Evidence snippets:
  • Snippet 1 (score: 0.677)
    > CellPhoneDB stores ligand-receptor interactions, as well as other properties of the interacting partners, including their subunit architecture and gene and protein identifiers. To create the content of the database, four main .csv data files are required: gene_input.csv, protein_input.csv, complex_input.csv and interaction_input.csv (Fig. 4).
    > gene_input Mandatory fields are 'gene_name', 'uniprot', 'hgnc_symbol' and 'ensembl'. This file is critical for establishing the link between the scRNA-seq data and the interaction pairs stored at the protein level. It includes the following gene and protein identifiers: (i) gene name ('gene_name'), (ii) UniProt identifier ('uniprot'), (iii) HUGO Nomenclature Committee (HGNC) symbol ('hgnc_symbol') and (iv) gene Ensembl identifier (ENSG) ('ensembl'). To create this file, lists of linked proteins and gene identifiers are downloaded from UniProt and merged using gene names. Several rules need to be considered when merging the files • UniProt annotation prevails over the gene Ensembl annotation when the same gene Ensembl identifier points toward different UniProt identifiers. • UniProt and Ensembl lists are also merged by their UniProt identifier, but this information is used only when the UniProt or Ensembl identifier is missing in the original list merged by gene name. • If the same gene name points toward different HGNC symbols, only the HGNC symbol matching the gene name annotation is considered. • Only one HLA isoform is considered in our interaction analysis, and it is stored in a manually HLA-curated list of genes, named HLA_curated.
    > protein_input Mandatory fields are 'uniprot' and 'protein_name'.

[18] Protein Localization Analysis of Essential Genes in Prokaryotes

  • Authors: Chong Peng, Feng Gao
  • Year: 2014
  • Venue: Scientific Reports
  • URL: https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  • DOI: 10.1038/srep06001
  • PMID: 25105358
  • PMCID: 4126397
  • Citations: 27
  • Summary: A comprehensive protein localization analysis of essential genes in 27 prokaryotes including 24 bacteria, 2 mycoplasmas and 1 archaeon has been performed and shows that proteins encoded by essential genes are enriched in internal location sites, while exist in cell envelope with a lower proportion compared with non-essential ones.
  • Evidence snippets:
  • Snippet 1 (score: 0.674)
    > Bioinformatics Databases. DEG is a database of essential genes (http://www. essentialgene.org/). The newly released DEG 10 has been developed to accommodate the quantitative and qualitative advancements brought by the progressive identification methods. Currently available records of both essential and nonessential genes among a wide range of organisms can be downloaded from DEG 10, making it possible to compare the two different types of genes in many aspects 21 .
    > 27 prokaryotic organisms including 24 bacteria, 2 mycoplasmas and Methanococcus maripaludis S2, the only record of the Archaea domain were selected to analyze the protein localization and GO distribution of the essential and nonessential genes. There are 31 bacterial records corresponding to 27 organisms in the database in total and 26 sets of data were selected in the current study. Streptococcus pneumonia was not chosen for the lack of non-essential genes. Since the essential genes were not genome-widely identified, it's not reasonable to regard the complementary set of essential genes as non-essential genes in Streptococcus pneumonia 29,30 . In the case of multiple records for one organism, the one with the most convincing experimental methods was chosen. The non-essential genes in Methanococcus maripaludis S2 and 13 bacteria such as Escherichia coli MG1655 are obtained based on the original literatures, while non-essential genes in other 12 organisms such as Bacillus subtilis 168 are the complementary set of essential genes. The information of the organisms used in the current study are displayed in Table 1.
    > The three model genomes' subcellular location information and the Gene Ontology (GO) terms used for the analysis in the current study were downloaded from the Universal Protein Resource (UniProt; http://www.uniprot.org). Maintained by the UniProt Consortium, UniProt is committed to providing biologists with a comprehensive, high-quality and freely accessible resource of protein sequences and functional annotation 27 . Among the wealth of annotation data, detailed GO annotation statements are included.

[19] MeiosisOnline: A Manually Curated Database for Tracking and Predicting Genes Associated With Meiosis

  • Authors: Xiaohua Jiang, Daren Zhao, Asim Ali, Bo Xu, Wei Liu et al.
  • Year: 2021
  • Venue: Frontiers in Cell and Developmental Biology
  • URL: https://www.semanticscholar.org/paper/aeb08368a5540685473578345739bb569103bd46
  • DOI: 10.3389/fcell.2021.673073
  • PMID: 34485275
  • PMCID: 8415030
  • Citations: 6
  • Summary: The developed MeiosisOnline provides the most updated and detailed information of experimental verified and predicted genes in meiosis, and will greatly help researchers in studying meiosis in an easy and efficient way.
  • Evidence snippets:
  • Snippet 1 (score: 0.674)
    > Annotation information for each gene in MeiosisOnline contains "basic information, " "function annotation and classification, " "protein-protein interaction (PPI) and gene expression."
    > (1) Basic information: gene name/synonyms, nucleotide sequences, etc., were extracted from GenBank3 and UniProt Knowledgebase.4 (2) Function annotation and classification: detailed functional information is also manually collected from literature reports. (i) Which meiotic stage is the gene involved? (ii) Did the gene function in one sex or both sexes? (iii) Whether deletion or mutation of the gene in model organism has a phenotype in fertility? (iv) Which protein complex of the gene is involved? (v) The cellular location and expression pattern in tissues or cell lines.
    > (vi) Experimental methods used for functional analysis.
    > (vii) The information of related literature and figures for illustrating the function of protein/gene. (viii) Gene ontology annotation for collected genes.
    > (3) Protein-protein interaction and gene expression: both verified and predicted PPI information were provided. Gene expression pattern in reproductive system was also provided graphically.

[20] The alpha-ketoacid dehydrogenase complexes of Drosophila melanogaster.

  • Authors: Steven J. Marygold
  • Year: 2024
  • Venue: microPublication Biology
  • URL: https://www.semanticscholar.org/paper/50942e603e0e14ee9195c0d7cb52db11a521f964
  • DOI: 10.17912/micropub.biology.001209
  • PMID: 38741935
  • PMCID: 11089389
  • Citations: 2
  • Summary: This work identifies and classify the genes encoding all Drosophila AKDHC subunits, update their functional annotations and integrate this work into the FlyBase database.
  • Evidence snippets:
  • Snippet 1 (score: 0.666)
    > Symbol: gene symbol in FlyBase -asterisk (*) indicates a gene with testis-specific expression. CG#: gene annotation ID in FlyBase. Function: component and associated EC number (where available/applicable). Key reference(s) for initial identification or genetic characterization (in a metabolic context): 1. Gruntenko et al. 1998;2. Chen et al. 2008;3. Yoon et al. 2017;4. Yap et al. 2021a;5. Yap et al. 2021b;6. Whittle et al. 2023;7. González Morales et al. 2023;8. Homem et al. 2014;9. Bonnay et al. 2020;10. Ivanova et al. 2004;11. Boyko et al. 2020;12. Liu et al. 2017;13. Li et al. 2020;14. Devilliers et al. 2021;15. Goyal et al. 2022;16. Huang et al. 2022;17. Plaçais et al. 2017;18. Dung et al. 2018;19. Rabah et al. 2023;20. Klenz et al. 1995;21. Katsube et al. 1997;22. Gándara et al. 2019;23. Lambrechts et al. 2019;24. Lee et al. 2022;25. Chen et al. 2006;26. Kim et al. 2023;27. Tsai et al. 2020. Human ortholog: gene symbol at HGNC, with % amino acid identity between the encoded protein and the Drosophila protein. Human disease: OMIM symbol for disease(s) associated with the human gene (Amberger et al. 2019) -see Extended Data File 1 for details.

Notes

  • This provider combines search_papers_by_relevance with snippet_search.
  • No synthesis or second-stage model call is performed.

Citations

  1. Dawn Cotter, A. Maer, C. Guda, Brian Saunders, S. Subramaniam (2005). LMPD: LIPID MAPS proteome database. Nucleic Acids Research. https://www.semanticscholar.org/paper/265c37b45326b7927e396484751e84e4aeff92d5
  2. Ralf C. Mueller, Nicolai Mallig, Jacqueline Smith, Lél Eöry, Richard I. Kuo et al. (2020). Avian Immunome DB: an example of a user-friendly interface for extracting genetic information. BMC Bioinformatics. https://www.semanticscholar.org/paper/b894d9ca8ea2d653bf1711a0c67dab71d054487c
  3. Samuel J. Modlin, A. Elghraoui, Deepika Gunasekaran, Alyssa M Zlotnicki, N. Dillon et al. (2021). Structure-Aware Mycobacterium tuberculosis Functional Annotation Uncloaks Resistance, Metabolic, and Virulence Genes. mSystems. https://www.semanticscholar.org/paper/76ff9a62b36b32cc10e46e71ffd4dd90344e4706
  4. Anshul Tiwari, Siddharth J Modi, A. Girme, L. Hingorani (2023). Network pharmacology-based strategic prediction and target identification of apocarotenoids and carotenoids from standardized Kashmir saffron (Crocus sativus L.) extract against polycystic ovary syndrome. Medicine. https://www.semanticscholar.org/paper/3e3253804574634d1968a0fd5b65dd1674bff6c6
  5. Monique Zahn-Zabal, C. Dessimoz, Natasha M. Glover (2020). Identifying orthologs with OMA: A primer. F1000Research. https://www.semanticscholar.org/paper/3b77eadcdd6979352c81d0876b0ed3a3ef4215d6
  6. Christina M McCosker, E. Unal, Alayna K Gigliotti, Wendy B Puryear, Jonathan A. Runstadler et al. (2025). Molecular Mechanisms Underlying Response to Influenza in Grey Seals (Halichoerus grypus), a Potential Wild Reservoir. Molecular Ecology. https://www.semanticscholar.org/paper/bebb135aae1c1182d098fce839c9a3df0cfb2b21
  7. H. Chiba, Hiroyo Nishide, I. Uchiyama (2015). Construction of an Ortholog Database Using the Semantic Web Technology for Integrative Analysis of Genomic Data. PLoS ONE. https://www.semanticscholar.org/paper/7cc805575c642aa8efdc1204383a7662965fbb60
  8. P. Hornbeck, J. Kornhauser, S. Tkachev, Bin Zhang, E. Skrzypek et al. (2011). PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Research. https://www.semanticscholar.org/paper/93e4acf4ba3bca1b379ae8292e73dddb344abd90
  9. Brigitte Waegele, I. Dunger, G. Fobo, Corinna Montrone, H. Mewes et al. (2008). CRONOS: the cross-reference navigation server. Bioinformatics. https://www.semanticscholar.org/paper/8c05b3aa0ba01c41ee97c2dc98ea7b5b14ce0e9c
  10. P. Pinto-Pinho, Joana Quelhas, Francis Impens, Sara Dufour, Delphi Van Haver et al. (2025). The Surface Proteome of Bovine Unsexed and Sexed Spermatozoa. Animals : an Open Access Journal from MDPI. https://www.semanticscholar.org/paper/66496158140e8f55a2c2ca8965bd298300ce9ab0
  11. Yanfei Liu, Yue Liu, Wantong Zhang, Mingyue Sun, Weiliang Weng et al. (2020). Network Pharmacology-Based Strategy to Investigate the Pharmacological Mechanisms of Ginkgo biloba Extract for Aging. Evidence-based Complementary and Alternative Medicine : eCAM. https://www.semanticscholar.org/paper/fadeff691eb41dff82e019cc3d38b846fdc605c4
  12. Anna K. Childers, S. Geib, S. Sim, Monica F. Poelchau, B. Coates et al. (2021). The USDA-ARS Ag100Pest Initiative: High-Quality Genome Assemblies for Agricultural Pest Arthropod Research. Insects. https://www.semanticscholar.org/paper/a33d31da6f5aee18501cfc2332ff50d5b7f23508
  13. V. Beisvåg, Frode K. R. Jünge, Hallgeir Bergum, Lars Jølsum, S. Lydersen et al. (2006). GeneTools – application for functional annotation and statistical hypothesis testing. BMC Bioinformatics. https://www.semanticscholar.org/paper/1d9e0c2f67acd5bf64c659f1f3f8624325b6be8a
  14. Xi-feng Fei, Xiangtong Xie, X. Ji, Haiyan Tian, F. Sun et al. (2022). Quantitative proteomic dataset of whole protein in three melanoma samples of 92.1, 92.1-A and 92.1-B. Data in Brief. https://www.semanticscholar.org/paper/33312bb6cc2cf985d32f3a31cf4fdee6b4e17385
  15. L. Zaslavsky, Tiejun Cheng, A. Gindulyte, Siqian He, Sunghwan Kim et al. (2021). Discovering and Summarizing Relationships Between Chemicals, Genes, Proteins, and Diseases in PubChem. Frontiers in Research Metrics and Analytics. https://www.semanticscholar.org/paper/57b86aef9aae576c2ae4199c0b74971f4c195211
  16. M. Pomaznoy, Brendan Ha, Bjoern Peters (2018). GOnet: a tool for interactive Gene Ontology analysis. BMC Bioinformatics. https://www.semanticscholar.org/paper/d984d075cb08f6c39a48bcf1f32e36f333a423d9
  17. M. Efremova, Miquel Vento-Tormo, S. Teichmann, R. Vento-Tormo (2020). CellPhoneDB: inferring cell–cell communication from combined expression of multi-subunit ligand–receptor complexes. Nature Protocols. https://www.semanticscholar.org/paper/6a3b3e4a2eebc3fad3b03ee6d8228264247abbc8
  18. Chong Peng, Feng Gao (2014). Protein Localization Analysis of Essential Genes in Prokaryotes. Scientific Reports. https://www.semanticscholar.org/paper/69181762648fd77a085b2f93618a71b43b62cf76
  19. Xiaohua Jiang, Daren Zhao, Asim Ali, Bo Xu, Wei Liu et al. (2021). MeiosisOnline: A Manually Curated Database for Tracking and Predicting Genes Associated With Meiosis. Frontiers in Cell and Developmental Biology. https://www.semanticscholar.org/paper/aeb08368a5540685473578345739bb569103bd46
  20. Steven J. Marygold (2024). The alpha-ketoacid dehydrogenase complexes of Drosophila melanogaster.. microPublication Biology. https://www.semanticscholar.org/paper/50942e603e0e14ee9195c0d7cb52db11a521f964

Falcon

(tyrB-deep-research-falcon.md)
Research report: **tyrB (PP_1972; UniProt Q88LG1)** in *Pseudomonas putida* KT2440 Falcon Edison Scientific Literature 19 citations 3 artifacts 2026-06-11T22:05:14.247139

The research report should be a detailed narrative explaining the function, biological processes, and localization of the gene product. Citations should be given for all claims.

You should prioritize authoritative reviews and primary scientific literature when conducting research. You can supplement
this with annotations you find in gene/protein databases, but these can be outdated or inaccurate.

We are specifically interested in the primary function of the gene - for enzymes, what reaction is catalyzed, and what is the substrate specificity? For transporters, what is the substrate? For structural proteins or adapters, what is the broader structural role? For signaling molecules, what is the role in the pathway.

We are interested in where in or outside the cell the gene product carries out its function.

We are also interested in the signaling or biochemical pathways in which the gene functions. We are less interested in broad pleiotropic effects, except where these elucidate the precise role.

Include evidence where possible. We are interested in both experimental evidence as well as inference from structure, evolution, or bioinformatic analysis. Precise studies should be prioritized over high-throughput, where available.

Research report: tyrB (PP_1972; UniProt Q88LG1) in Pseudomonas putida KT2440

0) Target verification (gene/protein identity)

The requested UniProt entry Q88LG1 corresponds to Pseudomonas putida KT2440 gene tyrB with locus tag PP_1972. In KT2440-focused literature analyzing aromatic amino acid catabolism, PP1972 is explicitly referred to as tyrB-1, while a paralogous aminotransferase gene (tyrB-2) is associated with PP_3590 (also discussed as amaC in later work), confirming that the report targets the correct gene/protein in the correct organism/strain (ATCC 47054/KT2440) and not a different “tyrB” from another organism. (herrera2010identificationandcharacterization pages 1-2, borchert2024machinelearninganalysis pages 7-11)

1) Key concepts and definitions (current understanding)

1.1 What “TyrB” typically denotes in bacteria

In many bacteria, TyrB denotes a PLP-dependent aminotransferase that catalyzes reversible transamination reactions, transferring an amino group between an amino acid and an α-keto acid. In Pseudomonas, enzymes annotated as TyrB-family aminotransferases have been studied primarily in the context of aromatic amino acid transformations, i.e., interconversion between aromatic amino acids (phenylalanine/tyrosine/tryptophan) and their corresponding aromatic 2-oxoacids. (szkop2013tyrb2andphhc pages 1-2)

1.2 General PLP-aminotransferase chemistry (mechanism-level view)

PLP-dependent transaminases operate through two half-reactions in which the cofactor cycles between pyridoxal-5′-phosphate (PLP) and pyridoxamine phosphate (PMP), with key intermediates (external aldimine, quinonoid, ketimine) formed during amino-group transfer. A frequent structural theme is an oligomeric enzyme (often a homodimer) with active sites formed by residues from both subunits. (menke2024proteinengineeringof pages 22-25, menke2024proteinengineeringof pages 25-28)

This mechanism-level understanding is important for functional annotation because it explains (i) why PLP is required, (ii) why α-keto acids such as 2-oxoglutarate or pyruvate commonly serve as amino acceptors, and (iii) why these enzymes can show substrate promiscuity that complicates gene-to-function assignment by annotation alone. (menke2024proteinengineeringof pages 22-25, menke2024proteinengineeringof pages 25-28)

2) Molecular function of PP_1972 (tyrB; Q88LG1)

2.1 Enzyme class and likely reaction

Direct purified-enzyme biochemistry for PP_1972/Q88LG1 was not found in the retrieved full texts; however, multiple KT2440 studies place PP_1972 among aromatic/tyrosine aminotransferase-like genes and test its role genetically. (herrera2010identificationandcharacterization pages 1-2, herrera2010identificationandcharacterization pages 4-5, borchert2024machinelearninganalysis pages 7-11)

The reaction class most consistent with the TyrB annotation in this KT2440 context is an aromatic amino acid transamination such as:

  • L-tyrosine + 2-oxoglutarate ⇌ 4-hydroxyphenylpyruvate + L-glutamate
  • L-phenylalanine + 2-oxoglutarate ⇌ phenylpyruvate + L-glutamate

Support for this reaction type comes from protein-level characterization of closely related P. putida aromatic aminotransferases (encoded by tyrB-2 and phhC), which preferentially catalyze transamination involving aromatic amino acids and aromatic 2-oxoacids, with PLP included as cofactor and 2-oxoglutarate used as amino acceptor in assays. (szkop2013tyrb2andphhc pages 1-2, szkop2013tyrb2andphhc media f8f3824d)

2.2 Substrate specificity (direct evidence from P. putida aromatic aminotransferases)

While the enzyme characterized biochemically in P. putida was not PP_1972, Table 2 from Szkop & Bielawski (2013) provides detailed substrate profiles for two P. putida aromatic aminotransferase isozymes, showing that they most efficiently catalyze reactions involving aromatic amino acids and aromatic 2-oxoacids, with L-phenylalanine and phenylpyruvate being the best substrates reported. Assays used 0.1 M phosphate buffer (pH 8.0), 3 mM 2-oxoglutarate, 10 µM PLP, and 35 °C—conditions consistent with fold-type I PLP aminotransferase enzymology. (szkop2013tyrb2andphhc pages 1-2, szkop2013tyrb2andphhc media f8f3824d)

These data support that TyrB-like enzymes in P. putida are plausibly aromatic aminotransferases rather than strictly tyrosine-specific enzymes, and that substrate promiscuity and isozyme redundancy should be expected in vivo. (szkop2013tyrb2andphhc pages 1-2, szkop2013tyrb2andphhc media f8f3824d)

3) Biological role and pathways in P. putida KT2440

3.1 Phenylalanine utilization routes and PhhR-controlled regulon

A central KT2440 pathway context is phenylalanine assimilation/catabolism regulated by PhhR, which induces the phhAB operon encoding a pterin-dependent phenylalanine hydroxylase system (PhhA catalytic hydroxylase; PhhB pterin cofactor regeneration). This provides a route for conversion of L-phenylalanine to L-tyrosine. (herrera2010identificationandcharacterization pages 1-2)

Herrera et al. (2010) place phenylalanine/tyrosine degradation into a larger catabolic funnel via p-hydroxyphenylpyruvate and homogentisate (hpd/hmg genes), which then feeds into central metabolism. (herrera2010identificationandcharacterization pages 4-5)

3.2 Genetic evidence for tyrB-family involvement in phenylalanine/tyrosine catabolism

In KT2440, mutants in TyrB-like genes exhibit measurable growth phenotypes on aromatic amino acids as nitrogen sources.

  • Wild-type KT2440 doubling time: ~8 h on phenylalanine; ~1.8 h on tyrosine.
  • tyrB-1 (PP_1972) mutant: phenylalanine growth near WT, but slower on tyrosine (~3.2 h).
  • tyrB-2 mutant: slower on phenylalanine (~12 h) and slower on tyrosine (~3.0 h).

This pattern supports that TyrB-family enzymes contribute to aromatic amino acid utilization, with a stronger phenotype for tyrB-2 in phenylalanine conditions in this dataset. (herrera2010identificationandcharacterization media d9fda959, herrera2010identificationandcharacterization pages 4-5)

3.3 Redundancy and reassignment of “primary” aromatic aminotransferase (major recent development)

A key development in 2024 is that machine-learning analysis of RB-TnSeq fitness compendia (ICA-derived “fModules”) was used to pinpoint genes involved in phenylalanine/tyrosine catabolism, followed by mutant validation.

Borchert et al. (2024, published March 2024; https://doi.org/10.1128/msystems.00942-23) report that disruption of tyrB (PP_1972) did not inhibit growth on L-phenylalanine or L-tyrosine as sole nitrogen sources, whereas disruption of amaC (PP_3590; sometimes called tyrB2) completely abrogated growth on these substrates. The authors therefore propose re-annotation of AmaC (PP_3590) as an L-tyrosine aminotransferase, implying that PP_1972 is not the primary enzyme for these growth phenotypes under the tested conditions and that previous “tyrB” annotation may overstate its physiological importance. (borchert2024machinelearninganalysis pages 7-11)

This reconciles earlier BarSeq-based observations that PP_1972 often shows weak fitness effects and that a PP_3590/PP_1972 double knockout did not cause phenylalanine auxotrophy, consistent with broader redundancy or alternative routes, while still allowing PP_1972 to contribute in specific environments or regulatory states. (schmidt2022nitrogenmetabolismin pages 8-10, schmidt2022nitrogenmetabolismin pages 10-12)

4) Cellular localization

No retrieved KT2440 primary source in this corpus provided a direct experimental localization (cytosol/periplasm) for PP_1972/Q88LG1.

For context on what localization evidence looks like for bacterial PLP-dependent aminotransferases, Ringel et al. (2017) show that a distinct periplasmic PLP-dependent transaminase (PtaA) can be demonstrated by subcellular fractionation and is a homodimer by SEC-MALS; however, this is a different enzyme in a different Pseudomonas species and should not be taken as evidence that PP_1972 is periplasmic. (ringel2017theperiplasmictransaminase pages 16-18)

Current best-supported statement from the retrieved KT2440 corpus: localization of PP_1972 remains unresolved here and should be taken from UniProt/InterPro experimental annotations if available, or validated experimentally.

5) Current applications and real-world implementations

5.1 Functional genomics for strain engineering and annotation (2024)

The 2024 RB‑TnSeq + ICA framework provides a practical route to re-annotate metabolic genes and identify engineering targets for P. putida as a chassis. In particular, the ability to distinguish PP_1972 (tyrB) from PP_3590 (AmaC) as the functionally dominant aminotransferase for phenylalanine/tyrosine utilization under defined conditions is directly actionable for (i) redirecting aromatic amino acid flux, and (ii) avoiding incorrect knockouts when designing production strains. (borchert2024machinelearninganalysis pages 7-11)

5.2 Engineering relevance via aromatic-stress tolerance and aromatic feedstocks

Borchert et al. (2024) also connect aminotransferase-linked modules to tolerance phenotypes during growth with high concentrations of hydroxycinnamates (e.g., ~60 mM in glucose + hydroxycinnamate tests; and higher concentrations as carbon sources in some conditions). These results highlight that aromatic amino acid and aromatic acid metabolism genes can have roles in stress tolerance, a key trait for industrial bioprocessing on lignin-derived aromatics. (borchert2024machinelearninganalysis pages 7-11)

5.3 Broader industrial deployment of PLP-transaminases (expert review)

Menke (2024) reviews industrial and engineering aspects of PLP-dependent amine transaminases (ATAs), emphasizing (i) their use in stereoselective synthesis of chiral amines, (ii) the importance of addressing unfavorable equilibria (e.g., via coproduct removal or sacrificial donors), and (iii) the rise of machine-learning-guided engineering. The review notes an industrial benchmark: (R)-ATA-catalyzed synthesis of (R)-sitagliptin with >99.95% optical purity, and also reports large engineering gains (e.g., up to 2000-fold improved catalytic activity in one redesign example), illustrating the real-world value of understanding transaminase substrate specificity and engineering it rationally. (menke2024proteinengineeringof pages 61-71)

Although this review is not specific to PP_1972, it provides authoritative context for why TyrB-family enzymes and related aminotransferases are frequently targeted in metabolic engineering and biocatalysis.

6) Summary of key quantitative evidence (recent and classic)

  1. Mutant growth phenotypes (KT2440; 2010): doubling times on phenylalanine and tyrosine as sole N sources show measurable aromatic-amino-acid-related phenotypes for tyrB mutants, with a stronger phenotype for tyrB-2 in phenylalanine conditions. (herrera2010identificationandcharacterization media d9fda959)
  2. Functional genomics/fitness and redundancy (2022): PP_1972 shows weak fitness effects in some aromatic N-source conditions and is genetically redundant with other aminotransferases in some contexts. (schmidt2022nitrogenmetabolismin pages 8-10, schmidt2022nitrogenmetabolismin pages 10-12)
  3. Reannotation-driven experimental validation (2024): PP_1972 knockout does not inhibit growth on L-Phe/L-Tyr as sole N sources, while PP_3590 (AmaC) knockout abolishes growth, supporting updated functional assignment of the dominant aromatic aminotransferase in those conditions. (borchert2024machinelearninganalysis pages 7-11)
  4. Biochemical assay conditions and substrate profiles (2013): P. putida aromatic aminotransferase isozymes show strongest activity for aromatic amino acids/2-oxoacids (e.g., L-Phe/phenylpyruvate) in PLP + 2-oxoglutarate-dependent assays. (szkop2013tyrb2andphhc media f8f3824d)

7) Practical functional-annotation conclusion for PP_1972 (tyrB; Q88LG1)

Most defensible functional annotation from the retrieved evidence:

  • Gene/protein: tyrB (PP_1972; UniProt Q88LG1) in P. putida KT2440.
  • Molecular function: PLP-dependent aminotransferase in the aromatic/tyrosine aminotransferase-like set of enzymes in Pseudomonas; likely catalyzes reversible transamination between aromatic amino acids (e.g., tyrosine/phenylalanine) and α-keto acids (commonly 2-oxoglutarate), but direct biochemical constants for PP_1972 are not established in the retrieved corpus. (herrera2010identificationandcharacterization pages 1-2, szkop2013tyrb2andphhc media f8f3824d)
  • Biological process/pathway: participates (directly or redundantly) in aromatic amino acid utilization/catabolism connected to phenylalanine → tyrosine conversion (PhhAB) and downstream degradation via p-hydroxyphenylpyruvate/homogentisate. Genetic data indicate PP_1972 contributes measurably to tyrosine growth in some conditions but is not the primary determinant of Phe/Tyr utilization in others. (herrera2010identificationandcharacterization pages 4-5, herrera2010identificationandcharacterization media d9fda959, borchert2024machinelearninganalysis pages 7-11)
  • Systems-level role: part of a redundant aminotransferase network in KT2440; recent (2024) work indicates that PP_3590 (AmaC) dominates growth phenotypes on Phe/Tyr as N sources, and PP_1972 disruption alone may not block those pathways under standard minimal-media conditions. (borchert2024machinelearninganalysis pages 7-11, schmidt2022nitrogenmetabolismin pages 10-12)
  • Cellular localization: not resolved from retrieved KT2440 evidence; should be verified experimentally or via curated UniProt localization annotations. (ringel2017theperiplasmictransaminase pages 16-18)

8) Evidence table (quick reference)

Aspect Evidence summary Key quantitative data
Target identity UniProt Q88LG1 corresponds to tyrB / PP_1972 in Pseudomonas putida KT2440; genome annotation in KT2440 literature lists PP1972 as tyrB-1, one of two tyrosine/aromatic aminotransferase-like genes in this strain (herrera2010identificationandcharacterization pages 1-2, herrera2010identificationandcharacterization pages 9-10) Locus tags/names reported as PP1972 / tyrB-1; paralog also noted as PP3590 / tyrB-2 (herrera2010identificationandcharacterization pages 1-2)
Predicted molecular function/class TyrB/PP_1972 is an aminotransferase in the PLP-dependent aromatic amino acid aminotransferase class; related P. putida aromatic aminotransferases preferentially transaminate aromatic amino acids with 2-oxoglutarate, with best substrates including L-phenylalanine and phenylpyruvate (szkop2013tyrb2andphhc pages 2-4, szkop2013tyrb2andphhc pages 1-2) Assays for related P. putida ArAT enzymes used 10 µM PLP and 3 mM 2-oxoglutarate; activity measured as release of 1 µmol IPyA min⁻¹ in L-tryptophan:2-oxoglutarate assays (szkop2013tyrb2andphhc pages 2-4)
Pathway role in aromatic amino acid metabolism In KT2440, phenylalanine can be degraded by the phenylalanine hydroxylase pathway (PhhAB → tyrosine → p-hydroxyphenylpyruvate → homogentisate), and KT2440 carries two TyrB-like aminotransferase genes. Mutant phenotypes support TyrB-family participation in phenylalanine/tyrosine catabolism, especially downstream aromatic transamination steps (herrera2010identificationandcharacterization pages 4-5, herrera2010identificationandcharacterization pages 1-2) Wild type doubling times on sole N source: phenylalanine ~8 h, tyrosine ~1.8 h; tyrB-1 mutant: phenylalanine ~WT, tyrosine ~3.2 h; tyrB-2 mutant: phenylalanine ~12 h, tyrosine ~3.0 h (herrera2010identificationandcharacterization pages 4-5, herrera2010identificationandcharacterization media d9fda959)
Evidence for redundancy Recent RB-TnSeq and prior knockout work indicate functional redundancy among aromatic aminotransferases in P. putida KT2440: PP_1972 has only weak single-gene phenotypes in some aromatic N-source conditions, and even combined loss with PP_3590 did not cause phenylalanine auxotrophy (schmidt2022nitrogenmetabolismin pages 8-10, schmidt2022nitrogenmetabolismin pages 10-12) BarSeq fitness effects for PP_1972 were small: phenylalanine -0.35 and pipecolate -0.15 in one report; another excerpt summarizes similarly weak effects and cites no phenylalanine auxotrophy in the PP_3590 PP_1972 double knockout (schmidt2022nitrogenmetabolismin pages 8-10, schmidt2022nitrogenmetabolismin pages 10-12)
Strength/limits of direct evidence for Q88LG1 Evidence for PP_1972/Q88LG1 specifically is mainly genetic/fitness-based in KT2440; direct biochemical characterization in P. putida has more clearly identified other aromatic aminotransferase isozymes (tyrB-2/phhC) than PP_1972 itself, so annotation of Q88LG1 is supported by homology plus mutant evidence rather than purified-enzyme kinetics (szkop2013tyrb2andphhc pages 1-2, schmidt2022nitrogenmetabolismin pages 10-12) No direct purified-enzyme kinetic constants for PP_1972/Q88LG1 were extracted from the cited KT2440 sources; strongest KT2440-specific quantitative data are mutant doubling times and RB-TnSeq fitness values (schmidt2022nitrogenmetabolismin pages 10-12, herrera2010identificationandcharacterization media d9fda959)

Table: This table summarizes the strongest available evidence for functional annotation of Pseudomonas putida KT2440 tyrB (PP_1972; UniProt Q88LG1), including its identity, predicted aminotransferase role, pathway context, redundancy, and the key quantitative phenotypes available from mutant and fitness studies.

9) Source list (with URLs and publication dates where available)

  • Herrera MC, Duque E, Rodríguez‑Herva JJ, Fernández‑Escamilla AM, Ramos JL. Identification and characterization of the PhhR regulon in Pseudomonas putida. Environmental Microbiology (Jun 2010). https://doi.org/10.1111/j.1462-2920.2009.02124.x (herrera2010identificationandcharacterization pages 1-2, herrera2010identificationandcharacterization media d9fda959)
  • Szkop M, Bielawski W. tyrB-2 and phhC genes of Pseudomonas putida encode aromatic amino acid aminotransferase isozymes: evidence at the protein level. Amino Acids (May 2013). https://doi.org/10.1007/s00726-013-1508-y (szkop2013tyrb2andphhc pages 1-2, szkop2013tyrb2andphhc media f8f3824d)
  • Schmidt M et al. Nitrogen metabolism in Pseudomonas putida: functional analysis using random barcode transposon sequencing. Applied and Environmental Microbiology (Apr 2022). https://doi.org/10.1128/aem.02430-21 (schmidt2022nitrogenmetabolismin pages 8-10, schmidt2022nitrogenmetabolismin pages 10-12)
  • Borchert AJ et al. Machine learning analysis of RB‑TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440. mSystems (Mar 2024). https://doi.org/10.1128/msystems.00942-23 (borchert2024machinelearninganalysis pages 7-11)
  • Menke M. Protein Engineering of Amine Transaminases and Methyltransferases using Machine Learning and High-Throughput Screening Tools. (2024; journal not captured in retrieved metadata). (menke2024proteinengineeringof pages 61-71, menke2024proteinengineeringof pages 25-28)
  • Direct PP_1972 enzymology: The retrieved corpus lacks purified-enzyme kinetics (Km/kcat) and direct substrate panels for PP_1972/Q88LG1 specifically; current enzymatic claims for PP_1972 are therefore best stated as inference from gene context + mutant/fitness phenotypes and from biochemical characterization of close isozymes in P. putida. (borchert2024machinelearninganalysis pages 7-11, szkop2013tyrb2andphhc media f8f3824d)
  • Localization: No KT2440 localization evidence for PP_1972 was found here. If localization is critical (e.g., for pathway compartmentalization or cofactor supply), it should be determined experimentally (cell fractionation, fluorescence tagging) or taken from curated UniProt/InterPro annotations.

References

  1. (herrera2010identificationandcharacterization pages 1-2): M. Carmen Herrera, Estrella Duque, José J. Rodríguez‐Herva, Ana M. Fernández‐Escamilla, and Juan L. Ramos. Identification and characterization of the phhr regulon in pseudomonas putida. Environmental microbiology, 12 6:1427-38, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2009.02124.x, doi:10.1111/j.1462-2920.2009.02124.x. This article has 43 citations and is from a domain leading peer-reviewed journal.

  2. (borchert2024machinelearninganalysis pages 7-11): Andrew J. Borchert, Alissa C. Bleem, Hyun Gyu Lim, Kevin Rychel, Keven D. Dooley, Zoe A. Kellermyer, Tracy L. Hodges, Bernhard O. Palsson, and Gregg T. Beckham. Machine learning analysis of rb-tnseq fitness data predicts functional gene modules in pseudomonas putida kt2440. Mar 2024. URL: https://doi.org/10.1128/msystems.00942-23, doi:10.1128/msystems.00942-23. This article has 13 citations and is from a peer-reviewed journal.

  3. (szkop2013tyrb2andphhc pages 1-2): Michał Szkop and Wiesław Bielawski. Tyrb-2 and phhc genes of pseudomonas putida encode aromatic amino acid aminotransferase isozymes: evidence at the protein level. Amino Acids, 45:351-358, May 2013. URL: https://doi.org/10.1007/s00726-013-1508-y, doi:10.1007/s00726-013-1508-y. This article has 5 citations and is from a peer-reviewed journal.

  4. (menke2024proteinengineeringof pages 22-25): M Menke. Protein engineering of amine transaminases and methyltransferases using machine learning and high-throughput screening tools. Unknown journal, 2024.

  5. (menke2024proteinengineeringof pages 25-28): M Menke. Protein engineering of amine transaminases and methyltransferases using machine learning and high-throughput screening tools. Unknown journal, 2024.

  6. (herrera2010identificationandcharacterization pages 4-5): M. Carmen Herrera, Estrella Duque, José J. Rodríguez‐Herva, Ana M. Fernández‐Escamilla, and Juan L. Ramos. Identification and characterization of the phhr regulon in pseudomonas putida. Environmental microbiology, 12 6:1427-38, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2009.02124.x, doi:10.1111/j.1462-2920.2009.02124.x. This article has 43 citations and is from a domain leading peer-reviewed journal.

  7. (szkop2013tyrb2andphhc media f8f3824d): Michał Szkop and Wiesław Bielawski. Tyrb-2 and phhc genes of pseudomonas putida encode aromatic amino acid aminotransferase isozymes: evidence at the protein level. Amino Acids, 45:351-358, May 2013. URL: https://doi.org/10.1007/s00726-013-1508-y, doi:10.1007/s00726-013-1508-y. This article has 5 citations and is from a peer-reviewed journal.

  8. (herrera2010identificationandcharacterization media d9fda959): M. Carmen Herrera, Estrella Duque, José J. Rodríguez‐Herva, Ana M. Fernández‐Escamilla, and Juan L. Ramos. Identification and characterization of the phhr regulon in pseudomonas putida. Environmental microbiology, 12 6:1427-38, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2009.02124.x, doi:10.1111/j.1462-2920.2009.02124.x. This article has 43 citations and is from a domain leading peer-reviewed journal.

  9. (schmidt2022nitrogenmetabolismin pages 8-10): Matthias Schmidt, Allison N. Pearson, Matthew R. Incha, Mitchell G. Thompson, Edward E. K. Baidoo, Ramu Kakumanu, Aindrila Mukhopadhyay, Patrick M. Shih, Adam M. Deutschbauer, Lars M. Blank, and Jay D. Keasling. Nitrogen metabolism in pseudomonas putida: functional analysis using random barcode transposon sequencing. Applied and Environmental Microbiology, Apr 2022. URL: https://doi.org/10.1128/aem.02430-21, doi:10.1128/aem.02430-21. This article has 36 citations and is from a peer-reviewed journal.

  10. (schmidt2022nitrogenmetabolismin pages 10-12): Matthias Schmidt, Allison N. Pearson, Matthew R. Incha, Mitchell G. Thompson, Edward E. K. Baidoo, Ramu Kakumanu, Aindrila Mukhopadhyay, Patrick M. Shih, Adam M. Deutschbauer, Lars M. Blank, and Jay D. Keasling. Nitrogen metabolism in pseudomonas putida: functional analysis using random barcode transposon sequencing. Applied and Environmental Microbiology, Apr 2022. URL: https://doi.org/10.1128/aem.02430-21, doi:10.1128/aem.02430-21. This article has 36 citations and is from a peer-reviewed journal.

  11. (ringel2017theperiplasmictransaminase pages 16-18): Michael T. Ringel, Gerald Dräger, and Thomas Brüser. The periplasmic transaminase ptaa of pseudomonas fluorescens converts the glutamic acid residue at the pyoverdine fluorophore to α-ketoglutaric acid. Journal of Biological Chemistry, 292:18660-18671, Nov 2017. URL: https://doi.org/10.1074/jbc.m117.812545, doi:10.1074/jbc.m117.812545. This article has 17 citations and is from a domain leading peer-reviewed journal.

  12. (menke2024proteinengineeringof pages 61-71): M Menke. Protein engineering of amine transaminases and methyltransferases using machine learning and high-throughput screening tools. Unknown journal, 2024.

  13. (herrera2010identificationandcharacterization pages 9-10): M. Carmen Herrera, Estrella Duque, José J. Rodríguez‐Herva, Ana M. Fernández‐Escamilla, and Juan L. Ramos. Identification and characterization of the phhr regulon in pseudomonas putida. Environmental microbiology, 12 6:1427-38, Jun 2010. URL: https://doi.org/10.1111/j.1462-2920.2009.02124.x, doi:10.1111/j.1462-2920.2009.02124.x. This article has 43 citations and is from a domain leading peer-reviewed journal.

  14. (szkop2013tyrb2andphhc pages 2-4): Michał Szkop and Wiesław Bielawski. Tyrb-2 and phhc genes of pseudomonas putida encode aromatic amino acid aminotransferase isozymes: evidence at the protein level. Amino Acids, 45:351-358, May 2013. URL: https://doi.org/10.1007/s00726-013-1508-y, doi:10.1007/s00726-013-1508-y. This article has 5 citations and is from a peer-reviewed journal.

Artifacts

Citations

  1. herrera2010identificationandcharacterization pages 1-2
  2. herrera2010identificationandcharacterization pages 4-5
  3. borchert2024machinelearninganalysis pages 7-11
  4. ringel2017theperiplasmictransaminase pages 16-18
  5. menke2024proteinengineeringof pages 61-71
  6. menke2024proteinengineeringof pages 22-25
  7. menke2024proteinengineeringof pages 25-28
  8. schmidt2022nitrogenmetabolismin pages 8-10
  9. schmidt2022nitrogenmetabolismin pages 10-12
  10. herrera2010identificationandcharacterization pages 9-10
  11. https://doi.org/10.1128/msystems.00942-23
  12. https://doi.org/10.1111/j.1462-2920.2009.02124.x
  13. https://doi.org/10.1007/s00726-013-1508-y
  14. https://doi.org/10.1128/aem.02430-21
  15. https://doi.org/10.1111/j.1462-2920.2009.02124.x,
  16. https://doi.org/10.1128/msystems.00942-23,
  17. https://doi.org/10.1007/s00726-013-1508-y,
  18. https://doi.org/10.1128/aem.02430-21,
  19. https://doi.org/10.1074/jbc.m117.812545,

📄 View Raw YAML

id: Q88LG1
gene_symbol: tyrB
product_type: PROTEIN
status: DRAFT
taxon:
  id: NCBITaxon:160488
  label: Pseudomonas putida (strain ATCC 47054 / DSM 6125 / CFBP 8728 / NCIMB 11950 / KT2440)
description: >-
  tyrB (PP_1972; also referred to as tyrB-1) is a cytoplasmic, pyridoxal
  5'-phosphate (PLP)-dependent aminotransferase of the class-I (fold-type I)
  aspartate aminotransferase superfamily. It catalyzes reversible transamination
  in which an amino group is transferred from an amino acid donor to a 2-oxoacid
  acceptor (commonly 2-oxoglutarate, yielding L-glutamate). It is annotated as an
  aromatic-amino-acid aminotransferase, interconverting aromatic amino acids
  (L-tyrosine, L-phenylalanine) and their cognate aromatic 2-oxoacids
  (4-hydroxyphenylpyruvate, phenylpyruvate), and contributes to aromatic amino
  acid biosynthesis and catabolism. Like many class-I PLP aminotransferases it is
  a homodimer with active sites formed at the subunit interface. In P. putida
  KT2440 it is one of several aminotransferase isozymes acting on aromatic amino
  acids; genetic studies show that loss of tyrB alone causes only mild
  aromatic-amino-acid utilization phenotypes because of redundancy with paralogous
  aminotransferases (notably PP_3590/AmaC and tyrB-2/phhC), so its physiological
  role overlaps with those enzymes.
existing_annotations:
- term:
    id: GO:0003824
    label: catalytic activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: enables
  review:
    summary: >-
      Generic root-level molecular function term. tyrB is an enzyme, so this is
      not wrong, but it is uninformatively general and is fully subsumed by the
      more specific transaminase/aminotransferase activity terms.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      "catalytic activity" is the MF root and conveys no specific information.
      The more precise terms GO:0008483 (transaminase activity) and GO:0004838
      (L-tyrosine:2-oxoglutarate transaminase activity) capture the actual
      function.
- term:
    id: GO:0004838
    label: L-tyrosine:2-oxoglutarate transaminase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: enables
  review:
    summary: >-
      Specific aromatic (tyrosine) aminotransferase activity assigned by
      TreeGrafter from the PTHR11879:SF37 "aromatic-amino-acid aminotransferase"
      subfamily. This is consistent with the protein family, the COG1448
      assignment, and with biochemical/genetic characterization of P. putida
      aromatic aminotransferases. This is the best representation of the gene's
      core molecular function.
    action: ACCEPT
    reason: >-
      Domain/family evidence (class-I PLP aminotransferase, PANTHER ArAT
      subfamily SF37) plus genetic evidence in KT2440 (tyrB mutants show a
      tyrosine-utilization phenotype; PMID:20050871) support tyrosine
      aminotransferase activity. The enzyme is likely promiscuous across aromatic
      amino acids (also acting on phenylalanine), but this term well captures the
      central characterized activity.
- term:
    id: GO:0005829
    label: cytosol
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: located_in
  review:
    summary: >-
      Cytosolic localization predicted by TreeGrafter. Soluble class-I PLP
      aminotransferases acting in amino acid metabolism are cytoplasmic enzymes;
      the sequence has no signal peptide or transmembrane region. Consistent with
      the expected localization.
    action: ACCEPT
    reason: >-
      Aromatic aminotransferases in this family are soluble cytoplasmic enzymes.
      No experimental KT2440 localization data exist, but the prediction is
      biologically appropriate and there is no evidence for periplasmic/membrane
      localization.
- term:
    id: GO:0006520
    label: amino acid metabolic process
  evidence_type: IEA
  original_reference_id: GO_REF:0000002
  qualifier: involved_in
  review:
    summary: >-
      Broad biological process term. tyrB participates in (aromatic) amino acid
      metabolism, so this is correct but general. A more specific process such as
      aromatic amino acid family metabolism / phenylalanine or tyrosine
      biosynthesis or catabolism would be more informative.
    action: MODIFY
    reason: >-
      The annotation is correct in essence but too high-level. tyrB acts
      specifically on aromatic amino acids (Tyr/Phe), so the more specific
      "aromatic amino acid family metabolic process" better reflects the
      characterized role while remaining defensible from family + genetic
      evidence. (Chorismate metabolic process is not appropriate: tyrB acts
      downstream of chorismate on the aromatic amino acids/2-oxoacids, not on
      chorismate itself.)
    proposed_replacement_terms:
    - id: GO:0009072
      label: aromatic amino acid family metabolic process
- term:
    id: GO:0008483
    label: transaminase activity
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: >-
      General transaminase (aminotransferase) activity. Correct and well
      supported by the class-I PLP-dependent aminotransferase family assignment,
      but less specific than GO:0004838. Useful as a parent term.
    action: KEEP_AS_NON_CORE
    reason: >-
      Accurately describes the enzymatic class but is a parent of the more
      specific aromatic aminotransferase term that represents the core function.
      Retain as supporting/non-core rather than as the primary MF.
- term:
    id: GO:0030170
    label: pyridoxal phosphate binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000120
  qualifier: enables
  review:
    summary: >-
      PLP cofactor binding. tyrB is a PLP-dependent enzyme (UniProt COFACTOR:
      pyridoxal 5'-phosphate; conserved PROSITE PS00105 class-I aminotransferase
      PLP-binding motif). This is a well-supported and informative molecular
      function annotation.
    action: ACCEPT
    reason: >-
      Strong family/motif evidence (IPR004838/IPR004839, PROSITE AA_TRANSFER_
      CLASS_1, UniProt cofactor annotation) for PLP binding, which is essential
      for the transamination mechanism.
- term:
    id: GO:0042802
    label: identical protein binding
  evidence_type: IEA
  original_reference_id: GO_REF:0000118
  qualifier: enables
  review:
    summary: >-
      Self-association annotation reflecting the homodimeric quaternary structure
      typical of class-I aminotransferases (UniProt SUBUNIT: Homodimer). While
      the homodimer assignment is reasonable, "identical protein binding" is an
      uninformative interaction term that does not convey biological function and
      is a frequent TreeGrafter over-propagation.
    action: MARK_AS_OVER_ANNOTATED
    reason: >-
      Homodimerization is a structural property rather than a distinct molecular
      function; the term adds little and is propagated electronically without
      direct evidence for this protein. Per curation guidance, generic
      "protein binding"-type terms are discouraged.
core_functions:
- description: >-
    PLP-dependent aromatic-amino-acid aminotransferase catalyzing reversible
    transamination between aromatic amino acids (L-tyrosine, L-phenylalanine) and
    their 2-oxoacids using 2-oxoglutarate/L-glutamate as the amino acceptor/donor
    pair, functioning in aromatic amino acid biosynthesis and catabolism.
  molecular_function:
    id: GO:0004838
    label: L-tyrosine:2-oxoglutarate transaminase activity
  supported_by:
  - reference_id: PMID:23685963
  - reference_id: PMID:20050871
  directly_involved_in:
  - id: GO:0009072
    label: aromatic amino acid metabolic process
proposed_new_terms: []
suggested_questions:
- question: >-
    What is the in vitro substrate range and kinetic preference of purified
    PP_1972 (Tyr vs Phe vs Trp; 2-oxoglutarate vs pyruvate as acceptor), given
    that direct enzymology exists for the paralogs but not for PP_1972 itself?
- question: >-
    What is the division of labor among the P. putida KT2440 aromatic
    aminotransferase isozymes (PP_1972/tyrB-1, PP_3590/AmaC, tyrB-2/phhC) in
    aromatic amino acid biosynthesis versus catabolism, and under what conditions
    is PP_1972 non-redundant?
suggested_experiments:
- description: >-
    Express and purify recombinant PP_1972 and determine kinetic constants
    (Km/kcat) against a panel of amino donors (Tyr, Phe, Trp, Asp) and 2-oxoacid
    acceptors to define substrate specificity directly.
- description: >-
    Construct single and combinatorial in-frame deletions of PP_1972, PP_3590,
    and tyrB-2/phhC and assay growth on aromatic amino acids as sole nitrogen and
    carbon sources to resolve the redundancy network and any condition-specific,
    non-redundant role of PP_1972.
references:
- id: GO_REF:0000002
  title: Gene Ontology annotation through association of InterPro records with GO terms
  findings: []
- id: GO_REF:0000118
  title: TreeGrafter-generated GO annotations
  findings: []
- id: GO_REF:0000120
  title: Combined Automated Annotation using Multiple IEA Methods
  findings: []
- id: PMID:23685963
  title: "tyrB-2 and phhC genes of Pseudomonas putida encode aromatic amino acid aminotransferase isozymes: evidence at the protein level"
  findings:
  - statement: >-
      P. putida aromatic aminotransferase isozymes preferentially transaminate
      aromatic amino acids and aromatic 2-oxoacids (best substrates L-phenylalanine
      and phenylpyruvate), using PLP cofactor and 2-oxoglutarate as amino acceptor.
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: >-
      PMID verified via PubMed (Szkop & Bielawski, Amino Acids 2013). Establishes
      aromatic aminotransferase activity for P. putida isozymes (paralogs of
      PP_1972), supporting the family-level molecular function assignment.
- id: PMID:20050871
  title: "Identification and characterization of the PhhR regulon in Pseudomonas putida"
  findings:
  - statement: >-
      Genetic study of aromatic amino acid catabolism in P. putida KT2440; tyrB-1
      (PP_1972) and tyrB-2 mutants show altered doubling times on tyrosine and
      phenylalanine as nitrogen sources, implicating tyrB-family aminotransferases
      in aromatic amino acid utilization.
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: >-
      PMID verified via PubMed (Herrera et al., Environ Microbiol 2010). Provides
      KT2440-specific genetic evidence linking PP_1972 to aromatic amino acid
      metabolism.
- id: PMID:38323821
  title: "Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440"
  findings:
  - statement: >-
      Disruption of tyrB (PP_1972) did not inhibit growth on L-phenylalanine or
      L-tyrosine as sole nitrogen sources, whereas disruption of AmaC (PP_3590)
      abolished growth; the authors propose PP_3590 as the dominant L-tyrosine
      aminotransferase, indicating PP_1972 is functionally redundant under those
      conditions.
  reference_review:
    relevance: HIGH
    correctness: VERIFIED
    review_notes: >-
      PMID verified via PubMed (Borchert et al., mSystems 2024). Important for
      interpreting the in-vivo, non-core/redundant role of PP_1972.
- id: PMID:35285712
  title: "Nitrogen metabolism in Pseudomonas putida: functional analysis using random barcode transposon sequencing"
  findings:
  - statement: >-
      RB-TnSeq fitness data show only weak single-gene fitness effects for
      PP_1972 on aromatic nitrogen sources, and a PP_3590/PP_1972 double knockout
      did not cause phenylalanine auxotrophy, consistent with redundancy among
      aromatic aminotransferases.
  reference_review:
    relevance: MEDIUM
    correctness: VERIFIED
    review_notes: >-
      PMID verified via PubMed (Schmidt et al., Appl Environ Microbiol 2022).
      Supports the redundancy interpretation; corroborating rather than primary.