Skip to content

Panther Gene Orthology Reference Ingest Guide

Source Information

InfoRes ID: infores:panther

Description: PANTHER (Protein ANalysis THrough Evolutionary Relationships) is developed and maintained by researchers at the University of Southern California in collaboration with the Gene Ontology Consortium and other partners. The core of PANTHER is a comprehensive, annotated library of gene family phylogenetic trees providing protein family and subfamily classifications, ortholog relationships, Gene Ontology annotations, biological pathways, protein class assignments, and evolutionary histories for genes across numerous species. The goal of Panther is to facilitate high-throughput analysis for comparative genomics, functional classification, evolutionary analysis, gene set enrichment, and interpretation of high-throughput genomic experiments. Knowledge is generated through computational phylogenetic analyses that classify proteins into evolutionary families, inference of orthology and gene function from evolutionary relationships, integration of Gene Ontology annotations and other external biological resources, and expert manual curation of selected protein families, evolutionary events, and biological pathways. All nodes in the tree have persistent identifiers that are maintained between versions of PANTHER, providing a stable substrate for annotations of protein properties like subfamily and function.

Citations: - https://doi.org/10.1002/pro.4218

Data Access Locations: - http://data.pantherdb.org/ftp/

Data Provision Mechanisms: file_download

Data Formats: csv

Data Versioning and Releases: Versioning by number. See https://www.pantherdb.org/data/

Ingest Information

Ingest Categories: primary_knowledge_provider

Utility: Homology relationships and association of GO and related annotation by orthology inference can be made in between human genes and non-human species like mouse, rat and many model species, based on the phenotypic characteristics of genes transitively inferred from experimental observations in the model species which cannot generally be easily or ethically replicated upon human beings.

Scope: Gene to gene genomic orthology relationships and associated annotations

Relevant Files

File Name Location Description
RefGenomeOrthologs.tar.gz http://data.pantherdb.org/ftp/ortholog/current_release/ Gene to Gene Orthology Relationships in reference genomes
PTHR_ http://data.pantherdb.org/ftp/sequence_classifications/current_release/PANTHER_Sequence_Classification_files/ Gene sequence annotation from specific genomes

Included Content

File Name Included Records Fields Used
RefGenomeOrthologs.tar.gz All records, with taxonomic filtering noted below. Gene, Ortholog, Type of ortholog, Panther Ortholog ID
PTHR_ All records for the specified taxon. Gene, Ortholog, Type of ortholog, Panther Ortholog ID

Filtered Content

File Name Filtered Records Rationale
RefGenomeOrthologs.tar.gz All records with Gene and Ortholog pairwise annotated with taxon name as 'HUMAN', 'MOUSE' or 'RAT' specific. Panther contains a huge number of records covering orthologs across 144 diverse species (as of September 2025), but our core interest in Translator focuses on genes in taxa close in evolutionary terms to human, thus having significant genetic, molecular and physiological annotation closer to human biology in character. It is also expected that the genes of these evolutionarily close species also already have a significant assignment of functional roles partially inferred from other model organisms (e.g. developmental gene functions mapped from fruit fly or nematode onto mouse genes, but also experimentally tested in mouse). The character of genetic and physiological systems are much more similar between humans and mice or rats, in particular, studied responses in pharmacology, metabolism, and the immune system.

Future Content Considerations

other: Additional model species may be included in the future.

edge_property_content: Are there any useful edge properties / EPC metrics that could be collected on orthology edges (e.g. data items output from the sequence alignment / phylogenetic analysis) that could be used to score these edges? Notes: this would require meta-analysis of the Panther HMM classification files (at https://data.pantherdb.org/ftp/hmm_classifications/current_release) and/or sequence classification files (at https://data.pantherdb.org/ftp/sequence_classifications/current_release).

node_property_content: Family node properties we might capture: HMM length, member count, others? Note: HMM data annotation would require meta-analysis of the Panther HMM library data at https://data.pantherdb.org/ftp/panther_library/current_release/

Target Information

Edge Types

Subject Categories Predicate Object Categories Knowledge Level Agent Type UI Explanation
biolink:Gene biolink:Gene knowledge_assertion manual_validation_of_automated_agent Panther gene orthology. To determine when two genes are orthologous, Panther uses advanced automated algorithms that rely only on sequence similarity, evolutionary tree topology, and evolutionary event labeling. Curators will manually review the data underlying orthology calls to improve accuracy.
biolink:Gene biolink:GeneFamily knowledge_assertion automated_agent Gene membership in Panther orthology family. Panther defines protein families using a statistical Hidden Markov Model (HMM) that represents its sequence signature. Individual protein sequences are compared to and scored against these signatures, with those passing a membership threshold being assigned to a family.

Node Types

Node Category Source Identifier Types Additional Notes
biolink:Gene HGNC, MGI, RGD, ENSEMBL
biolink:GeneFamily PANTHER.FAMILY

Future Modeling Considerations

other: The Monarch Initiative ingest of Panther data (https://github.com/monarch-initiative/pantherdb-orthologs-ingest) sometimes attempts to map eccentric identifiers into the NCBI gene identifier space. The initial iteration of the Panther ingest in Translator does not attempt to do this at this time, but rather simply uses the given id.

Provenance Information

Contributors: - Richard Bruskiewich - data modelling, domain expertise, code author - Kevin Schaper - Phase 2 legacy code expert - Evan Morris - Phase 2 legacy code expert - Chunlei Wu - Phase 2 legacy code expert - Matt Brush - data modelling

Artifacts: - Ingest Survey (https://docs.google.com/spreadsheets/d/1YlpI5bjGNGR5JC9VWxZJ7dd87hS_b4BMZv5geYe2NCk/)