Skip to content

Human Phenotype Ontology Annotations Reference Ingest Guide

Source Information

InfoRes ID: infores:hpo-annotations

Description: The Human Phenotype Ontology (HPO) annotation dataset is produced and maintained by the Human Phenotype Ontology Consortium, led by investigators at the Charité – Universitätsmedizin Berlin and collaborating institutions. Its purpose is to provide standardized annotations linking human diseases and genes to phenotypic abnormalities using Human Phenotype Ontology terms, thereby supporting rare disease diagnosis, genomic variant interpretation, disease gene discovery, and computational phenotype analysis. The dataset contains disease–phenotype annotations, gene–phenotype associations, phenotype frequencies, onset information, modifiers, evidence metadata, and supporting references. Knowledge is generated through expert manual curation of the biomedical literature, clinical case reports, and authoritative disease resources, together with integration of annotations from external databases such as rare disease and clinical genetics resources. Additional computational pipelines harmonize annotations to ontology terms and propagate annotations through the ontology hierarchy while preserving provenance and evidence for each association. Here we create Biolink associations between genes, phenotypic features and diseases, together with their evidence, and age of onset and frequency (if known). Disease annotations here are also cross-referenced to the MONarch Disease Ontology (MONDO). There are three HPOA ingests: 'disease-to-phenotype' (nodes including disease modes of inheritance when given, plus edges), 'gene-to-phenotype' and 'gene-to-disease'.

Citations: - https://doi.org/10.1093/nar/gkaa1043

Data Access Locations: - https://hpo.jax.org/data/annotations

Data Provision Mechanisms: file_download, api_endpoint

Data Formats: tsv, other

Data Versioning and Releases: GitHub managed releases at https://github.com/obophenotype/human-phenotype-ontology/releases HPO is released +/- 5 times each year. Versioning is based on the month and year of the release.

Ingest Information

Ingest Categories: primary_knowledge_provider

Utility: The Human Phenotype Ontology was initially developed at the Charité University Hospital in Berlin, Germany by Peter N Robinson and Sebastian Köhler with support from Sandra Doelken and Sebastian Bauer. Very soon after the initial publication of the HPO in 2008, a collaboration was initiated with researchers from the Lawrence Berkeley National Laboratory (Chris Mungall, Suzi Lewis, and Nicole Washington), the University of Cambridge (Michael Ashburner Paul Schofield, George Gkoutos), and the University of Oregon (Melissa Haendel, Monte Westerfield). These groups have developed a scheme for Linking human diseases to animal models using ontology-based phenotype annotation. In parallel, the DECIPHER project at the Wellcome Trust Sanger Institute led by Helen Firth and Matt Hurles, conducted clinical workshops to extend the HPO and have provided an enormous amount of support and input in subsequent years. Knowledge asserted by HPO phenotype/disease/gene annotation aligns well with research questions for target users of the Biomedical Data Translator.

Scope: Covers curated Disease, Phenotype and Genes relationships annotated with Human Phenotype Ontology terms.

Relevant Files

File Name Location Description
phenotype.hpoa https://hpo.jax.org/data/annotations disease to HPO phenotype annotations, including inheritance information
genes_to_disease.txt https://hpo.jax.org/data/annotations gene to HPO disease annotations
genes_to_phenotype.txt https://hpo.jax.org/data/annotations gene to HPO phenotype annotations

Included Content

File Name Included Records Fields Used
phenotype.hpoa Disease to Phenotype relationships (i.e., rows with 'aspect' == 'P') database_id, qualifier, hpo_id, reference, evidence, onset, frequency, sex, aspect
phenotype.hpoa Disease "Mode of Inheritance" relationships (i.e., rows with 'aspect' == 'I') represented as node properties rather than edges database_id, qualifier, hpo_id, reference, evidence, onset, frequency, sex, aspect
genes_to_disease.txt Mendelian Gene to Disease relationships (i.e., rows with 'association_type' == 'MENDELIAN') ncbi_gene_id, gene_symbol, association_type, disease_id, source
genes_to_disease.txt Polygenic Gene to Disease relationships (i.e., rows with 'association_type' == 'POLYGENIC') ncbi_gene_id, gene_symbol, association_type, disease_id, source
genes_to_phenotype.txt Records where we determine that the reported G-P association was inferred over a G-D associated type with the value "MENDELIAN" ncbi_gene_id, gene_symbol, hpo_id, hpo_name, frequency, disease_id

Filtered Content

File Name Filtered Records Rationale
phenotype.hpoa Negated edges from phenotype.hpoa (i.e. edges with qualifier field value 'NOT'). Negated edge associations are not included as a matter of general policy within Translator at present.
genes_to_disease.txt Records where we determine that the reported G-P association was inferred over a G-D associated type with the value "UNKNOWN" Records of this type are not included in the ingest because they are Gene-Disease relationships for which the supporting evidence is missing, weak or inconclusive.
genes_to_phenotype.txt Records where we determine that the reported G-P association was inferred over a G-D associated type with the value "POLYGENIC" or "UNKNOWN" HPO will infer a Gene-Phenotype association G1-P1 in cases where G1 causes, contributes_to, or is associated with D1, and D1 is associated with a Phenotype P1. This logic holds for Mendelian disease where a single gene is causal and thus responsible for all associated phenotypes. It does not necessarily hold for Polygenic or Unknown diseases where the gene may be one of many contributing factors, and thus does not necessarily contribute to or have an association with each phenotype of the disease.

Future Content Considerations

edge_content: Consider bringing back G-P associations based on inferences over Unknown Diseases if we establish a confidence annotation paradigm that lets us indicate these inferences to be weaker than those inferred over Mendelian or Polygenic diseases where the Gene is individually causal for, or contributing to, the disease. - Relevant files: genes_to_disease.txt

edge_content: Consider bringing back G-P associations based on inferences over Polygenic or Unknown Diseases if we establish a confidence annotation paradigm that lets us indicate these inferences to be weaker than those inferred over Mendelian diseases where the Gene is individually causal for the disease and all of its phenotypes. - Relevant files: genes_to_phenotype.txt

Target Information

Edge Types

Subject Categories Predicate Object Categories Knowledge Level Agent Type UI Explanation
biolink:Disease biolink:PhenotypicFeature knowledge_assertion text_mining_agent, manual_agent, manual_validation_of_automated_agent Source HPOA data provide Disease-Phenotype associations that are produced by HPO curators through manual review of clinical data and published evidence. The HPOA record used to create this Translator edge reports that a Phenotype is observed to manifest in a particular Disease, and may include information about the frequency and/or biological context (e.g. sex, onset, frequency) of this manifestation. This relationship is represented using the Biolink 'has phenotype' predicate - with optional qualifiers describing frequency/context information where provided.
biolink:Gene biolink:Disease knowledge_assertion manual_agent Source HPOA data provide Gene-Disease associations that are manually curated from sources like Orphanet and MIM2Gene and DECIPHER by HPO. The record used to create this Translator edge reports that a Gene is associated with a Disease when there is evidence that variants in the gene may cause or contribute to its manifestation. The Biolink predicate used to represent this relationship depends on the type of Disease. The qualified predicate is selected on the basis of whether the 'association_type' reported by HPOA is 'MENDELIAN' or 'POLYGENIC'. For Mendelian disorders (where the "association_type" field = 'MENDELIAN') direct causation is assumed hence the qualified_predicate is 'causes'). For polygenic disorders, the qualified predicate is rather set to 'contributes_to' since the subject gene of the relationship may only be one of a set of genes contributing to the phenotype. Qualifier is used to indicate that it is a 'genetic_variant_form' of the gene that participates in these relationships, not the Gene in general.
biolink:Gene biolink:PhenotypicFeature logical_entailment text_mining_agent, manual_agent, manual_validation_of_automated_agent A left outer join of HPOA phenotype, gene-to-disease and gene-to-phenotype data provides for the assignment of Gene-Phenotype associations that are manifested in a Disease associated with a particular Gene.
The record used to create this edge specifically reports a Gene-Phenotype association where the Gene has variants causal for a Mendelian disorder in which the Phenotype is manifested - from which it is logically entailed that variants in the Gene are also causal for the Phenotype (as such disorders have a single causal gene).
Given the uncertainties of directly mapping genes onto phenotypic features for either polygenic or diseases of unknown genetic etiology, this edge type is only generated for Mendelian disorders. The Mendelian causality of the relationship is therefore represented using the Biolink 'causes' predicate.
Use of the qualifier 'genetic_variant_form' indicates that it is a variant of the gene that is causing the Phenotype (not necessarily other variants, including wild type variants of the Gene).
Agent type of the relationship is inherited from the HPOA phenotype.hpoa file.

Node Types

Node Category Source Identifier Types Additional Notes
biolink:Disease OMIM, ORPHANET, DECIPHER
biolink:PhenotypicFeature HP
biolink:Gene NCBIGene

Future Modeling Considerations

spoq_pattern: Consider alternate patterns for representing G-causes-D and G-contributes_to-D associations where we place more semantics into predicates, per https://github.com/NCATSTranslator/Data-Ingest-Coordination-Working-Group/issues/22 Should we consider creating support paths in our data/graphs, for the G-D-P hops over which HPO infers G-P associations? (e.g. GENE1 -causes-> DISEASE1 -has_phenotype-> PHENO1 ----> GENE1 -causes-> PHENO1)

spoq_pattern: Negated edges from phenotype.hpoa (i.e. edges with qualifier field value 'NOT') are filtered out in the current ingest, but could be easily re-introduced at a later date, should they be of interest.

Additional Notes: ['The HPOA ingest of the Disease.node_properties.inheritance value is currently thought to be mildly unreliable due to the way node property merging is (not) currently implemented in the ingest. This is a known issue (Translator Ingest PR https://github.com/NCATSTranslator/translator-ingests/issues/259) which will be addressed in a future release of the pipeline and/or ingest.\n']

Provenance Information

Contributors: - Richard Bruskiewich - data modeling, domain expertise, code author - Kevin Schaper - code author - Sierra Moxon - data modeling, domain expertise, code support - Matthew Brush - data modeling, domain expertise

Artifacts: - Ingest Survey (https://docs.google.com/spreadsheets/d/1iLpJ3vULH058rBxlGv_soUyZw107OORZ-cUOgo3Lz3M/)