Human Phenotype Ontology Annotations Reference Ingest Guide¶
Source Information¶
InfoRes ID: infores:hpo-annotations
Description: The Human Phenotype Ontology (HPO) annotation dataset is produced and maintained by the Human Phenotype Ontology Consortium, led by investigators at the Charité – Universitätsmedizin Berlin and collaborating institutions. Its purpose is to provide standardized annotations linking human diseases and genes to phenotypic abnormalities using Human Phenotype Ontology terms, thereby supporting rare disease diagnosis, genomic variant interpretation, disease gene discovery, and computational phenotype analysis. The dataset contains disease–phenotype annotations, gene–phenotype associations, phenotype frequencies, onset information, modifiers, evidence metadata, and supporting references. Knowledge is generated through expert manual curation of the biomedical literature, clinical case reports, and authoritative disease resources, together with integration of annotations from external databases such as rare disease and clinical genetics resources. Additional computational pipelines harmonize annotations to ontology terms and propagate annotations through the ontology hierarchy while preserving provenance and evidence for each association. Here we create Biolink associations between genes, phenotypic features and diseases, together with their evidence, and age of onset and frequency (if known). Disease annotations here are also cross-referenced to the MONarch Disease Ontology (MONDO). There are three HPOA ingests: 'disease-to-phenotype' (nodes including disease modes of inheritance when given, plus edges), 'gene-to-phenotype' and 'gene-to-disease'.
Citations: - https://doi.org/10.1093/nar/gkaa1043
Data Access Locations: - https://hpo.jax.org/data/annotations
Data Provision Mechanisms: file_download, api_endpoint
Data Formats: tsv, other
Data Versioning and Releases: GitHub managed releases at https://github.com/obophenotype/human-phenotype-ontology/releases HPO is released +/- 5 times each year. Versioning is based on the month and year of the release.
Ingest Information¶
Ingest Categories: primary_knowledge_provider
Utility: The Human Phenotype Ontology was initially developed at the Charité University Hospital in Berlin, Germany by Peter N Robinson and Sebastian Köhler with support from Sandra Doelken and Sebastian Bauer. Very soon after the initial publication of the HPO in 2008, a collaboration was initiated with researchers from the Lawrence Berkeley National Laboratory (Chris Mungall, Suzi Lewis, and Nicole Washington), the University of Cambridge (Michael Ashburner Paul Schofield, George Gkoutos), and the University of Oregon (Melissa Haendel, Monte Westerfield). These groups have developed a scheme for Linking human diseases to animal models using ontology-based phenotype annotation. In parallel, the DECIPHER project at the Wellcome Trust Sanger Institute led by Helen Firth and Matt Hurles, conducted clinical workshops to extend the HPO and have provided an enormous amount of support and input in subsequent years. Knowledge asserted by HPO phenotype/disease/gene annotation aligns well with research questions for target users of the Biomedical Data Translator.
Scope: Covers curated Disease, Phenotype and Genes relationships annotated with Human Phenotype Ontology terms.
Relevant Files¶
| File Name | Location | Description |
|---|---|---|
| phenotype.hpoa | https://hpo.jax.org/data/annotations | disease to HPO phenotype annotations, including inheritance information |
| genes_to_disease.txt | https://hpo.jax.org/data/annotations | gene to HPO disease annotations |
| genes_to_phenotype.txt | https://hpo.jax.org/data/annotations | gene to HPO phenotype annotations |
Included Content¶
| File Name | Included Records | Fields Used |
|---|---|---|
| phenotype.hpoa | Disease to Phenotype relationships (i.e., rows with 'aspect' == 'P') | database_id, qualifier, hpo_id, reference, evidence, onset, frequency, sex, aspect |
| phenotype.hpoa | Disease "Mode of Inheritance" relationships (i.e., rows with 'aspect' == 'I') represented as node properties rather than edges | database_id, qualifier, hpo_id, reference, evidence, onset, frequency, sex, aspect |
| genes_to_disease.txt | Mendelian Gene to Disease relationships (i.e., rows with 'association_type' == 'MENDELIAN') | ncbi_gene_id, gene_symbol, association_type, disease_id, source |
| genes_to_disease.txt | Polygenic Gene to Disease relationships (i.e., rows with 'association_type' == 'POLYGENIC') | ncbi_gene_id, gene_symbol, association_type, disease_id, source |
| genes_to_phenotype.txt | Records where we determine that the reported G-P association was inferred over a G-D associated type with the value "MENDELIAN" | ncbi_gene_id, gene_symbol, hpo_id, hpo_name, frequency, disease_id |
Filtered Content¶
| File Name | Filtered Records | Rationale |
|---|---|---|
| phenotype.hpoa | Negated edges from phenotype.hpoa (i.e. edges with qualifier field value 'NOT'). | Negated edge associations are not included as a matter of general policy within Translator at present. |
| genes_to_disease.txt | Records where we determine that the reported G-P association was inferred over a G-D associated type with the value "UNKNOWN" | Records of this type are not included in the ingest because they are Gene-Disease relationships for which the supporting evidence is missing, weak or inconclusive. |
| genes_to_phenotype.txt | Records where we determine that the reported G-P association was inferred over a G-D associated type with the value "POLYGENIC" or "UNKNOWN" | HPO will infer a Gene-Phenotype association G1-P1 in cases where G1 causes, contributes_to, or is associated with D1, and D1 is associated with a Phenotype P1. This logic holds for Mendelian disease where a single gene is causal and thus responsible for all associated phenotypes. It does not necessarily hold for Polygenic or Unknown diseases where the gene may be one of many contributing factors, and thus does not necessarily contribute to or have an association with each phenotype of the disease. |
Future Content Considerations¶
edge_content: Consider bringing back G-P associations based on inferences over Unknown Diseases if we establish a confidence annotation paradigm that lets us indicate these inferences to be weaker than those inferred over Mendelian or Polygenic diseases where the Gene is individually causal for, or contributing to, the disease. - Relevant files: genes_to_disease.txt
edge_content: Consider bringing back G-P associations based on inferences over Polygenic or Unknown Diseases if we establish a confidence annotation paradigm that lets us indicate these inferences to be weaker than those inferred over Mendelian diseases where the Gene is individually causal for the disease and all of its phenotypes. - Relevant files: genes_to_phenotype.txt
Target Information¶
Edge Types¶
| Subject Categories | Predicate | Object Categories | Knowledge Level | Agent Type | UI Explanation |
|---|---|---|---|---|---|
| biolink:Disease | biolink:PhenotypicFeature | knowledge_assertion | text_mining_agent, manual_agent, manual_validation_of_automated_agent | Source HPOA data provide Disease-Phenotype associations that are produced by HPO curators through manual review of clinical data and published evidence. The HPOA record used to create this Translator edge reports that a Phenotype is observed to manifest in a particular Disease, and may include information about the frequency and/or biological context (e.g. sex, onset, frequency) of this manifestation. This relationship is represented using the Biolink 'has phenotype' predicate - with optional qualifiers describing frequency/context information where provided. | |
| biolink:Gene | biolink:Disease | knowledge_assertion | manual_agent | Source HPOA data provide Gene-Disease associations that are manually curated from sources like Orphanet and MIM2Gene and DECIPHER by HPO. The record used to create this Translator edge reports that a Gene is associated with a Disease when there is evidence that variants in the gene may cause or contribute to its manifestation. The Biolink predicate used to represent this relationship depends on the type of Disease. The qualified predicate is selected on the basis of whether the 'association_type' reported by HPOA is 'MENDELIAN' or 'POLYGENIC'. For Mendelian disorders (where the "association_type" field = 'MENDELIAN') direct causation is assumed hence the qualified_predicate is 'causes'). For polygenic disorders, the qualified predicate is rather set to 'contributes_to' since the subject gene of the relationship may only be one of a set of genes contributing to the phenotype. Qualifier is used to indicate that it is a 'genetic_variant_form' of the gene that participates in these relationships, not the Gene in general. | |
| biolink:Gene | biolink:PhenotypicFeature | logical_entailment | text_mining_agent, manual_agent, manual_validation_of_automated_agent | A left outer join of HPOA phenotype, gene-to-disease and gene-to-phenotype data provides for the assignment of Gene-Phenotype associations that are manifested in a Disease associated with a particular Gene. | |
| The record used to create this edge specifically reports a Gene-Phenotype association where the Gene has variants causal for a Mendelian disorder in which the Phenotype is manifested - from which it is logically entailed that variants in the Gene are also causal for the Phenotype (as such disorders have a single causal gene). | |||||
| Given the uncertainties of directly mapping genes onto phenotypic features for either polygenic or diseases of unknown genetic etiology, this edge type is only generated for Mendelian disorders. The Mendelian causality of the relationship is therefore represented using the Biolink 'causes' predicate. | |||||
| Use of the qualifier 'genetic_variant_form' indicates that it is a variant of the gene that is causing the Phenotype (not necessarily other variants, including wild type variants of the Gene). | |||||
| Agent type of the relationship is inherited from the HPOA phenotype.hpoa file. |
Node Types¶
| Node Category | Source Identifier Types | Additional Notes |
|---|---|---|
| biolink:Disease | OMIM, ORPHANET, DECIPHER | |
| biolink:PhenotypicFeature | HP | |
| biolink:Gene | NCBIGene |
Future Modeling Considerations¶
spoq_pattern: Consider alternate patterns for representing G-causes-D and G-contributes_to-D associations where we place more semantics into predicates, per https://github.com/NCATSTranslator/Data-Ingest-Coordination-Working-Group/issues/22 Should we consider creating support paths in our data/graphs, for the G-D-P hops over which HPO infers G-P associations? (e.g. GENE1 -causes-> DISEASE1 -has_phenotype-> PHENO1 ----> GENE1 -causes-> PHENO1)
spoq_pattern: Negated edges from phenotype.hpoa (i.e. edges with qualifier field value 'NOT') are filtered out in the current ingest, but could be easily re-introduced at a later date, should they be of interest.
Additional Notes: ['The HPOA ingest of the Disease.node_properties.inheritance value is currently thought to be mildly unreliable due to the way node property merging is (not) currently implemented in the ingest. This is a known issue (Translator Ingest PR https://github.com/NCATSTranslator/translator-ingests/issues/259) which will be addressed in a future release of the pipeline and/or ingest.\n']
Provenance Information¶
Contributors: - Richard Bruskiewich - data modeling, domain expertise, code author - Kevin Schaper - code author - Sierra Moxon - data modeling, domain expertise, code support - Matthew Brush - data modeling, domain expertise
Artifacts: - Ingest Survey (https://docs.google.com/spreadsheets/d/1iLpJ3vULH058rBxlGv_soUyZw107OORZ-cUOgo3Lz3M/)