ChEMBL Reference Ingest Guide¶
Source Information¶
InfoRes ID: infores:chembl
Description: ChEMBL is a large-scale, manually curated database of bioactive molecules with drug-like properties, maintained by the European Bioinformatics Institute (EMBL-EBI). It brings together chemical, bioactivity, and genomic data to aid the translation of genomic information into effective new drugs. ChEMBL captures curated drug mechanisms of action, quantitative bioactivity measurements from assays reported in the literature, drug metabolism relationships, and the protein target composition of complexes and families. Molecules are identified by ChEMBL compound IDs and targets by ChEMBL target IDs (mapped to UniProtKB where applicable).
Citations: - Barbara Zdrazil, Eloy Felix, Fiona Hunter, Emma J Manners, James Blackshaw, Sybilla Corbett, Marleen de Veij, Harris Ioannidis, David Mendez Lopez, Juan F Mosquera, Maria Paula Magarinos, Nicolas Bosc, Ricardo Arcila, Tevik Kizilören, Anna Gaulton, A PatrĂcia Bento, Melissa. F Adasme, Pater Monecke, Gregory A Landrum, Andrew R Leach. Nucleic Acids Res. 2023: gkad1004. doi: 10.1093/nar/gkad1004 - Davies M, Nowotka M, Papadatos G, Dedman N, Gaulton A, Atkinson F, Bellis L, Overington JP. Nucleic Acids Res. 2015; 43(W1):W612-20, doi: 10.1093/nar/gkv352
Data Access Locations: - https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/latest/
Data Provision Mechanisms: file_download, database_dump
Data Formats: mysql, postgresql, sqlite
Data Versioning and Releases: ChEMBL is released roughly twice per year (approximately semiannual). Releases are versioned with an incrementing integer (e.g. ChEMBL 36). This ingest targets ChEMBL 36. Release notes and the latest distribution are available at https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/latest/
Ingest Information¶
Ingest Categories: primary_knowledge_provider
Utility: ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
Scope: Drugs and probes mechanisms, metabolism, bioactivity, gene targets
Relevant Files¶
| File Name | Location | Description |
|---|---|---|
| chembl_36_sqlite.tar.gz | https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/latest/ | ChEMBL SQLite database |
Included Content¶
| File Name | Included Records | Fields Used |
|---|---|---|
| chembl_36_sqlite.tar.gz | Drug mechanisms from the following tables: drug_mechanism, molecule_dictionary, target_dictionary, binding_sites, compound_records, docs, source, target_components, component_sequences, variant_sequences. | drug_mechanism: mec_id, molregno, mechanism_of_action, action_type, direct_interaction, mechanism_comment, selectivity_comment, binding_site_comment, variant_id; molecule_dictionary: molregno, chembl_id; target_dictionary: tid, chembl_id (AS target_chembl_id), pref_name (AS target_name), target_type, organism (AS target_organism); binding_sites: site_id, site_name; compound_records: record_id, doc_id, src_id; docs: doc_id, chembl_id (AS document_chembl_id); source: src_id, src_description (AS source_description); target_components: tid, component_id; component_sequences: component_id, component_type, accession, description, tax_id, organism; variant_sequences: variant_id, mutation, accession (AS mutation_accession) |
| chembl_36_sqlite.tar.gz | Gene targets - same as drug mechanisms with targets mapped to genes | drug_mechanism: mec_id, molregno, mechanism_of_action, action_type, direct_interaction, mechanism_comment, selectivity_comment, binding_site_comment, variant_id; molecule_dictionary: molregno, chembl_id; target_dictionary: tid, chembl_id (AS target_chembl_id), pref_name (AS target_name), target_type, organism (AS target_organism); binding_sites: site_id, site_name; compound_records: record_id, doc_id, src_id; docs: doc_id, chembl_id (AS document_chembl_id); source: src_id, src_description (AS source_description); target_components: tid, component_id; component_sequences: component_id, component_type, accession, description, tax_id, organism; variant_sequences: variant_id, mutation, accession (AS mutation_accession) |
| chembl_36_sqlite.tar.gz | Metabolism information from the following tables: metabolism, molecule_dictionary, compound_records, compound_structures, target_dictionary, metabolism_refs | metabolism: drug_record_id, substrate_record_id, metabolite_record_id, met_id, enzyme_name, met_conversion, met_comment, organism, tax_id, enzyme_tid; molecule_dictionary: molregno, chembl_id, pref_name; compound_records: molregno, record_id, compound_name; target_dictionary: tid, target_type, chembl_id; compound_structures: molregno, standard_inchi, standard_inchi_key, canonical_smiles; metabolism_refs: met_id, ref_type, ref_id, ref_url |
| chembl_36_sqlite.tar.gz | Bioactivity and assay information from the following tables: activities, molecule_dictionary, assays, bioassay_ontology, target_dictionary, target_components, component_sequences, cell_dictionary, assay_type, tissue_dictionary, docs, source, ligand_eff, relationship_type, confidence_score_lookup, confidence_score_lookup | activities: activity_id, standard_type, standard_relation, standard_value, standard_units, pchembl_value, activity_comment, data_validity_comment, standard_text_value, standard_upper_value, uo_units, potential_duplicate, action_type, src_id, doc_id; molecule_dictionary: molregno, chembl_id; assays: assay_id, chembl_id, description, assay_organism, assay_cell_type, assay_subcellular_fraction, bao_format, assay_category, assay_tax_id, assay_tissue, relationship_type, confidence_score, curated_by, src_id, assay_type, cell_id, tissue_id, tid; bioassay_ontology: bao_id, label; target_dictionary: tid, chembl_id, pref_name, organism, target_type; target_components: tid, component_id; component_sequences: component_id, component_type, accession; cell_dictionary: cell_id, chembl_id, cell_name, cell_description, cell_source_tissue, cell_source_organism, cell_source_tax_id, clo_id, efo_id, cellosaurus_id, cl_lincs_id, cell_ontology_id; assay_type: assay_type, assay_desc; tissue_dictionary: tissue_id, chembl_id, pref_name; docs: doc_id, chembl_id, journal, title, year, authors, pubmed_id, doi; source: src_id, src_description; ligand_eff: activity_id, bei, le, lle, sei; relationship_type: relationship_type, relationship_desc; confidence_score_lookup: confidence_score, description, target_mapping; curation_lookup: curated_by, description |
Filtered Content¶
| File Name | Filtered Records | Rationale |
|---|---|---|
| chembl_36_sqlite.tar.gz | Gene targets - removed targets that cannot be mapped to genes (like cell-lines) | targets that cannot be mapped to genes |
| chembl_36_sqlite.tar.gz | Bioactivity records that are not flagged as 'Active' and lack an action_type, and any activity record whose (molecule, target, action) already corresponds to a curated drug_mechanism record (mec_id is not null) or has already been emitted, are excluded. | Only bioactivities with a clear activity outcome are retained; records already represented by a curated drug mechanism edge are dropped to avoid duplicate edges, and repeated (chemical, target, action) combinations are deduplicated. |
Future Content Considerations¶
edge_property_content: ChEMBL drug mechanism and bioactivity records carry descriptive detail - binding site name/comment, mechanism-of-action description and comment, selectivity comment, mutation and mutation accession, and assay description - that is currently held as TODOs in the transform code and not yet emitted as edge properties, pending suitable Biolink Model slots. - Relevant files: chembl_36_sqlite.tar.gz
edge_property_content: Quantitative bioactivity values from the activities table (standard_type, standard_value, standard_units, pchembl_value) are used to select and characterize edges but are not yet represented as edge properties. Consider capturing them to support potency-aware queries. - Relevant files: chembl_36_sqlite.tar.gz
node_property_content: Molecule routes of delivery (derived from the oral/parenteral/topical flags) and additional molecule_dictionary properties (e.g. max_phase, therapeutic_flag, first_approval, withdrawn_flag) are computed or available but not yet emitted as node properties, pending suitable Biolink slots. - Relevant files: chembl_36_sqlite.tar.gz
edge_content: An enzyme/context qualifier linking metabolism edges to the mediating enzyme is computed in code but not yet attached, pending Biolink support for the relevant statement-level qualifier. - Relevant files: chembl_36_sqlite.tar.gz
Target Information¶
Edge Types¶
| Subject Categories | Predicate | Object Categories | Knowledge Level | Agent Type | UI Explanation |
|---|---|---|---|---|---|
| biolink:ChemicalEntity | biolink:Protein, biolink:ProteinFamily, biolink:MacromolecularComplex, biolink:NucleicAcidEntity, biolink:MolecularEntity, biolink:MolecularActivity, biolink:SmallMolecule, biolink:CellularComponent, biolink:CellularOrganism, biolink:CellLine | knowledge_assertion | manual_agent, automated_agent | This edge asserts that a chemical affects a gene or protein target, based on ChEMBL's manually curated drug mechanisms of action and, in some cases, curated bioactivity measurements. The specific ChEMBL 'mechanism of action' (e.g. 'agonist', 'inhibitor', 'stabiliser') determines the qualified predicate and the aspect, direction, and causal-mechanism qualifiers applied to the edge (e.g. an inhibitor causes decreased activity via inhibition). Edges from curated drug mechanisms are manually curated by ChEMBL; edges derived from autocurated bioactivity assays are marked as automated. | |
| biolink:Protein, biolink:MacromolecularComplex, biolink:NucleicAcidEntity, biolink:MolecularEntity, biolink:SmallMolecule | biolink:ChemicalEntity | knowledge_assertion | manual_agent, automated_agent | This edge asserts that a gene or protein (for example an enzyme) acts on and chemically transforms a chemical, based on ChEMBL's curated drug mechanisms and bioactivities. It is created for ChEMBL mechanisms where the protein acts on the chemical (e.g. 'hydrolytic enzyme', 'oxidative enzyme', 'proteolytic enzyme', 'reducing agent', 'methylating agent'). The gene/protein is the subject and the chemical the object; object aspect and direction qualifiers describe what happens to the chemical (e.g. increased hydrolysis or oxidation). | |
| biolink:ChemicalEntity | biolink:Protein, biolink:ProteinFamily, biolink:MacromolecularComplex, biolink:NucleicAcidEntity, biolink:MolecularEntity, biolink:MolecularActivity, biolink:SmallMolecule, biolink:CellularComponent, biolink:CellularOrganism, biolink:CellLine | knowledge_assertion | manual_agent, automated_agent | This edge asserts a direct physical interaction between a chemical and a gene/protein target, based on ChEMBL's curated drug mechanisms and bioactivities. It is created for ChEMBL mechanisms that imply a physical interaction without a specified functional effect (e.g. 'binding agent', 'chelating agent', 'cross-linking agent', 'sequestering agent'), using the 'directly physically interacts with' predicate. ChEMBL mechanisms classified as 'OTHER', where no specific interaction type is known, produce the more generic 'interacts_with' edge. | |
| biolink:Protein, biolink:ProteinFamily | biolink:ChemicalEntity | knowledge_assertion | manual_agent, automated_agent | This edge asserts that a protein or protein family acts on a chemical as its substrate, based on ChEMBL drug mechanisms and bioactivities where the ChEMBL mechanism is 'substrate'. The protein/protein family is the subject and the chemical substrate is the object. Edges from curated drug mechanisms are manual; edges derived from autocurated bioactivity assays are marked as automated. | |
| biolink:MacromolecularComplex, biolink:ProteinFamily, biolink:MolecularActivity, biolink:Protein | biolink:Protein | knowledge_assertion | manual_agent | This edge describes the protein composition of a ChEMBL target that is a complex, protein family, or selectivity/interaction group, asserting that the target 'has part' a particular constituent protein. It is derived from ChEMBL's curated target-to-component mappings (target_components and component_sequences), where each protein component of a complex or family target is linked to that target. | |
| biolink:ChemicalEntity | biolink:ChemicalEntity | knowledge_assertion | manual_agent | This edge asserts that a chemical (a drug or substrate) has a particular metabolite, based on ChEMBL's manually curated drug metabolism records. It links a substrate chemical to its metabolite; where the parent drug differs from the metabolized substrate, an additional edge from the parent drug to the metabolite is also created. A species context qualifier records the organism in which the metabolic conversion was observed. |
Node Types¶
| Node Category | Source Identifier Types | Additional Notes |
|---|---|---|
| biolink:CellLine | ChEMBL | |
| biolink:CellularComponent | ChEMBL | |
| biolink:CellularOrganism | ChEMBL | |
| biolink:ChemicalEntity | ChEMBL | |
| biolink:MacromolecularComplex | ChEMBL | |
| biolink:MolecularActivity | ChEMBL | |
| biolink:MolecularEntity | ChEMBL | |
| biolink:NucleicAcidEntity | ChEMBL | |
| biolink:Protein | ChEMBL, UniProtKB | |
| biolink:ProteinFamily | ChEMBL | |
| biolink:SmallMolecule | ChEMBL |
Future Modeling Considerations¶
edge_properties: Add descriptive edge properties for drug mechanism and bioactivity edges (mechanism-of-action description, selectivity comment, binding site, mutation, assay description) once corresponding Biolink Model slots exist; these are currently TODOs in the transform code.
qualifiers: Attach an enzyme/context qualifier to metabolism (has_metabolite) edges to record the mediating enzyme, once Biolink supports the relevant statement-level qualifier.
edge_properties: Consider capturing quantitative bioactivity values (e.g. pchembl_value, standard_value/units) as edge properties on chemical-target edges derived from the activities table.
Provenance Information¶
Contributors: - Vlado Dancik - code author, domain expertise - Kevin Schaper - code support - Evan Morris - code support - Sierra Moxon - data modeling, code support - Matthew Brush - data modeling, domain expertise
Artifacts: - Ingest Survey: https://docs.google.com/spreadsheets/d/1CENLoKukPCHW2SimabAm6HfXeldeRYmPTV_--Fs9whg/edit?gid=0#gid=0