Skip to content

ChEMBL Reference Ingest Guide

Source Information

InfoRes ID: infores:chembl

Description: ChEMBL is a large-scale, manually curated database of bioactive molecules with drug-like properties, maintained by the European Bioinformatics Institute (EMBL-EBI). It brings together chemical, bioactivity, and genomic data to aid the translation of genomic information into effective new drugs. ChEMBL captures curated drug mechanisms of action, quantitative bioactivity measurements from assays reported in the literature, drug metabolism relationships, and the protein target composition of complexes and families. Molecules are identified by ChEMBL compound IDs and targets by ChEMBL target IDs (mapped to UniProtKB where applicable).

Citations: - Barbara Zdrazil, Eloy Felix, Fiona Hunter, Emma J Manners, James Blackshaw, Sybilla Corbett, Marleen de Veij, Harris Ioannidis, David Mendez Lopez, Juan F Mosquera, Maria Paula Magarinos, Nicolas Bosc, Ricardo Arcila, Tevik Kizilören, Anna Gaulton, A Patrícia Bento, Melissa. F Adasme, Pater Monecke, Gregory A Landrum, Andrew R Leach. Nucleic Acids Res. 2023: gkad1004. doi: 10.1093/nar/gkad1004 - Davies M, Nowotka M, Papadatos G, Dedman N, Gaulton A, Atkinson F, Bellis L, Overington JP. Nucleic Acids Res. 2015; 43(W1):W612-20, doi: 10.1093/nar/gkv352

Data Access Locations: - https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/latest/

Data Provision Mechanisms: file_download, database_dump

Data Formats: mysql, postgresql, sqlite

Data Versioning and Releases: ChEMBL is released roughly twice per year (approximately semiannual). Releases are versioned with an incrementing integer (e.g. ChEMBL 36). This ingest targets ChEMBL 36. Release notes and the latest distribution are available at https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/latest/

Ingest Information

Ingest Categories: primary_knowledge_provider

Utility: ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.

Scope: Drugs and probes mechanisms, metabolism, bioactivity, gene targets

Relevant Files

File Name Location Description
chembl_36_sqlite.tar.gz https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/latest/ ChEMBL SQLite database

Included Content

File Name Included Records Fields Used
chembl_36_sqlite.tar.gz Drug mechanisms from the following tables: drug_mechanism, molecule_dictionary, target_dictionary, binding_sites, compound_records, docs, source, target_components, component_sequences, variant_sequences. drug_mechanism: mec_id, molregno, mechanism_of_action, action_type, direct_interaction, mechanism_comment, selectivity_comment, binding_site_comment, variant_id; molecule_dictionary: molregno, chembl_id; target_dictionary: tid, chembl_id (AS target_chembl_id), pref_name (AS target_name), target_type, organism (AS target_organism); binding_sites: site_id, site_name; compound_records: record_id, doc_id, src_id; docs: doc_id, chembl_id (AS document_chembl_id); source: src_id, src_description (AS source_description); target_components: tid, component_id; component_sequences: component_id, component_type, accession, description, tax_id, organism; variant_sequences: variant_id, mutation, accession (AS mutation_accession)
chembl_36_sqlite.tar.gz Gene targets - same as drug mechanisms with targets mapped to genes drug_mechanism: mec_id, molregno, mechanism_of_action, action_type, direct_interaction, mechanism_comment, selectivity_comment, binding_site_comment, variant_id; molecule_dictionary: molregno, chembl_id; target_dictionary: tid, chembl_id (AS target_chembl_id), pref_name (AS target_name), target_type, organism (AS target_organism); binding_sites: site_id, site_name; compound_records: record_id, doc_id, src_id; docs: doc_id, chembl_id (AS document_chembl_id); source: src_id, src_description (AS source_description); target_components: tid, component_id; component_sequences: component_id, component_type, accession, description, tax_id, organism; variant_sequences: variant_id, mutation, accession (AS mutation_accession)
chembl_36_sqlite.tar.gz Metabolism information from the following tables: metabolism, molecule_dictionary, compound_records, compound_structures, target_dictionary, metabolism_refs metabolism: drug_record_id, substrate_record_id, metabolite_record_id, met_id, enzyme_name, met_conversion, met_comment, organism, tax_id, enzyme_tid; molecule_dictionary: molregno, chembl_id, pref_name; compound_records: molregno, record_id, compound_name; target_dictionary: tid, target_type, chembl_id; compound_structures: molregno, standard_inchi, standard_inchi_key, canonical_smiles; metabolism_refs: met_id, ref_type, ref_id, ref_url
chembl_36_sqlite.tar.gz Bioactivity and assay information from the following tables: activities, molecule_dictionary, assays, bioassay_ontology, target_dictionary, target_components, component_sequences, cell_dictionary, assay_type, tissue_dictionary, docs, source, ligand_eff, relationship_type, confidence_score_lookup, confidence_score_lookup activities: activity_id, standard_type, standard_relation, standard_value, standard_units, pchembl_value, activity_comment, data_validity_comment, standard_text_value, standard_upper_value, uo_units, potential_duplicate, action_type, src_id, doc_id; molecule_dictionary: molregno, chembl_id; assays: assay_id, chembl_id, description, assay_organism, assay_cell_type, assay_subcellular_fraction, bao_format, assay_category, assay_tax_id, assay_tissue, relationship_type, confidence_score, curated_by, src_id, assay_type, cell_id, tissue_id, tid; bioassay_ontology: bao_id, label; target_dictionary: tid, chembl_id, pref_name, organism, target_type; target_components: tid, component_id; component_sequences: component_id, component_type, accession; cell_dictionary: cell_id, chembl_id, cell_name, cell_description, cell_source_tissue, cell_source_organism, cell_source_tax_id, clo_id, efo_id, cellosaurus_id, cl_lincs_id, cell_ontology_id; assay_type: assay_type, assay_desc; tissue_dictionary: tissue_id, chembl_id, pref_name; docs: doc_id, chembl_id, journal, title, year, authors, pubmed_id, doi; source: src_id, src_description; ligand_eff: activity_id, bei, le, lle, sei; relationship_type: relationship_type, relationship_desc; confidence_score_lookup: confidence_score, description, target_mapping; curation_lookup: curated_by, description

Filtered Content

File Name Filtered Records Rationale
chembl_36_sqlite.tar.gz Gene targets - removed targets that cannot be mapped to genes (like cell-lines) targets that cannot be mapped to genes
chembl_36_sqlite.tar.gz Bioactivity records that are not flagged as 'Active' and lack an action_type, and any activity record whose (molecule, target, action) already corresponds to a curated drug_mechanism record (mec_id is not null) or has already been emitted, are excluded. Only bioactivities with a clear activity outcome are retained; records already represented by a curated drug mechanism edge are dropped to avoid duplicate edges, and repeated (chemical, target, action) combinations are deduplicated.

Future Content Considerations

edge_property_content: ChEMBL drug mechanism and bioactivity records carry descriptive detail - binding site name/comment, mechanism-of-action description and comment, selectivity comment, mutation and mutation accession, and assay description - that is currently held as TODOs in the transform code and not yet emitted as edge properties, pending suitable Biolink Model slots. - Relevant files: chembl_36_sqlite.tar.gz

edge_property_content: Quantitative bioactivity values from the activities table (standard_type, standard_value, standard_units, pchembl_value) are used to select and characterize edges but are not yet represented as edge properties. Consider capturing them to support potency-aware queries. - Relevant files: chembl_36_sqlite.tar.gz

node_property_content: Molecule routes of delivery (derived from the oral/parenteral/topical flags) and additional molecule_dictionary properties (e.g. max_phase, therapeutic_flag, first_approval, withdrawn_flag) are computed or available but not yet emitted as node properties, pending suitable Biolink slots. - Relevant files: chembl_36_sqlite.tar.gz

edge_content: An enzyme/context qualifier linking metabolism edges to the mediating enzyme is computed in code but not yet attached, pending Biolink support for the relevant statement-level qualifier. - Relevant files: chembl_36_sqlite.tar.gz

Target Information

Edge Types

Subject Categories Predicate Object Categories Knowledge Level Agent Type UI Explanation
biolink:ChemicalEntity biolink:Protein, biolink:ProteinFamily, biolink:MacromolecularComplex, biolink:NucleicAcidEntity, biolink:MolecularEntity, biolink:MolecularActivity, biolink:SmallMolecule, biolink:CellularComponent, biolink:CellularOrganism, biolink:CellLine knowledge_assertion manual_agent, automated_agent This edge asserts that a chemical affects a gene or protein target, based on ChEMBL's manually curated drug mechanisms of action and, in some cases, curated bioactivity measurements. The specific ChEMBL 'mechanism of action' (e.g. 'agonist', 'inhibitor', 'stabiliser') determines the qualified predicate and the aspect, direction, and causal-mechanism qualifiers applied to the edge (e.g. an inhibitor causes decreased activity via inhibition). Edges from curated drug mechanisms are manually curated by ChEMBL; edges derived from autocurated bioactivity assays are marked as automated.
biolink:Protein, biolink:MacromolecularComplex, biolink:NucleicAcidEntity, biolink:MolecularEntity, biolink:SmallMolecule biolink:ChemicalEntity knowledge_assertion manual_agent, automated_agent This edge asserts that a gene or protein (for example an enzyme) acts on and chemically transforms a chemical, based on ChEMBL's curated drug mechanisms and bioactivities. It is created for ChEMBL mechanisms where the protein acts on the chemical (e.g. 'hydrolytic enzyme', 'oxidative enzyme', 'proteolytic enzyme', 'reducing agent', 'methylating agent'). The gene/protein is the subject and the chemical the object; object aspect and direction qualifiers describe what happens to the chemical (e.g. increased hydrolysis or oxidation).
biolink:ChemicalEntity biolink:Protein, biolink:ProteinFamily, biolink:MacromolecularComplex, biolink:NucleicAcidEntity, biolink:MolecularEntity, biolink:MolecularActivity, biolink:SmallMolecule, biolink:CellularComponent, biolink:CellularOrganism, biolink:CellLine knowledge_assertion manual_agent, automated_agent This edge asserts a direct physical interaction between a chemical and a gene/protein target, based on ChEMBL's curated drug mechanisms and bioactivities. It is created for ChEMBL mechanisms that imply a physical interaction without a specified functional effect (e.g. 'binding agent', 'chelating agent', 'cross-linking agent', 'sequestering agent'), using the 'directly physically interacts with' predicate. ChEMBL mechanisms classified as 'OTHER', where no specific interaction type is known, produce the more generic 'interacts_with' edge.
biolink:Protein, biolink:ProteinFamily biolink:ChemicalEntity knowledge_assertion manual_agent, automated_agent This edge asserts that a protein or protein family acts on a chemical as its substrate, based on ChEMBL drug mechanisms and bioactivities where the ChEMBL mechanism is 'substrate'. The protein/protein family is the subject and the chemical substrate is the object. Edges from curated drug mechanisms are manual; edges derived from autocurated bioactivity assays are marked as automated.
biolink:MacromolecularComplex, biolink:ProteinFamily, biolink:MolecularActivity, biolink:Protein biolink:Protein knowledge_assertion manual_agent This edge describes the protein composition of a ChEMBL target that is a complex, protein family, or selectivity/interaction group, asserting that the target 'has part' a particular constituent protein. It is derived from ChEMBL's curated target-to-component mappings (target_components and component_sequences), where each protein component of a complex or family target is linked to that target.
biolink:ChemicalEntity biolink:ChemicalEntity knowledge_assertion manual_agent This edge asserts that a chemical (a drug or substrate) has a particular metabolite, based on ChEMBL's manually curated drug metabolism records. It links a substrate chemical to its metabolite; where the parent drug differs from the metabolized substrate, an additional edge from the parent drug to the metabolite is also created. A species context qualifier records the organism in which the metabolic conversion was observed.

Node Types

Node Category Source Identifier Types Additional Notes
biolink:CellLine ChEMBL
biolink:CellularComponent ChEMBL
biolink:CellularOrganism ChEMBL
biolink:ChemicalEntity ChEMBL
biolink:MacromolecularComplex ChEMBL
biolink:MolecularActivity ChEMBL
biolink:MolecularEntity ChEMBL
biolink:NucleicAcidEntity ChEMBL
biolink:Protein ChEMBL, UniProtKB
biolink:ProteinFamily ChEMBL
biolink:SmallMolecule ChEMBL

Future Modeling Considerations

edge_properties: Add descriptive edge properties for drug mechanism and bioactivity edges (mechanism-of-action description, selectivity comment, binding site, mutation, assay description) once corresponding Biolink Model slots exist; these are currently TODOs in the transform code.

qualifiers: Attach an enzyme/context qualifier to metabolism (has_metabolite) edges to record the mediating enzyme, once Biolink supports the relevant statement-level qualifier.

edge_properties: Consider capturing quantitative bioactivity values (e.g. pchembl_value, standard_value/units) as edge properties on chemical-target edges derived from the activities table.

Provenance Information

Contributors: - Vlado Dancik - code author, domain expertise - Kevin Schaper - code support - Evan Morris - code support - Sierra Moxon - data modeling, code support - Matthew Brush - data modeling, domain expertise

Artifacts: - Ingest Survey: https://docs.google.com/spreadsheets/d/1CENLoKukPCHW2SimabAm6HfXeldeRYmPTV_--Fs9whg/edit?gid=0#gid=0