Preprint
Machine Learning

ChEMBL: a large-scale bioactivity database for drug discovery

Anna Gaulton(European Bioinformatics Institute), Louisa J. Bellis(European Bioinformatics Institute), A. Patrícia Bento(European Bioinformatics Institute), Jon Chambers(European Bioinformatics Institute), Mark Davies(European Bioinformatics Institute), Anne Hersey(European Bioinformatics Institute), Yvonne Light(European Bioinformatics Institute), Shaun McGlinchey(European Bioinformatics Institute), David Michalovich(European Bioinformatics Institute), Bissan Al‐Lazikani(European Bioinformatics Institute), John P. Overington(European Bioinformatics Institute)
September 23, 2011Nucleic Acids Research4,495 citations

4.5k

Citations

298

Influential Citations

Nucleic Acids Research

Venue

2011

Year

Abstract

ChEMBL is an Open Data database containing binding, functional and ADMET information for a large number of drug-like bioactive compounds. These data are manually abstracted from the primary published literature on a regular basis, then further curated and standardized to maximize their quality and utility across a wide range of chemical biology and drug-discovery research problems. Currently, the database contains 5.4 million bioactivity measurements for more than 1 million compounds and 5200 protein targets. Access is available through a web-based interface, data downloads and web services at: https://www.ebi.ac.uk/chembldb.

Analysis

Why This Paper Matters

ChEMBL addresses a critical bottleneck in drug discovery: the lack of large-scale, high-quality, and openly accessible bioactivity data. Prior to ChEMBL, most bioactivity data was scattered across proprietary databases or buried in unstructured literature, making it difficult for researchers to build predictive models or perform large-scale analyses. By providing a centralized, manually curated repository, ChEMBL democratizes access to drug-target interaction data, accelerating both academic research and industrial drug development.

The database's emphasis on open data and standardization is particularly significant for the machine learning community. High-quality training data is essential for developing accurate predictive models, and ChEMBL has become a standard benchmark for tasks such as bioactivity prediction, compound-target interaction modeling, and ADMET property estimation. Its regular updates and broad coverage of targets and compounds ensure its continued relevance.

Technical Contributions

  • Manual Curation Pipeline: Data is systematically extracted from primary literature by trained curators, ensuring accuracy and consistency.
  • Standardization: Compounds and targets are standardized using controlled vocabularies and chemical structure normalization, enabling cross-study comparisons.
  • Comprehensive Coverage: Includes binding affinities (e.g., IC50, Ki), functional assays (e.g., agonist/antagonist), and ADMET properties (e.g., solubility, toxicity).
  • Open Access: Data is freely available via web interface, bulk downloads, and RESTful web services, facilitating integration into computational workflows.
  • Scale: With 5.4 million measurements, 1 million compounds, and 5200 targets, ChEMBL provides one of the largest public bioactivity datasets.

Results

The paper reports that ChEMBL contains 5.4 million bioactivity measurements for more than 1 million distinct compounds and 5200 protein targets. These data are manually abstracted from the primary literature and curated to maximize quality. The database is accessible at https://www.ebi.ac.uk/chembldb. No specific performance metrics or comparisons to other databases are provided in the abstract.

Significance

ChEMBL has had a transformative impact on computational drug discovery and cheminformatics. It has enabled the development of machine learning models for predicting drug-target interactions, compound bioactivity, and ADMET properties. The database is widely used in both academia and industry, serving as a benchmark for evaluating new algorithms and as a training resource for deep learning models. Its open data philosophy has inspired similar initiatives and fostered a culture of data sharing in the drug discovery community.