Preprint
Machine Learning

Machine learning in chemoinformatics and drug discovery

Yu‐Chen Lo(Stanford University), Stefano Rensi(Stanford University), Wen Torng(Stanford University), Russ B. Altman(Stanford University)
May 8, 2018Drug Discovery Today964 citations

964

Citations

16

Influential Citations

Drug Discovery Today

Venue

2018

Year

Abstract

Chemoinformatics is an established discipline focusing on extracting, processing and extrapolating meaningful data from chemical structures. With the rapid explosion of chemical 'big' data from HTS and combinatorial synthesis, machine learning has become an indispensable tool for drug designers to mine chemical information from large compound databases to design drugs with important biological properties. To process the chemical data, we first reviewed multiple processing layers in the chemoinformatics pipeline followed by the introduction of commonly used machine learning models in drug discovery and QSAR analysis. Here, we present basic principles and recent case studies to demonstrate the utility of machine learning techniques in chemoinformatics analyses; and we discuss limitations and future directions to guide further development in this evolving field.

Analysis

Why This Paper Matters

This paper is a pivotal review that bridges the gap between chemoinformatics and machine learning, published at a time when chemical big data from high-throughput screening (HTS) and combinatorial synthesis were rapidly expanding. It provides a structured framework for understanding how chemical structures are processed and how ML models can be applied to extract meaningful biological insights. For AI practitioners, it offers a clear entry point into the domain of drug discovery, highlighting the unique challenges of representing chemical data and the potential of ML to accelerate drug design.

The paper's significance is underscored by its high citation count (964), indicating its widespread use as a reference in both academia and industry. It systematically covers the chemoinformatics pipeline—from data acquisition to feature engineering—and then maps various ML models (e.g., random forests, neural networks) to specific drug discovery tasks. This dual focus makes it valuable for both cheminformaticians looking to adopt ML and ML researchers seeking domain context.

Technical Contributions

The paper's technical contributions include:

  • Pipeline overview: Detailed description of data processing layers, including compound representation (2D/3D descriptors, fingerprints) and data curation.
  • ML model taxonomy: Categorization of models such as support vector machines, random forests, and deep learning, with guidance on their applicability to QSAR and virtual screening.
  • Case studies: Real-world examples demonstrating ML's utility in predicting biological activity, ADMET properties, and target interactions.
  • Future directions: Discussion of challenges like data imbalance, model interpretability, and integration of physics-based methods.

Results

As a review, the paper does not present new experimental metrics. Instead, it synthesizes findings from prior studies, showing that ML models often outperform traditional statistical methods in QSAR tasks. It highlights specific successes, such as improved hit identification in HTS campaigns and more accurate property prediction, though exact numbers are not provided in the abstract. The paper's value lies in its qualitative assessment of model performance and practical recommendations.

Significance

The paper has had a lasting impact on the AI-for-drug-discovery field by providing a common language and framework for researchers. It helped legitimize ML as an indispensable tool in chemoinformatics, encouraging further investment in deep learning approaches. Its discussion of limitations—such as the need for larger, higher-quality datasets and better model interpretability—has guided subsequent research directions. For AI practitioners, it remains a foundational reference for understanding the intersection of chemical data and machine learning, and it underscores the importance of domain-aware feature engineering in achieving meaningful results.