Preprint
Reinforcement Learning

Principles of early drug discovery

J. P. HUGHES, S. Rees(GlaxoSmithKline (United Kingdom)), S. Barret Kalindjian(King's College London), K.L. Philpott(King's College London)
November 22, 2010British Journal of Pharmacology2,842 citations

2.8k

Citations

80

Influential Citations

British Journal of Pharmacology

Venue

2010

Year

Abstract

Developing a new drug from original idea to the launch of a finished product is a complex process which can take 12-15 years and cost in excess of $1 billion. The idea for a target can come from a variety of sources including academic and clinical research and from the commercial sector. It may take many years to build up a body of supporting evidence before selecting a target for a costly drug discovery programme. Once a target has been chosen, the pharmaceutical industry and more recently some academic centres have streamlined a number of early processes to identify molecules which possess suitable characteristics to make acceptable drugs. This review will look at key preclinical stages of the drug discovery process, from initial target identification and validation, through assay development, high throughput screening, hit identification, lead optimization and finally the selection of a candidate molecule for clinical development.

Analysis

Why This Paper Matters

This review is a cornerstone reference for anyone entering or working in pharmaceutical R&D. It demystifies the complex, multi-year journey from a biological target to a clinical candidate, a process that is both time-consuming and financially intensive. By consolidating key preclinical stages—target identification, assay development, high-throughput screening, hit identification, and lead optimization—the paper provides a clear roadmap that bridges academic research and industrial application. Its high citation count (2842) underscores its utility as a teaching tool and a practical guide for drug discovery teams.

The paper matters because it highlights the critical decision points where resources are committed and where failures often occur. Understanding these early steps is essential for improving success rates in an industry where only a fraction of candidates make it to market. For AI practitioners, this context is vital: machine learning models are increasingly applied to predict target-druggability, optimize screening libraries, and prioritize leads, making this foundational knowledge indispensable.

Technical Contributions

The paper's main technical contribution is its structured synthesis of the drug discovery pipeline. Key innovations include:

  • Target Identification and Validation: Discusses criteria for selecting biological targets (e.g., disease linkage, druggability) and methods like gene knockout, RNA interference, and chemical probes.
  • Assay Development: Covers design of biochemical and cell-based assays for high-throughput screening, including considerations for robustness, reproducibility, and physiological relevance.
  • High-Throughput Screening (HTS): Describes automation, miniaturization, and data analysis techniques to test millions of compounds efficiently.
  • Hit Identification and Lead Optimization: Explains how hits are confirmed, validated, and iteratively modified to improve potency, selectivity, and pharmacokinetic properties.
  • Candidate Selection: Outlines the criteria for advancing a molecule into clinical trials, including safety, efficacy, and formulation.

Results

As a review, the paper does not present novel experimental results. Instead, it aggregates industry benchmarks: typical timelines of 12-15 years and costs exceeding $1 billion per drug. It notes that only about 10% of candidates entering clinical trials ultimately receive approval. The paper emphasizes that early-stage decisions heavily influence downstream success, though no specific metrics are provided for individual stages.

Significance

The broader impact of this paper lies in its role as a standard reference for drug discovery education and practice. It has informed the design of academic curricula, industry training programs, and collaborative frameworks between academia and pharma. For the AI field, this paper provides the domain context needed to develop machine learning models for target prediction, virtual screening, and lead optimization. Its clear articulation of the pipeline helps AI researchers identify where their tools can add the most value—such as improving hit rates in HTS or predicting ADMET properties—thereby accelerating the translation of computational methods into real-world drug development.