ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
212
Citations
0
Influential Citations
Chemical Science
Venue
2024
Year
Large language models (LLMs) have emerged as powerful tools in chemistry, significantly impacting molecule design, property prediction, and synthesis optimization. This review highlights LLM capabilities in these domains and their potential to accelerate scientific discovery through automation. We also review LLM-based autonomous agents: LLMs with a broader set of tools to interact with their surrounding environment. These agents perform diverse tasks such as paper scraping, interfacing with automated laboratories, and synthesis planning. As agents are an emerging topic, we extend the scope of our review of agents beyond chemistry and discuss across any scientific domains. This review covers the recent history, current capabilities, and design of LLMs and autonomous agents, addressing specific challenges, opportunities, and future directions in chemistry. Key challenges include data quality and integration, model interpretability, and the need for standard benchmarks, while future directions point towards more sophisticated multi-modal agents and enhanced collaboration between agents and experimental methods. Due to the quick pace of this field, a repository has been built to keep track of the latest studies: https://github.com/ur-whitelab/LLMs-in-science.
This review arrives at a critical juncture where large language models are transitioning from pure language tasks to scientific discovery. By systematically covering LLM applications in chemistry—molecule design, property prediction, and synthesis optimization—it provides a much-needed map for researchers navigating this fast-moving field. The inclusion of autonomous agents, which combine LLMs with tools like web scraping and lab automation, extends the relevance beyond chemistry to any domain where AI can interact with physical or digital environments.
The paper's emphasis on challenges such as data quality, model interpretability, and the absence of standard benchmarks is particularly valuable. These are the bottlenecks that currently limit real-world adoption. The authors also point to future directions like multi-modal agents, which could integrate text, images, and experimental data, and enhanced collaboration between agents and experimental methods. This forward-looking perspective helps practitioners prioritize research investments.
The paper does not present new experimental results but synthesizes findings from 212 cited works. It reports that LLMs have shown strong performance in property prediction and synthesis planning, though quantitative benchmarks are still emerging. The review notes that autonomous agents have successfully automated paper scraping and lab interfacing, but standardized evaluation metrics are lacking. The repository is actively maintained, reflecting the field's dynamism.
For AI practitioners, this review serves as both a primer and a roadmap. It clarifies where LLMs and agents currently excel (e.g., property prediction, synthesis planning) and where they fall short (e.g., interpretability, data integration). The identification of standard benchmarks as a key need is a call to action for the community. By linking to a living repository, the paper ensures its relevance over time. This work will likely influence how researchers design and evaluate LLM-based tools for scientific discovery, accelerating the adoption of AI in chemistry and beyond.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba