ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
737
Citations
29
Influential Citations
IEEE Access
Venue
2024
Year
Large Language Models (LLMs) recently demonstrated extraordinary capability, including natural language processing (NLP), language translation, text generation, question answering, etc. Moreover, LLMs are a new and essential part of computerized language processing, having the ability to understand complex verbal patterns and generate coherent and appropriate replies for the situation. Though this success of LLMs has prompted a substantial increase in research contributions, rapid growth has made it difficult to understand the overall impact of these improvements. Since a lot of new research on LLMs is coming out quickly, it is getting tough to get an overview of all of them in a short note. Consequently, the research community would benefit from a short but thorough review of the recent changes in this area. This article thoroughly overviews LLMs, including their history, architectures, transformers, resources, training methods, applications, impacts, challenges, etc. This paper begins by discussing the fundamental concepts of LLMs with its traditional pipeline of the LLMs training phase. It then provides an overview of the existing works, the history of LLMs, their evolution over time, the architecture of transformers in LLMs, the different resources of LLMs, and the different training methods that have been used to train them. It also demonstrated the datasets utilized in the studies. After that, the paper discusses the wide range of applications of LLMs, including biomedical and healthcare, education, social, business, and agriculture. It also illustrates how LLMs create an impact on society and shape the future of AI and how they can be used to solve real-world problems. Then it also explores open issues and challenges to deploying LLMs in real-world scenario. Our review paper aims to help practitioners, researchers, and experts thoroughly understand the evolution of LLMs, pre-trained architectures, applications, challenges, and future goals.
This survey arrives at a critical juncture in AI research, where the rapid proliferation of large language models has made it difficult for even experts to maintain a holistic view. By consolidating knowledge on LLM history, architectures, training methods, and applications, the paper provides a much-needed map of the field. Its comprehensive scope—from foundational transformer concepts to domain-specific uses in healthcare and agriculture—makes it an essential starting point for newcomers and a useful reference for seasoned practitioners.
The paper's emphasis on open issues and challenges, such as deployment robustness, bias, and computational costs, highlights the practical hurdles that must be overcome for LLMs to achieve widespread real-world impact. This focus on real-world applicability distinguishes it from purely theoretical surveys.
As a review paper, the primary result is a structured synthesis of the LLM literature. The paper does not report new experimental metrics but instead aggregates findings from hundreds of prior works. It notes that transformer-based architectures dominate the field, with models like GPT, BERT, and their variants achieving state-of-the-art performance on NLP benchmarks. The survey also quantifies the growth in LLM research, citing 737 citations as evidence of the field's rapid expansion.
This paper serves as a foundational reference for the AI community, enabling researchers and practitioners to quickly grasp the current state of LLMs and identify promising research directions. By highlighting both achievements and unresolved challenges, it encourages a balanced approach to LLM development—one that pursues performance gains while addressing ethical and practical concerns. Its broad coverage ensures relevance across multiple subfields, from NLP to applied AI in specialized domains.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba