Attention Is All You Need
Ashish Vaswani, Noam Shazeer et al.
62
Citations
1
Influential Citations
Electronics
Venue
2024
Year
Large language model-related technologies have shown astonishing potential in tasks such as machine translation, text generation, logical reasoning, task planning, and multimodal alignment. Consequently, their applications have continuously expanded from natural language processing to computer vision, scientific computing, and other vertical industry fields. This rapid surge in research work in a short period poses significant challenges for researchers to comprehensively grasp the research dynamics, understand key technologies, and develop applications in the field. To address this, this paper provides a comprehensive review of research on large language models. First, it organizes and reviews the research background and current status, clarifying the definition of large language models in both Chinese and English communities. Second, it analyzes the mainstream infrastructure of large language models and briefly introduces the key technologies and optimization methods that support them. Then, it conducts a detailed review of the intersections between large language models and interdisciplinary technologies such as contrastive learning, knowledge enhancement, retrieval enhancement, hallucination dissolution, recommendation systems, reinforcement learning, multimodal large models, and agents, pointing out valuable research ideas. Finally, it organizes the deployment and industry applications of large language models, identifies the limitations and challenges they face, and provides an outlook on future research directions. Our review paper aims not only to provide systematic research but also to focus on the integration of large language models with interdisciplinary technologies, hoping to provide ideas and inspiration for researchers to carry out industry applications and the secondary development of large language models.
This survey arrives at a critical juncture where the rapid proliferation of large language models has made it difficult for researchers to maintain a holistic view of the field. By systematically organizing the sprawling literature, the authors provide a much-needed map of LLM research, from foundational architectures to cutting-edge interdisciplinary integrations. The paper's explicit focus on bridging Chinese and English research communities is particularly valuable, as it helps overcome language barriers that often fragment scientific discourse.
The paper's comprehensive scope—covering not only core LLM technologies but also their intersections with contrastive learning, knowledge enhancement, retrieval augmentation, hallucination mitigation, recommendation systems, reinforcement learning, multimodal models, and agents—makes it a one-stop resource for anyone entering the field or seeking to expand their research directions. This breadth is both a strength and a challenge, as it necessarily sacrifices depth in some areas.
As a review paper, the primary result is a structured synthesis of the current state of LLM research. The paper does not present new experimental metrics but instead catalogs existing findings and identifies promising research directions. It highlights that LLMs have expanded from NLP to computer vision, scientific computing, and vertical industries, demonstrating their transformative potential.
This survey serves as a foundational reference for AI practitioners and researchers, enabling them to quickly orient themselves in the rapidly evolving LLM landscape. By systematically mapping the integration of LLMs with interdisciplinary technologies, it provides a roadmap for future innovations. The paper's emphasis on practical deployment and industry applications ensures its relevance beyond academia, potentially accelerating the adoption of LLMs in real-world systems. Its identification of current limitations—such as hallucination, computational cost, and alignment challenges—also helps focus future research efforts on the most pressing problems.
Ashish Vaswani, Noam Shazeer et al.
Jakubův, Jan, Chvalovský, Karel et al.
Pauli Virtanen, Ralf Gommers et al.
Tom B. Brown, Benjamin Mann et al.