Attention Is All You Need
Ashish Vaswani, Noam Shazeer et al.
4.7k
Citations
210
Influential Citations
Frontiers of Computer Science
Venue
2026
Year
Abstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities, LLMs necessitate new frameworks for understanding their development, behavior, and societal impact. This survey systematically reviews recent advancements in LLM techniques across four key dimensions: (1) pre-training methodologies, which establish core model capabilities through large-scale self-supervised training, architectural innovations, and data curation strategies; (2) post-training techniques, including supervised fine-tuning and reinforcement learning, which adapt foundational models to downstream tasks and enhance their alignment and safety; (3) utilization strategies, such as in-context learning, prompt engineering, and agentic reasoning, that optimize real-world deployment and enable effective interaction with external environments; and (4) evaluation methods, encompassing benchmarks for key ability dimensions such as core language capabilities, reasoning, and safety, which support comprehensive and reliable assessment of model performance. Additionally, we identify critical research issues, including those concerning theoretical foundations, efficient scaling, alignment, and agentic capability, and highlight the open challenges they present. By synthesizing state-of-the-art insights and emerging trends, this survey aims to provide a systematic and comprehensive framework for understanding the trajectory, current limitations, and future directions of LLM progress.
This survey arrives at a critical juncture in AI research, where large language models have become central to both academic inquiry and industrial deployment. With 4719 citations, it has already established itself as a foundational reference. The paper's systematic organization across pre-training, post-training, utilization, and evaluation offers a much-needed structured lens for navigating the rapidly expanding LLM landscape. By explicitly addressing open challenges such as theoretical foundations and efficient scaling, it helps the community prioritize research directions.
The paper's main technical contribution is its comprehensive taxonomy of LLM techniques:
As a survey, the paper does not report new experimental results. Instead, it synthesizes existing findings across the field, noting that LLMs have achieved unprecedented performance on language understanding, generation, and reasoning tasks. It highlights that evaluation benchmarks now cover dimensions like safety and alignment, reflecting the growing emphasis on responsible AI. The paper does not provide specific numerical comparisons but references the broader trend of scaling laws and emergent abilities.
This survey serves as a critical roadmap for AI practitioners and researchers. By framing LLM progress in terms of pre-training, post-training, utilization, and evaluation, it clarifies the distinct phases of model development and deployment. Its identification of open challenges—especially alignment and agentic capability—points to where future breakthroughs are needed. The paper's high citation count underscores its role as a key reference, likely influencing both academic research and industry best practices for years to come.
Ashish Vaswani, Noam Shazeer et al.
Jakubův, Jan, Chvalovský, Karel et al.
Pauli Virtanen, Ralf Gommers et al.
Tom B. Brown, Benjamin Mann et al.