ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs'metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion. An organized list of papers can be found at https://github.com/yale-nlp/LLM-Metacognition.
Metacognition—the ability to monitor and control one's own cognitive processes—is a cornerstone of human intelligence and is increasingly recognized as essential for building capable, transparent AI systems. While large language models (LLMs) have achieved remarkable performance across diverse tasks, their ability to exhibit metacognitive behaviors such as self-assessment, uncertainty estimation, and strategic reasoning remains poorly understood. This paper addresses this critical gap by providing the first comprehensive overview of the emerging field of metacognition in LLMs. It systematically organizes the current state of knowledge, making it an indispensable resource for researchers and practitioners aiming to enhance AI reliability and self-awareness.
The significance of this work lies in its timely synthesis of a fragmented research area. As LLMs are deployed in high-stakes applications like healthcare, law, and education, understanding their metacognitive limitations becomes crucial for safety and trust. By taxonomizing methods, benchmarks, and techniques, the paper lays the groundwork for systematic progress, helping the community identify what works, what doesn't, and where to focus future efforts.
The paper's main technical contributions are its comprehensive taxonomy and structured review of the literature. Key innovations include:
As a survey, the paper does not present new experimental results. Instead, it aggregates findings from prior studies, noting that LLMs exhibit some metacognitive abilities—e.g., they can provide confidence scores that correlate with accuracy—but these abilities are often brittle and context-dependent. For instance, calibration can be improved with specific prompts but degrades under distribution shift. The paper highlights that no single method consistently outperforms others, and many benchmarks lack standardization, making cross-study comparisons difficult.
This survey has broad implications for the AI field. By systematically organizing knowledge on LLM metacognition, it provides a roadmap for developing more reliable and transparent AI systems. It underscores the need for robust evaluation frameworks and encourages the community to treat metacognition as a first-class capability rather than an emergent byproduct. The paper also highlights open challenges, such as defining metacognition in non-human agents and ensuring that metacognitive behaviors generalize across tasks. Ultimately, this work could accelerate progress toward AI systems that can self-monitor, self-correct, and communicate their uncertainty—key steps toward trustworthy AI.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba