ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
5
Citations
1
Influential Citations
arXiv.org
Venue
2025
Year
The recent advent of reasoning models like OpenAI's o1 was met with excited speculation by the AI community about the mechanisms underlying these capabilities in closed models, followed by a rush of replication efforts, particularly from the open source community. These speculations were largely settled by the demonstration from DeepSeek-R1 that chains-of-thought and reinforcement learning (RL) can effectively replicate reasoning on top of base LLMs. However, it remains valuable to explore alternative methods for theoretically eliciting reasoning that could help elucidate the underlying mechanisms, as well as providing additional methods that may offer complementary benefits. Here, we build on the long-standing literature in cognitive psychology and cognitive architectures, which postulates that reasoning arises from the orchestrated, sequential execution of a set of modular, predetermined cognitive operations. Crucially, we implement this key idea within a modern agentic tool-calling framework. In particular, we endow an LLM with a small set of"cognitive tools"encapsulating specific reasoning operations, each executed by the LLM itself. Surprisingly, this simple strategy results in considerable gains in performance on standard mathematical reasoning benchmarks compared to base LLMs, for both closed and open-weight models. For instance, providing our"cognitive tools"to GPT-4.1 increases its pass@1 performance on AIME2024 from 32% to 53%, even surpassing the performance of o1-preview. In addition to its practical implications, this demonstration contributes to the debate regarding the role of post-training methods in eliciting reasoning in LLMs versus the role of inherent capabilities acquired during pre-training, and whether post-training merely uncovers these latent abilities.
This paper addresses a critical question in AI: how to elicit robust reasoning from large language models (LLMs). While recent work like DeepSeek-R1 has shown that reinforcement learning (RL) and chain-of-thought (CoT) can replicate reasoning, the authors explore an alternative, cognitively inspired approach. By drawing on decades of cognitive psychology and cognitive architectures, they propose that reasoning arises from the orchestrated execution of modular, predetermined cognitive operations. This perspective is not only theoretically interesting but also yields practical gains.
The significance lies in the simplicity and effectiveness of the method. Rather than requiring complex RL training or massive datasets, the authors simply endow an LLM with a small set of "cognitive tools"—each a reasoning operation executed by the LLM itself. This approach is lightweight, model-agnostic, and works for both closed and open-weight models. The paper also contributes to the ongoing debate about whether post-training methods uncover latent abilities from pre-training or genuinely add new capabilities.
This work has both practical and theoretical implications. Practically, it offers a simple, cost-effective way to boost LLM reasoning without retraining. Theoretically, it supports the view that reasoning in LLMs can be elicited by providing appropriate scaffolding, suggesting that these models already possess latent reasoning capabilities from pre-training. The paper also opens avenues for integrating cognitive science insights into AI system design, potentially leading to more interpretable and controllable reasoning processes.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba