ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
Trains an LLM using reinforcement learning with a verifiable reward (e.g., a binary correct/incorrect signal) on only a single training example, surprisingly achieving comparable performance to training on much larger datasets in mathematical reasoning tasks.
This paper challenges a core assumption in LLM training: that large datasets are necessary for effective reinforcement learning with verifiable rewards. By showing that a single example can suffice for mathematical reasoning, it opens the door to extremely data-efficient fine-tuning. This is particularly significant for domains where collecting large labeled datasets is expensive or impractical.
The finding also has implications for rapid iteration and personalization of LLMs. If a model can be adapted to a new task with just one example, it reduces the barrier to entry for specialized applications. However, the paper's brevity leaves open questions about the robustness of this approach across different tasks and model scales.
The paper reports that one-shot RLVR achieves performance comparable to training on much larger datasets in mathematical reasoning tasks. No specific metrics (e.g., accuracy percentages, benchmark names) are provided in the abstract, so the exact magnitude of the result is unclear. The claim is surprising and warrants further investigation.
If validated, this work could transform how LLMs are fine-tuned for tasks with verifiable rewards (e.g., math, code, logic). It suggests that the RLVR signal is powerful enough to generalize from minimal data, potentially reducing training costs and enabling on-the-fly adaptation. However, the lack of detail in the abstract means the community should treat this as a preliminary finding until full results are available.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba