LIMO: Less is More for Reasoning logo

LIMO: Less is More for Reasoning

Free

Less is More for Reasoning

FreeFree tier
Inputs: textOutputs: text
Type
Open Source

About LIMO: Less is More for Reasoning

LIMO (Less is More for Reasoning) is a research model that demonstrates sophisticated mathematical reasoning in large language models can emerge with minimal training data. Through simple supervised fine-tuning, LIMO achieves state-of-the-art results on AIME24 (63.3%) and MATH500 (95.6%), outperforming previous models that used 100x more data. It proposes the Less-Is-More Reasoning Hypothesis, suggesting that complex reasoning can be elicited with few cognitive template demonstrations if the model has a comprehensive pre-trained knowledge base.

Key Features

Achieves 63.3% accuracy on AIME24 with minimal training data
95.6% accuracy on MATH500
Uses only 1% of training data compared to prior approaches
Strong out-of-distribution generalization (45.8% absolute improvement across benchmarks)
Proposes Less-Is-More Reasoning Hypothesis
Based on simple supervised fine-tuning

Pros & Cons

Pros
  • Requires significantly less training data than previous methods
  • Achieves high accuracy on challenging math benchmarks
  • Exhibits strong generalization to unseen tasks
  • Provides theoretical hypothesis for reasoning emergence
Cons
  • Focuses primarily on mathematical reasoning; may not generalize to other domains
  • Requires a strong pre-trained foundation model with comprehensive knowledge
  • Limited to supervised fine-tuning; not tested with other methods

Best For

Mathematical reasoning tasksLLM fine-tuning researchDemonstrating data efficiency in reasoningOut-of-distribution generalization benchmarks

FAQ

What is the LIMO hypothesis?
The Less-Is-More Reasoning Hypothesis states that in foundation models where domain knowledge is comprehensively encoded during pre-training, sophisticated reasoning can emerge through minimal but strategically designed demonstrations of cognitive processes.
How much training data does LIMO use?
LIMO uses only 1% of the training data required by previous fine-tuned models.