Rethinking Generalization in Reasoning SFT (April 2026)
FreeChallenges "SFT memorizes, RL generalizes" — reasoning SFT with long CoT does generalize cross-domain, conditional on optimization dynamics; discovers safety-reasoning tradeoff (reasoning improves but safety degrades); 152 HF likes
About Rethinking Generalization in Reasoning SFT (April 2026)
This preprint challenges the prevailing narrative that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes in LLM post-training. Through a conditional analysis of reasoning SFT with long chain-of-thought (CoT) supervision, the authors demonstrate that cross-domain generalization is not absent but conditional—jointly shaped by optimization dynamics, training data quality and structure, and base-model capability. Key findings include a dip-and-recovery pattern (performance first degrades then improves with extended training), the importance of verified long-CoT traces for consistent gains, and an asymmetric generalization where reasoning improves but safety degrades. The work reframes the question from whether reasoning SFT generalizes to under what conditions and at what cost.
Key Features
Pros & Cons
- Provides nuanced empirical analysis challenging oversimplified narratives
- Identifies practical conditions (optimization, data, model) that enable generalization
- Reveals dip-and-recovery pattern, warning against premature evaluation
- Highlights asymmetric generalization, raising important safety considerations
- Based on rigorous experimental methodology with long-CoT supervision
- Preprint under review, not yet peer-reviewed
- Safety degradation is a nontrivial cost of reasoning improvement
- Generalization findings may depend on specific task domains and model scales
- Requires access to verified long-CoT data, which may not be widely available