PreprintarXiv.org2025
Emergent Hierarchical Reasoning
Haozhe Wang, Qixin Xu, Che Liu, et al.
This paper identifies a two-phase reasoning hierarchy in LLMs trained with RL and proposes HICRA, a credit assignment algorithm that boosts performance by focusing optimization on high-level planning tokens.
41Sep 3, 2025Large Language ModelsReinforcement Learning
arXiv