QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory
FreePost-training recipe for long-context reasoning and memory management
About QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory
QwenLong-L1.5 is a post-training recipe that significantly enhances the long-context reasoning capabilities of the Qwen3-30B-A3B-Thinking base model. It introduces three key technical innovations: (1) a Long-Context Data Synthesis Pipeline that generates challenging multi-hop reasoning tasks by decomposing documents into atomic facts and composing verifiable questions; (2) Stabilized Reinforcement Learning for Long-Context Training using task-balanced sampling, task-specific advantage estimation, and Adaptive Entropy-Controlled Policy Optimization (AEPO); and (3) a Memory-Augmented Architecture for Ultra-Long Contexts that integrates single-pass reasoning with iterative memory-based processing for sequences exceeding 4 million tokens. The model achieves performance comparable to GPT-5 and Gemini-2.5-Pro on long-context reasoning benchmarks, and shows gains in scientific reasoning, memory tool use, and extended dialogue.
Key Features
Pros & Cons
- Significant improvement in long-context reasoning over baseline Qwen3-30B-A3B-Thinking
- Achieves performance comparable to state-of-the-art models GPT-5 and Gemini-2.5-Pro
- Memory architecture enables processing of ultra-long sequences beyond 4M tokens
- Systematic data synthesis approach provides high-quality training data for complex reasoning
- Enhanced performance transfers to general domains like scientific reasoning and extended dialogue