MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention logo

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Free
FreeFree tier
Inputs: textOutputs: text
Type
Open Source
Company
MiniMax

About MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

MiniMax-M1 is an open-weight, large-scale hybrid-attention reasoning model developed by MiniMax. It is built on a hybrid Mixture-of-Experts (MoE) architecture with a lightning attention mechanism, based on the MiniMax-Text-01 model (456B total parameters, 45.9B activated per token). The model natively supports a context length of 1 million tokens, 8 times that of DeepSeek-R1, and enables efficient scaling of test-time compute. It is trained using large-scale reinforcement learning (RL) on diverse problems including sandbox-based and real-world software engineering environments. The training process is accelerated by the proposed CISPO RL algorithm and completed on 512 H800 GPUs in three weeks with a rental cost of $534,700. Two versions are released with 40K and 80K thinking budgets, achieving comparable or superior performance to models like DeepSeek-R1 and Qwen3-235B on standard benchmarks, particularly excelling in complex software engineering tasks.

Key Features

Hybrid Mixture-of-Experts (MoE) architecture with lightning attention mechanism
1 million token context length (8x DeepSeek-R1)
Efficient scaling of test-time compute via lightning attention
Trained with large-scale reinforcement learning on diverse problems including software engineering
CISPO reinforcement learning algorithm for improved training efficiency
Full training on 512 H800 GPUs in three weeks at $534,700 rental cost
Two versions with 40K and 80K thinking budgets released

Pros & Cons

Pros
  • Open-weight model with state-of-the-art performance
  • Massive 1 million token context length
  • Efficient training with reduced cost and time
  • Strong performance on complex software engineering benchmarks

Best For

Complex reasoning tasks requiring long context processingSoftware engineering tasks (e.g., sandbox-based and real-world environments)

FAQ

What is MiniMax-M1?
MiniMax-M1 is the world's first open-weight, large-scale hybrid-attention reasoning model, developed by MiniMax. It uses a hybrid Mixture-of-Experts architecture with lightning attention and supports 1 million token context.
How was MiniMax-M1 trained?
It was trained using large-scale reinforcement learning on diverse problems including sandbox-based and real-world software engineering environments, with the CISPO algorithm to enhance training efficiency. The full training on 512 H800 GPUs completed in three weeks at a cost of $534,700.
What context length does MiniMax-M1 support?
MiniMax-M1 natively supports a context length of 1 million tokens, which is 8 times the context size of DeepSeek-R1.