Preprint2026
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun, et al.
Full-stack optimization for post-training trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and domain-specialized OR models outperforming GPT-5.4-Mini.
0Jul 22, 2026Reasoning
arXiv