Post-training sparse attention with double sparsity
Unknown
Double Sparsity reduces KV cache access in LLMs via post-training sparse attention combining token and channel sparsity.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Double Sparsity reduces KV cache access in LLMs via post-training sparse attention combining token and channel sparsity.
Jinuk Kim, Marwa El Halabi, Wonpyo Park, et al.
GuidedQuant improves post-training quantization of large language models by using end-loss guidance to better preserve model accuracy.
Unknown
Proposes post-training world models via reinforcement learning to improve their generality and task performance.
Dongfang Li, Xiaodong Luo, Ruoyu Sun, et al.
Full-stack optimization for post-training trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and domain-specialized OR models outperforming GPT-5.4-Mini.
Wayne Xin Zhao, Kun Zhou, Junyi Li, et al.
A comprehensive survey of LLM techniques across pre-training, post-training, utilization, and evaluation, synthesizing state-of-the-art insights and open challenges.