Preprint2024
Post-training sparse attention with double sparsity
Unknown
Double Sparsity reduces KV cache access in LLMs via post-training sparse attention combining token and channel sparsity.
0Jan 1, 2024Attention Mechanisms
arXiv
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Double Sparsity reduces KV cache access in LLMs via post-training sparse attention combining token and channel sparsity.
Unknown
Seerattention introduces a block-sparse attention kernel that achieves a 7.3× speedup at 90% sparsity on 128k sequences, enabling efficient long-context LLMs.