Dynamic sparse attention for scalable transformer acceleration
Unknown
Proposes Dynamic Sparse Attention (DSA) to efficiently exploit dynamic sparse patterns in attention for scalable transformer acceleration.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Proposes Dynamic Sparse Attention (DSA) to efficiently exploit dynamic sparse patterns in attention for scalable transformer acceleration.
Unknown
FastKV decouples context reduction from KV cache compression to accelerate both prefill and decoding phases in LLM inference.
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
Unknown
This survey comprehensively reviews knowledge distillation techniques for model compression and acceleration.
Hongzheng Chen, Jiahao Zhang, Yixiao Du, et al.
This paper investigates FPGA-based spatial acceleration for LLM inference, achieving up to 13.4x speedup over prior FPGA accelerators and 5.7x energy efficiency vs. A100 GPU.