Preprint2024
Sparser is faster and less is more: Efficient sparse attention for long-range transformers
Chao Lou, Zixia Jia, Zilong Zheng, et al.
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
0Jun 1, 2024TransformersAttention Mechanisms
arXiv