ConferenceInternational Conference on Machine Learning2024
Star Attention: Efficient LLM Inference over Long Sequences
Shantanu Acharya, Fei Jia, Boris Ginsburg
Star Attention enables efficient LLM inference over long sequences via a two-phase block-sparse approximation that reduces memory and time by up to 11x while preserving 97-100% accuracy.
34Nov 26, 2024Large Language ModelsTransformers
arXiv