Sparser is faster and less is more: Efficient sparse attention for long-range transformers
Unknown
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
Unknown
This paper introduces a content-based sparse attention mechanism that combines the benefits of content-based and local temporal sparse attention for improved efficiency in Transformer models.
Unknown
A comprehensive survey of speculative decoding techniques for efficient large language model inference.
Unknown
This paper identifies and addresses suboptimal fine-tuning in LoRA for wide models by proposing Lora+, a method that improves parameter efficiency and performance.
Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, et al.
A controlled study using a procedural testbed reveals that data distribution balance and caption quality critically impact text-to-video model generalization and training efficiency.
Unknown
This paper proposes LongRAG, a framework that leverages long-context LLMs to improve the efficiency and effectiveness of retrieval-augmented generation.
Unknown
Lightrag introduces a graph-based indexing approach to enhance efficiency in retrieval-augmented generation for LLMs.
Unknown
This paper proposes encoder-free vision-language models that directly process visual features without a separate vision encoder, simplifying architecture and improving efficiency.
Unknown
This paper establishes a quantitative foundation showing that small language models can outperform large ones on task-specific efficiency metrics.
Unknown
This paper introduces ShorterBetter, a method to train reasoning models to dynamically determine optimal inference length, improving efficiency without sacrificing accuracy.
Unknown
AdaptThink uses reinforcement learning to teach reasoning models when to engage in extended thinking, improving efficiency and accuracy.
Yuanchun Li, Hao Wen, Weijun Wang, et al.
This paper surveys and analyzes Personal LLM Agents, focusing on their capability, efficiency, and security to guide future development.