How feature learning can improve neural scaling laws
Unknown
Develops a solvable model showing how feature learning improves neural scaling laws beyond the kernel limit.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Develops a solvable model showing how feature learning improves neural scaling laws beyond the kernel limit.
Unknown
This paper proposes a solvable model that explains the empirical neural scaling laws observed in large language models.
Unknown
Proposes a smoothly broken power law functional form (Broken Neural Scaling Law) to accurately model and extrapolate scaling behaviors in deep learning.
Unknown
This paper revisits neural scaling laws for language and vision models, finding that scaling trends hold across modalities but with different exponents and saturation points.
Unknown
This paper demonstrates that data pruning can improve neural scaling laws, reducing the resource costs of deep learning.
Unknown
This paper introduces a dynamical model that reproduces key observations about neural scaling laws, explaining why performance scales with training time and model size.
Unknown
This paper explains neural scaling laws by deriving them from realistic data assumptions within learning theory.
Unknown
Ostquant improves LLM quantization by applying orthogonal and scaling transformations to better fit weight distributions, reducing accuracy loss.
Junsong Chen, Jincheng Yu, Yitong Li, et al.
SANA-Video 2.0 introduces hybrid linear-softmax attention and attention residuals to achieve softmax-level video quality with linear-complexity scaling, enabling 720p generation on a single GPU.
Unknown
A systematic overview of parameter-efficient fine-tuning methods, covering over 50 papers from 2019 to 2024.
Weilin Cai, Juyong Jiang, Fan Wang, et al.
A comprehensive survey on mixture of experts (MoE) as an effective method for scaling model capacity with minimal computation overhead.
Unknown
This survey comprehensively reviews mixture of experts (MoE) methods for scaling large language models efficiently, covering architectures, training, and applications.