Preprint2024
Minitron
Unknown
Pruning and knowledge distillation can compress LLMs 2-4x with up to 40x compute savings and improved performance.
0Jul 1, 2024
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.