Model quantization and hardware acceleration for vision transformers: A comprehensive survey
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
Unknown
Intactkv improves large language model quantization by preserving pivot tokens, achieving state-of-the-art results.
Jinuk Kim, Marwa El Halabi, Wonpyo Park, et al.
GuidedQuant improves post-training quantization of large language models by using end-loss guidance to better preserve model accuracy.
Unknown
This paper introduces Llmc, a versatile toolkit for benchmarking LLM quantization methods, enabling standardized evaluation and comparison.
Tianyi Zhang, Anshumali Shrivastava
Leanquant introduces a loss-error-aware grid quantization method for LLMs that achieves accurate and scalable compression by minimizing quantization error with respect to the model's loss function.
Yuhang Li, Mingzhu Shen, Yan Ren, et al.
Mqbench provides a reproducible and deployable benchmark for model quantization to accelerate deep learning inference.
Unknown
Ostquant improves LLM quantization by applying orthogonal and scaling transformations to better fit weight distributions, reducing accuracy loss.
Unknown
Proposes a retraining-free model quantization method using one-shot weight-coupling learning to achieve efficient compression without fine-tuning.
Unknown
This survey provides a comprehensive overview of low-bit model quantization techniques for deep neural networks, including a curated list of resources.
Unknown
This survey comprehensively reviews model quantization techniques for deep neural networks in image classification, covering methods, challenges, and future directions.
Panagiotis Papantonakis, Georgios Kopanas, Bernhard Kerbl, et al.
This paper reduces the memory footprint of 3D Gaussian splatting by 27x via resolution-aware pruning, adaptive spherical harmonic coefficients, and codebook quantization.