Attention Is All You Need
Ashish Vaswani, Noam Shazeer et al.
71
Citations
8
Influential Citations
—
Venue
2021
Year
… been developed to accelerate the deep learning inference, including model quantization [1… In this work, we focus on model quantization for efficient inference. Quantization targets to …
Model quantization is critical for deploying deep learning models on resource-constrained devices, yet the field lacks standardized evaluation. Mqbench addresses this by providing a reproducible benchmark that allows practitioners to compare quantization methods fairly. This is significant because inconsistent evaluation has hindered progress and practical adoption.
The paper's emphasis on deployability ensures that benchmark results translate to real-world performance, not just academic metrics. This bridges the gap between research and production, making it highly relevant for AI engineers.
The benchmark reveals that post-training quantization often matches quantization-aware training in accuracy for 8-bit precision, but degrades for lower bit-widths. It also shows that hardware-specific optimizations can yield up to 2x speedup without accuracy loss.
Mqbench sets a new standard for quantization research, encouraging reproducible and practical contributions. It helps the AI community identify the most effective quantization techniques, accelerating the deployment of efficient models in edge devices and data centers alike.
Ashish Vaswani, Noam Shazeer et al.
Jakubův, Jan, Chvalovský, Karel et al.
Pauli Virtanen, Ralf Gommers et al.
Tom B. Brown, Benjamin Mann et al.