Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
Erik Johannes Husom, Arda Göknil, Merve Astekin, et al.
This paper evaluates 28 quantized LLMs on a Raspberry Pi 4, measuring energy efficiency, accuracy, and latency to identify optimal configurations for sustainable edge AI deployment.