
Same Model, Different Answers: How Inference Backends and Quantization Quietly Break Tool Calls
A new investigation reveals that the same LLM weights can produce different tokens depending on the inference backend, quantization format, or CUDA kernel used. Tests on Qwen3.6-27B show token flips leading to incorrect tool calls and botched Cisco commands. KV-cache quantization and fine-tune comparisons further highlight significant performance variations across setups.
