I measure how much KV cache quantization (F16/Q8/Q4) hurts Qwen3.8-27B output quality via KL divergence, and how that compares to the impact of quantizing the model weights themselves.
Curious about this picture? These are Tulipa gesneriana, red tulips. This picture was taken on the northwest coast of the Washington state, and then a voronoi blur applied on it.