Meta's paper "Quantized Reasoning Models Think They Need to Think Longer, but They Do Not" didn't cover the quantizations llama.cpp actually ships. Here's what happens when we apply its overthinking-marker logit penalties across GGUF quantization levels.
Ran Qwen3.5-4B (bartowski GGUF) on
50 random MATH-500 problems (HuggingFaceH4/MATH-500),
applying a penalty of -2 (via --logit-bias) to 50 tokens the Meta paper identified as
"overthinking" markers — hedges, self-doubt, and backtracking phrases like perhaps, wait,
reconsider, and incorrect.
Full flag list used:
--logit-bias 466-2 --logit-bias 694-2 --logit-bias 1362-2 \ --logit-bias 1412-2 --logit-bias 1921-2 --logit-bias 1990-2 \ --logit-bias 2086-2 --logit-bias 2361-2 --logit-bias 2441-2 \ --logit-bias 2493-2 --logit-bias 2892-2 --logit-bias 3222-2 \ --logit-bias 3315-2 --logit-bias 3384-2 --logit-bias 3404-2 \ --logit-bias 3482-2 --logit-bias 3655-2 --logit-bias 4213-2 \ --logit-bias 4370-2 --logit-bias 4598-2 --logit-bias 4611-2 \ --logit-bias 4808-2 --logit-bias 5752-2 --logit-bias 6970-2 \ --logit-bias 7014-2 --logit-bias 7643-2 --logit-bias 8106-2 \ --logit-bias 10179-2 --logit-bias 10451-2 --logit-bias 11746-2 \ --logit-bias 13264-2 --logit-bias 13428-2 --logit-bias 14673-2 \ --logit-bias 15029-2 --logit-bias 16036-2 --logit-bias 21143-2 \ --logit-bias 21979-2 --logit-bias 33955-2 --logit-bias 35999-2 \ --logit-bias 36563-2 --logit-bias 37201-2 --logit-bias 37781-2 \ --logit-bias 41484-2 --logit-bias 62586-2 --logit-bias 66073-2 \ --logit-bias 73071-2 --logit-bias 84485-2 --logit-bias 85152-2 \ --logit-bias 95500-2
These correspond to the paper's overthinking markers:
Surprisingly, even BF16 gets better accuracy when the penalties are applied — so it's not just a quantization artifact. And the reasoning-token count drops everywhere, meaning the models stop overthinking.
| Format | Accuracy: baseline → penalty | Reasoning tokens |
|---|
Every quantization level improves when overthinking tokens are penalized.
Q3_K_M and Q2_K gain the most — quantization damage appears partially recoverable this way.
Percent decrease in reasoning tokens. BF16 thinks 19.4% less — the least quantized model was overthinking the most.