My model was fine at 13 GB and useless at 7 — and perplexity never noticed.
Every published quantization ladder gives you file sizes and a fidelity score. Neither tells you the thing you actually choose on: whether that rung can still do anything. So I measured it, rung by rung, on tasks scored by running the code the model writes. There is a cliff. It is sharp, it is not where I expected, and the metrics everyone publishes are blind to it.
A quantized model is a compression trade: you give up precision to fit the thing on your machine. The community publishes that trade as a table of file sizes, usually with perplexity or KL-divergence against the original [1]. Smaller file, slightly worse number, pick your row.
That table answered a question I was not asking. I wanted to know which row still works as an agent — reads a file, runs a command, writes the fix. So I built the smallest honest test I could: multi-step tool tasks where a script verifies the artifact the model produced. No judge model, no partial credit. Either the file has the right content or it does not.
The cliff
Between 2.93 and 2.23 bits per weight this model goes from doing essentially everything to doing about a third. Above the line it is an agent; below it, a fluent text generator that does not reliably act.
Two details make that table worth more than its headline:
The rung just under the top is the one to take. At 2.93 bpw the file is 28 % smaller than the 3.40-bpw build and scores the same. That is the whole compression win, and it sits one step below where I would have stopped if I had been reading fidelity numbers.
The rung below that buys nothing. 2.66 bpw is 2 % smaller and measurably worse — six tasks where the bigger build wins, zero the other way, p = 0.031 on a paired test. It is strictly dominated. You would never see that in a size table.
Why perplexity misses it
This is the part I did not expect, and it took an outside reviewer telling me twice before I believed it.
The 3.40-bpw build disagrees with the full-precision original about the next token — the actual argmax — for 9.5 % of positions. Over a rollout of ~1500 tokens, the chance of getting through with zero disagreements is e−150. So its trajectories diverge from the original essentially always. And it solves 40 out of 40 tasks.
A 10 % per-token divergence rate is compatible with a perfect task score. Fidelity to the original and capability at the task are different axes. At the one point where I have both numbers, fidelity is measurably non-predictive.
There is a mechanical reason on top of the empirical one. The standard tool for this, llama-perplexity --kl-divergence, is teacher-forced: at every scored position it resets the context to the reference prefix. Error accumulation is exactly zero by construction. It cannot observe a trajectory drifting and failing to recover, because no trajectory is ever generated.
The failure is not a wrong answer
I classified every failure instead of just counting them, and the shape surprised me. Below the cliff, across three different builds:
| bits/weight | no tool call at all | acts, wrong result | runs out of turns | passes |
|---|---|---|---|---|
| 2.23 | 54.2 % | 10.0 % | 0.0 % | 35.8 % |
| 1.99 | 55.8 % | 9.2 % | 2.5 % | 32.5 % |
| 1.94 | 53.8 % | 12.5 % | 3.1 % | 30.6 % |
In more than half of the cases the model produces no action whatsoever. It reasons about the task, in detail, coherently — and stops. And that rate is flat across the whole band: compressing further does not make it worse, because it is already gone.
I spent a while trying to fix that, on the theory that it was a prompting or serving artifact rather than lost capability. Four attempts, all negative: a repetition penalty (the reasoning does sometimes degenerate into counting loops, but penalising it does not help), the tool chat-template on versus off, an explicit "you must call a tool" system instruction, and raising the generation cap. None moved it. On toy one-line prompts the same build emits tool calls about 75 % of the time; on the real tasks, 46 %. The failure scales with difficulty, which is what lost capability looks like and what a formatting bug does not.
What I got wrong, which is most of this
The result above is four numbers. Getting to them took a week, and almost all of that week was me measuring in the wrong place.
I compared five builds that were all below the cliff. Different bit allocations, different sources, one copied tensor-for-tensor from a well-regarded community release. Nothing was ever statistically significant, and I concluded the benchmark lacked resolution. The benchmark was fine — it separates 100 % from 37 % without effort, and 98.8 % from 90 % at p = 0.031. There was simply nothing to measure down there. Allocation stops mattering once the model is broken, and I had spent days optimising allocation inside that region.
I read a curve into two points. When the first fidelity numbers came in I wrote that per-token flip rate explained the capability cliff. It does not, and the very first data point — 9.5 % flips, perfect score — already said so. I noticed the tension, called it a tension, and then talked myself out of it when a second point pointed the way I wanted.
And I conceded a statistical argument I should have held. Ranking three builds by size held in three of four rounds, which a sign test on rounds calls p = 0.016. It is a unit-of-analysis error: four rounds over a fixed 40-task bank are not four independent replicates, so that standard error contains decoding noise and no item-sampling noise at all — Clark's language-as-fixed-effect fallacy [4]. The per-task test, p = 0.85, is the defensible one. I had it right, was pushed, and folded.
The protocol, which is the transferable part
Five rules, each one earned by getting it wrong first:
- Objective checkers only. The checker reads the artifact or runs the produced code against an independently computed reference. Never a literal you typed — that is how a bug in your expected value becomes a "model failure".
- Use the sampling profile the model declares, not greedy. On this family, temperature 0 produces a structural 0 %: it stops emitting tool calls entirely. A harness defaulting to greedy would report every rung as broken and the cliff would vanish into a flat floor.
- Two rounds minimum. One unchanged build re-run four times scored 17, 12, 16, 15 out of 40. That five-task spread is wider than most differences you will want to report [no. 18].
- Cap generation. Without a cap, one runaway response eats the request deadline and is scored as a model failure. With a cap, a task that "failed" twice passes cleanly.
- Classify failures, do not just count them. The table above is the most informative thing in this post, and a pass rate alone would have hidden all of it.
What this does not settle
Bits per weight is not the variable I manipulated, and that is the honest limit here. Each rung changes attention, feed-forward, state-space projections, embeddings and the output head all at once; "2.93 bpw" is a scalar summary of a five-dimensional move. So I can say where this ladder breaks. I cannot yet say which part breaking causes it — that needs one family pushed down at a time with the rest held high, and it is the next thing I intend to run.
One model, one family, one benchmark. Whether the threshold generalises is untested, and the mechanism is open: the reasoning-path story is plausible, has a plausible competitor in every direction, and my attempt to measure it picked a protocol that could not see it.
References
- Dettmers, T. & Zettlemoyer, L. (2023). The case for 4-bit precision: k-bit Inference Scaling Laws. ICML 2023. arXiv:2212.09720 — the accuracy-per-bit framing that makes sub-3-bit quantization worth attempting, and the source of the convention of reporting ladders by bits per weight.
- Gu, A. & Dao, T. (2024). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. COLM 2024. arXiv:2312.00752 — the selective state-space formulation behind the hybrid architecture measured here; most of this model's layers are of this kind rather than attention.
- Frantar, E., Ashkboos, S., Hoefler, T. & Alistarh, D. (2023). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. ICLR 2023. arXiv:2210.17323 — second-order error compensation, the family of techniques that would be the principled way to buy back what the low rungs lose. Not applied here.
- Clark, H.H. (1973). The language-as-fixed-effect fallacy: A critique of language statistics in psychological research. Journal of Verbal Learning and Verbal Behavior, 12(4), 335–359. doi:10.1016/S0022-5371(73)80014-3 — why treating repeated runs over a fixed item set as independent replicates understates the standard error. The reason the p = 0.016 in this post is discarded and the p = 0.85 kept.
- McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2), 153–157. doi:10.1007/BF02295996 — the paired test behind every build-versus-build comparison here; appropriate because all rungs run the identical task set.
- Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. TACL 12, 157–173. arXiv:2307.03172 — prior evidence that an intact-but-differently-positioned context changes behaviour; the same lesson as no. 17, that a metric can be blind to a factor it holds fixed.
The harness, the checkers and the full task bank are being released so these numbers can be checked — publishing a bank also makes it contaminable, so once it is out, treat models trained after that date with suspicion on it. Every figure here comes from at least two independent rounds, produced by the analysis script rather than transcribed by hand.
Previous: benchmarking a file against itself · back to the index.