Bitácora · Kiko Cisneros
quantization · measurement no. 18

I spent two days benchmarking a file against itself.

I had a hypothesis, a pre-registered decision rule, a control condition, and a clean negative result. I wrote it into the lab notebook as a closed question. Then I opened the two files with a hex tool and found they were byte-for-byte identical. What that mistake uncovered was worse, and more useful: the run-to-run noise in my benchmark is larger than every difference I had been reporting from it.


Two entries ago, in no. 16, I wrote about building four versions of a context engine and never once comparing any of them against doing nothing. This is the same failure wearing a better disguise, and I walked straight into it with every safeguard I had learned since.

The setup, which was correct

The model is a 27B hybrid — most of its layers are state-space rather than attention [1] — squeezed to roughly 1.75 bits per weight [2]. At that pressure it degrades badly, and the question was where the damage lands. My hypothesis: the state-space dynamics are the bottleneck, so spending 8 bits on those two small parameter sets should buy back capability.

I did this properly. A control condition — the same quantization without the change. A decision rule fixed before looking at anything: ≥45 % means the dynamics were the bottleneck, ≤30 % means refuted. Both conditions in the same batch, because comparing against a number from another day is how you fool yourself.

The result came in at 20 % versus 20 %. Refuted, cleanly. I even had corroborating detail: the two conditions passed different cases — five in common, three unique each — which is exactly the signature of "no effect, just sampling noise". I wrote the negative result into the notebook and moved on.

The check I did not run

A day later, chasing an unrelated question, I dumped the tensor tables of both files side by side.

baseline.gguf ssm_alpha · ssm_beta → Q8_0 ×48 experiment.gguf ssm_alpha · ssm_beta → Q8_0 ×48 vs Tensors whose type differs 0 Tensor data (sha256 at 5 / 30 / 60 / 90 %, offset-aligned) identical Difference in file size 256 bytes — metadata only
The "experimental" file and its control. The baseline already had the change I thought I was introducing.

Zero tensors differed. The weights hashed identically at every offset I sampled. The two files differed by 256 bytes of metadata and nothing else.

The baseline already had the change. The default quantization recipe was already spending 8 bits on those parameters, in all 48 layers, and had been all along. My experimental condition changed nothing, so of course it measured the same. The pre-registration was impeccable. The control was correct. The artifact did not exist.

A pre-registered decision rule is worthless without a manipulation check. Before looking at any outcome, show that the experimental artifact actually differs from the control — a diff, a hash, a size. If the diff is empty, there is no experiment, and every downstream statistic is a description of noise.

Then the floor gave way

Having decided to re-run the comparison properly, I did something I had never bothered to do: I ran the same model through the same 40 cases four separate times.

Seventeen. Twelve. Sixteen. Fifteen. Same weights, same harness, same settings — the sampling temperature is 1.0, which is what the model itself declares in its metadata, so each run is a fresh draw. The spread of one model against itself is five cases wide.

Now go back and look at every comparison I had been making. Our recipe versus theirs: one case. Our old build versus the new one: four cases. The graded allocation versus the uniform one: five cases. Every difference I had been interpreting fits inside the noise band of a single model measured twice.

The paired test says the same thing without the drama. Across four models on the same 40 cases, not one comparison reaches significance: the closest rival pair comes out at p = 1.000 — eight cases won by each side [3].

So I bought more resolution the only way that does not require inventing new test cases: four independent runs of each model, turning every case from a coin flip into a rate. 160 observations per model instead of 40. The apparent edge — my build led in all four rounds, by one or two cases each time — evaporated on contact with the aggregate: 15 cases where mine does better, 13 where theirs does, 12 tied. p = 0.851. The consistent-looking lead was the shape noise takes when you only look at totals.

Update (Sep 2026): there was a reason every comparison in this post came out inconclusive: all five builds being compared sat below a capability cliff, in a region where bit allocation no longer changes anything. The benchmark was fine — measured above the cliff it separates 100 % from 37 % trivially. → no. 19

What was underneath all of it

This matters beyond one lab notebook, because there was a story built on those numbers. For weeks I had been operating on "the well-known community build beats ours — better quality in a smaller file", and had spent real effort chasing that gap. That belief rested on two percentages: 56 % versus 60 %.

Those two numbers came from different runs, on different days, on either side of a server upgrade, and at a sampling temperature of zero — which, for this particular model family, structurally prevents it from emitting tool calls at all. Re-measured properly, in one batch, with the sampling profile the model itself asks for, the gap disappears: ours and theirs are indistinguishable on quality. Their file is 13 % smaller, which is a real advantage and the one I should have been chasing all along.

Three traps in the plumbing, for anyone doing this

Chasing this turned up defects that had nothing to do with models, and every one of them produced a number that looked publishable:

What I actually changed

Two things, and neither is a clever idea.

Every comparison now carries a manipulation check: before any outcome is read, the artifact and its control are diffed, and the diff goes in the log next to the result. And no single run counts as a measurement — the noise floor has to be established by re-running one condition against itself, because until you know how wide that is, you cannot know whether anything you are reporting is real.

The uncomfortable part is that both of these are things I would have told someone else to do. What I was missing was not the principle. It was applying it to the step I considered too simple to check.

References

  1. Gu, A. & Dao, T. (2024). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. COLM 2024. arXiv:2312.00752 — the selective state-space formulation behind the hybrid architecture measured here, and the source of the alpha/beta dynamics parameters the failed experiment was aimed at.
  2. Dettmers, T. & Zettlemoyer, L. (2023). The case for 4-bit precision: k-bit Inference Scaling Laws. ICML 2023. arXiv:2212.09720 — the accuracy-per-bit framing that makes sub-2-bit quantization worth attempting at all, and the reason file size is the axis worth optimising.
  3. McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2), 153–157. doi:10.1007/BF02295996 — the paired test used for every model-versus-model comparison here; appropriate because all conditions run the identical case set.
  4. Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. Annals of Statistics 7(1), 1–26. doi:10.1214/aos/1176344552 — the resampling procedure behind the pre-registered "fewer than 90 % of resamples is not conclusive" rule used in no. 17 and here.
  5. Nosek, B.A., Ebersole, C.R., DeHaven, A.C. & Mellor, D.T. (2018). The preregistration revolution. PNAS 115(11), 2600–2606. doi:10.1073/pnas.1708274114 — pre-registration as a guard against post-hoc reasoning. This post is a note on its limit: it constrains the analysis, not the existence of the manipulation.

All figures here come from runs on a public benchmark against public model builds, produced by the analysis script rather than transcribed by hand. The identity of the two "different" files was established with gguf-dump tensor tables plus offset-aligned sha256 sampling, both reproducible.


Previous: the gaps were doing more damage than the picking · back to the index.