I spent months picking better context. It turns out the gaps were doing more damage than the picking.
Every system that trims a long conversation has a scorer: the part that decides which chunks survive. That is where the effort goes, mine included. So I ran the full grid — three ways of choosing, each with and without one small change to what happens after choosing. The scorer turned out to be worth almost nothing on its own. The other factor, which I had never once measured, was worth as much as retrieving perfectly.
The previous entry, no. 16, ended badly for me on purpose. I had built four generations of a context engine, each one beating the last, and then added a baseline with no ideas in it at all — keep the newest messages, drop the rest — and watched a good chunk of my cleverness evaporate. The one thing that survived was worth keeping: a well-chosen 8 % of a conversation beats handing the model the whole thing.
This post is what happened when I kept pulling on that thread. It is a better result than the one that broke, and it arrived by the same route: measuring the thing nobody measures.
The shape of the problem
Modern long-context serving does something that sounds obvious once you say it out loud. The context is cut into fixed-size blocks. Each block gets a score against the current question. You keep the best ones until a token budget is full, and you drop the rest.
Then comes the part that gets no attention at all. The blocks you kept are fed to the model at the positions they originally had. Block 4, block 31, block 88. Between them are the positions of everything you threw away — holes.
Nobody reports what those holes cost, because nobody treats layout as a variable. It is just what the plumbing happens to do — even though position is known to change whether evidence gets used at all [5]. I had never measured it either, across four generations of engine and a withdrawn upstream proposal [9].
So I made it a variable
Three ways of choosing the blocks. Lexical scoring with BM25 [1], term weighting and all [2]. Rank fusion of BM25 with dense embeddings [3] — the "better scorer". And an oracle that is handed the blocks containing the gold evidence, which is the ceiling this literature reports against.
Each of those, twice: once with the holes, once re-packed. Six conditions, one run, 200 questions from a long-conversation benchmark [4], every comparison paired question-by-question. The token budget is identical everywhere — 8 % of the full context — so no condition ever gets more to read than another.
The decision rule was fixed before I looked at any of it: a difference supported by fewer than 90 % of paired bootstrap resamples gets reported as not conclusive, never as a win. That rule cost me the headline I was expecting, twice.
Update (Sep 2026): a later entry puts a floor under all of this: running one unchanged model through the same benchmark four times moves the score by five cases, which is wider than most differences reported here. Single-run comparisons in this notebook should be read with that in mind. → no. 18
The numbers
Read the first pair first. Rank fusion recovers substantially more of the gold evidence than plain lexical scoring — 53.7 % up to 61.7 % — which is exactly what a better scorer is supposed to do. And it buys +0.23 F1, favoured in 60 % of resamples. Not conclusive. The retrieval improved and the answers did not.
Now read down the column. Re-packing lifts every single arm: +3.06 for lexical, +4.65 for fusion, +3.09 for the oracle. Holes win in 0 % of resamples in all three cases. Whatever this effect is, it is not a property of one selection rule.
Fixing the layout is worth about as much as retrieving perfectly. With positions held constant, going from a lexical scorer all the way to gold evidence is worth +3.68 F1. Re-packing alone, with no better retrieval at all, is worth +3.06 to +4.65.
The part that reframes it
Here is the finding I did not go looking for, and the one I think is actually useful to someone else.
With the holes in place, rank fusion is worth +0.23 — nothing. With both conditions re-packed, the same selection improvement is worth +1.83. The retrieval gain was real the whole time. It simply could not be collected while the layout was broken.
That offers an explanation for something I suspect is common and rarely written down: you build a scorer that demonstrably retrieves better evidence, you can prove it on recall, and the end-to-end metric barely moves. The improvement may not be missing. It may be masked.
The ceiling isn't the ceiling
There is a second consequence, and it is uncomfortable for a whole genre of results.
The oracle — perfect evidence, 99.9 % recall — scores 17.97 with holes and 21.06 re-packed. The number this field quotes as the upper bound is itself paying the positional cost. Every paper that reports "we close N % of the oracle gap" is measuring against a target that is partly an artefact of layout.
And it explains something that looked like a mistake in my own data before I had the full grid. A mediocre selection, re-packed, scores 19.21 — above the oracle's 17.97. For a while I thought I had beaten perfect retrieval, which is the kind of result that should make you go back and look for your bug. I hadn't. The oracle was just handicapped in the same way as everything else.
What this does not say
Selection still matters. With positions fixed, lexical → gold is +3.68 F1, and that is real. Streaming methods renumber positions too, but for a different reason and without isolating the factor [7]; and block granularity itself is under question from the opposite direction [8]. The claim here is narrower and, I think, more useful: a factor that is usually held fixed and almost never reported is worth about the same as one the whole field optimises — and it is far cheaper to obtain, because it needs no better retriever, only a different arrangement of what you already chose.
Three honest limits. One reader, one dataset — a small model on a long-conversation benchmark; whether the magnitude survives at scale is untested and I do not claim it does. The adversarial category scores 0.0 in every condition, so it contributes nothing and I exclude it from the per-category figures rather than let it flatter the averages. And re-packing is not free in every serving stack: renumbering positions means the surviving blocks' cache entries have to be re-materialised, and an engine that cannot do that cheaply will not get this for free. I measured the quality effect, not the systems cost.
The biggest gap is one I can name precisely and have not closed: I do not know which property of contiguity does the work. Absolute position, relative order, or simply the absence of discontinuities — the design here cannot separate them. Rotary position encoding [6] is the mechanism through which the indices reach attention, but naming a mechanism is not measuring one. The per-category breakdown hints at order (the gains concentrate in temporal and multi-hop questions, which are the ones that depend on it) but hinting is not showing.
Why this is the same lesson as last time
In no. 16 the mistake was comparing every version against my previous version, so I never learned what a program with no ideas would score. This time the mistake was subtler and had the same shape: I compared scorers against scorers, for months, while the thing sitting downstream of all of them was never in the comparison at all.
Both were fixed the same way — by adding the condition that made me look bad. The withdrawn upstream proposal [9] was the first version of that habit: I closed it in public because I had finally measured it against doing nothing and it lost. This is what turned up by continuing to look.
References
- Robertson, S.E., Walker, S., Jones, S., Hancock-Beaulieu, M.M. & Gatford, M. (1994). Okapi at TREC-3. Proceedings of the Third Text REtrieval Conference (TREC-3), 109–126. — the lexical scorer used as the baseline selection rule in every condition here.
- Spärck Jones, K. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1), 11–21. doi:10.1108/eb026526 — IDF, the term weighting inside the lexical arm.
- Cormack, G.V., Clarke, C.L.A. & Buettcher, S. (2009). Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods. SIGIR '09, 758–759. doi:10.1145/1571941.1572114 — the rank-fusion method used for the "better scorer" arm; chosen because it needs no score calibration between the lexical and dense rankers.
- Maharana, A., Lee, D.-H., Tulyakov, S., Bansal, M., Barbieri, F. & Fang, Y. (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents. ACL 2024. arXiv:2402.17753 — the benchmark and its own F1 script, used unmodified, and the source of the five question categories. Licensed CC BY-NC: used to measure, not shipped.
- Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. TACL 12, 157–173. arXiv:2307.03172 — position affects whether evidence is used at all, independent of whether it is present. The prior that made me suspect layout was worth measuring.
- Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B. & Liu, Y. (2024). RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing 568. arXiv:2104.09864 — rotary position encoding, the mechanism through which block indices reach attention at all. Cited as the candidate explanation for the effect, not as something this post demonstrates.
- Xiao, G., Tian, Y., Chen, B., Han, S. & Lewis, M. (2024). Efficient Streaming Language Models with Attention Sinks. ICLR 2024. arXiv:2309.17453 — the closest published relative: it also renumbers positions after dropping tokens, but for streaming rather than query-driven selection, and it does not isolate the layout factor from the selection factor.
- Wu, W., Pan, Z., Wang, C., et al. (2024). TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection. arXiv:2411.02886 — questions block-level selection from the opposite direction: critical tokens are not contiguous, so block granularity drags in noise. Complementary to this post, which holds selection fixed and varies layout.
- Cisneros, K. (2026). [RFC] ACE — Attention-Weighted Context Eviction for multi-turn tool results. vllm-project/vllm PR #42645, closed by the author. github.com/vllm-project/vllm/pull/42645 — the earlier proposal this line came from, withdrawn after it lost to a no-op baseline. Its closing comment is the starting point of no. 16.
All six conditions were measured in a single run against a public model and a public benchmark, paired per question; no fine-tuning is involved anywhere. Every figure quoted here is produced by the analysis script, not transcribed by hand — a check that failed twice while drafting, which is why it exists.
Previous: four versions, none measured against doing nothing · back to the index.