The Prefix Breaks at the First Edit: How the Skill Harness’s Overnight Self-Rewrites Steal Cache Reuse and Serving Cost
Read this if you run unattended agents, or if you own the serving cost of the H200 instance that carries them. The fixed header an agent resends every turn is fixed only within a version. When the harness rewrites itself every night, the cache is invalidated from the first edited entry, and the fraction that survives a single rewrite is less than 2 in 100.
A visual metaphor for the article’s key idea.
In Plain Terms
Think of a print shop. The agent requests the same manuscript every turn. The print shop works by a rule: anything that matches a previous edition page by page, starting from page 1, is not reprinted. If even one page changes, everything after it is treated as new manuscript and reprinted. That is exactly what the serving engine’s prefix cache does.
The manuscript is the skill manifest the agent sends every turn: 200 entries, 6,000 tokens. The harness is a publisher that reprints part of the pages every night. It rewrites description text, changes chapter order, and adjusts control parameters. The point where reprinting begins is called the first break, and it is set by the first page that changed.
When the publisher makes a lot of changes every night, the first break lands around 4 pages on average under the reference configuration the paper uses to price things. Everything from there to page 200 is reprinted, and the surviving fraction is less than 2 in 100. Every lever this article hands you is something the print shop itself decides: how to bind the manuscript, how often to reprint, what to ship, and what thickness to hold. No new printing machine, no model swap.
Infographic generated by NotebookLM from the sources.
A Manifest That Changes Every Night
The fixed header is three pieces: the system prompt, the tool schema, and the skill manifest. The serving instance runs a prefill on the first turn, the process of computing from the start, and after that it uses the prefix cache. A cached block is reused only if it matches the served header starting from token 0. This rule is called coherent semantics.
The manifest lists the 200 entries the routing layer can use, in order. The subject of this paper is the movement of the manifest itself. An overnight self-evolution loop rewrites description text, reorders entries, and adjusts control parameters.
The fleet’s recorded history backs this up. The registry grew from 1,600 to 2,029 over 3 weeks.
The recorded edit sweeps show only 11 description files rewritten. Identifiers and fixed parameters were left untouched.
Two earlier studies priced the static side. Ordering within a version is worth up to 12~89% of input cost, and index staleness has its own decay law and an optimal re-embedding cadence. What nightly rewrites do to the serving instance’s cache was not priced. The layer this paper prices is exactly that.
The reference configuration used for pricing is an assumption, not a measurement. Each apply, the event that pushes a rewrite onto the instance, rewrites 21~31% of descriptions and moves 14~17 entries in rank. That is a stress point set several times above the recorded sweep rate.
The First-Break Law
The law comes out in closed form. The expected length of the prefix shared by two consecutive versions is fixed by the first break. If the probability of an entry being rewritten is β and the number of entries is N, the expected number of entries changed per apply is the product βN. Once βN is bigger than a handful of fingers, the surviving prefix stops growing no matter how large the manifest gets. It becomes a fixed token budget at the front of the manifest.
The survival rate falls steeply as βN grows. Under the reference configuration, 50 entries change per apply and only 1.6% survives. In plain terms, that is like 196 of 200 pages being reprinted every night.
As the expected number of entries changed per apply (βN) grows, the fraction of the cached manifest that survives one apply shrinks. At βN=1 about 63%, at βN=10 about 10%, and at the reference point (βN=50) 96 of the 6,000 tokens, 1.6%, survive. Graph from an interpretive model, not measurements.
Two properties worth noting here. First, the loss does not shrink just because the cache is large or many loops are shared. It is a different channel from the within-night drift priced earlier. Second, in the saturation region the loss is independent of manifest size. Whether it is 2,000 tokens or 48,000, almost the whole thing is reprinted every apply, and only the bill grows with the size.
Rebinding to Confine Reprinting to the Body
The first lever is how to bind the manuscript. The current layout weaves each entry’s name and description together in order of access frequency. So when one description changes, the break lands at the head of that entry. The paper proposes a layout where changes can never reach the front matter. It is called the prefix-stable layout (PSL).
It has two layers. First, all stable fields, identifiers, names, and fixed parameters, are concatenated in identifier order and placed at the front. This is the skeleton, 2,400 of the 6,000 tokens. Skeleton entries barely change, so they never break. Behind them sits a detail block collecting the description text. The swap rule sorts the detail block starting with the entries least likely to be rewritten.
That confines invalidation inside the detail block. Under the reference configuration, the invalidation span per apply drops from 5,904 to 3,546, a 1.7x reduction.
The effect grows with the skeleton’s share. Move long parameter blocks into the skeleton and the reduction becomes 2.5x.
Invalidation positions for a 6,000-token manifest. The stable skeleton is 2,400 tokens, the detail block is 3,600. The current layout weaves stable and volatile fields per entry, so the first break invalidates almost the entire manifest. The prefix-stable layout draws the byte-fixed skeleton first and traps invalidation in the detail block. Graph from an interpretive model, not measurements.
The Price of One Apply and the Freeze Cadence
An apply is a visible event. The first turn right after an apply prefills 5,904 tokens as a cache miss, and that turn’s input cost becomes 2.3x the warm baseline.
Every loop whose turns fall inside the re-edit window suffers the same miss. The paper calls this the collision factor. The faster the loop, the larger the factor. In plain terms, the faster the loop, the more crowds gather at the re-edit counter and the first day’s bill and latency spike.
| Loop interval | Collision factor | Time to first token (baseline 1.44s) | Bill per apply |
|---|---|---|---|
| 900s | 1.02 | 2.46s (1.7x) | $0.0033 |
| 60s | 1.37 | 2.97s (2.1x) | $0.0043 |
| 10s | 3.21 | 5.69s (3.9x) | $0.0100 |
Collision factor, time to first token, and bill per apply by loop interval. The warm baseline latency is 1.44s. All model predictions, not measurements.
How often to reprint is decided by the freeze cadence. If the cadence is k, k nights of edits are bundled into a single apply. The per-night cost is the sum of two terms. The amortized re-edit cost shrinks as k grows, and the manifest staleness cost grows as k grows. Set the cadence to 7 and the nightly tax comes down to roughly half of what it is at cadence 1.
| Freeze cadence k | Current layout (tokens/night) | Prefix-stable layout (tokens/night) | Prefix-stable layout + gate (tokens/night) |
|---|---|---|---|
| 1 | 5,904 | 3,546 | 1,773 |
| 2 | 2,978 | 1,788 | 894 |
| 4 | 1,495 | 898 | 449 |
| 7 | 856 | 514 | 257 |
| 14 | 428 | 257 | 129 |
| 28 | 214 | 129 | 64 |
Amortized invalidation span S(k)/k by freeze cadence, in tokens/night, under a 900s loop interval. All model predictions, not measurements.
Per-night decay law: the graph of the amortized invalidation span S(k)/k by freeze cadence k. The invalidation span is capped by the layout, and the per-night cost decays hyperbolically in k. The prefix-stable layout pulls the floor down multiplicatively, and the quality gate halves the event rate. Meaning: drawn from an interpretive model. It does not contain measured data.*
Minimizing the sum gives the square-root law. The optimal cadence k* is set by how expensive losing one routing point is in follow-up rework. Under the reference configuration, k* falls somewhere between about 1 and 12 depending on the rework cost. Read loosely: if a routing point is expensive, reprint daily; if it is cheap, bundle it into one weekly apply.
The cost landscape is flat. Within 2x of k*, no cadence costs more than 25% extra. So the operator does not need to chase the exact value; measure the rework cost and pick a value that fits the calendar.
Re-freezing the manifest and re-embedding the search index are the same job. Since both artifacts have to catch up whenever the registry moves, it is better to apply them together. Then you open only one collision window.
There is a hardware asymmetry too. During the re-edit window, the old and new manifest versions coexist in the cache, and on smaller-capacity machines the effective retention time shortens. On 32B/H100 each window cuts 12.9% of the retention time. The event cuts retention the most exactly where the static levers are weakest.
The Gate and the Pin
The quality gate composes with everything. A divergence gate and an adoption gate judge whether the night’s rewrite actually improved routing. A rejected night produces no apply: no version change, no invalidation, no event. Under the current layout, churn that yields no gain is just a tax, and the gate turns it into nothing. If the adoption rate is 0.5, the nightly serving tax is halved.
The pool pin addresses the second, permanent channel. If the in-scope pool follows the registry, the 27% growth becomes a 1,585-token increase in the manifest, and +4.5% per-turn input is locked in. It is a ratchet that does not release over time. With a pool pin in place, a new skill can only enter by pushing out an existing one. Then the ratchet stops turning.
Stack the levers and the churn channel drops from 0.07% to 0.01% of the per-night instance cost. The ratchet goes to zero. These are all deployment-time numbers. No engine change, no training, no change to routing outcomes.
What You Gain by Measuring This Channel Separately
What changes for the platform that carries this loop is visibility into the serving bill. The harness’s own nightly self-edits have been quietly growing it. Now you can measure the invalidation rate per apply, the optimal freeze cadence, and even a layout that traps invalidation in the detail block. They are levers the operator sets directly, without changing routing outcomes.
The significance goes beyond the platform. Self-improving agent infrastructure running 24/7 commonly rewrites itself every night, but the hidden tax in that serving layer was invisible. This paper makes that cost predictable, alongside energy. You can keep the quality gain and bind the price you pay to a stable-prefix price only.
Scientifically, it isolates prefix instability as an independent serving-cost channel, distinct from both index staleness and the within-version ordering effect. It is the first closed-form link between self-evolution and cache-reuse dynamics. The edit-freeze cadence becomes a new variable in the routing/serving objective.
Infographic generated by NotebookLM from the sources.
What Not to Trust
Every number in this article is a model prediction under the 10 stated assumptions, and none of them is a measurement. The paper specifies a re-run protocol on a dedicated H200 to verify these predictions. It compares a side that freezes the manifest against one that changes it nightly, using pre-registered falsification conditions. This protocol is still at the specification stage.
The first is the churn rate itself. The reference configuration, 21~31% rewrites and 14~17 departures, is an assumption several times above the recorded sweeps. The protocol re-pins to the recorded sweep rate before trusting re-run results. Even so, the conclusion does not change. Even at the recorded rate ceiling, the survival rate is on the order of 1 in 10.
The second is the independence assumption. The first-break law assumes entries are rewritten independently. If the curator rewrites a related skill family all at once, the first break comes earlier than the law predicts. Skeleton protection works only insofar as that concentrated rewrite does not touch the identifiers. The recorded sweeps did not touch them. But that is only one sweep we have seen.
The third is the quality assumption. Because the freeze cadence inherits the staleness gradient, a model parameter, k* is conditional on the rework cost the operator measures. The prefix-stable layout re-renders the manifest in identifier order, and position-sensitivity exposure is largest when the manifest is 48,000 tokens. Savings should only be booked after subtracting the measured quality loss.
The fourth is the engine semantics. Everything holds only inside strict coherent prefix reuse. An engine that patches fields in place or follows cache directives has a different cost structure, so the first-break law no longer applies. The savings trade off against the price of pre-assumed value risk.
All three figures are graphs drawn from an interpretive model. They are not measured curves.
The paper’s detail page is here: The Unstable Prefix: Measuring KV-Cache Reuse Decay and Serving-Cost Inflation Under Overnight Self-Evolving Skill Harnesses on Self-Hosted H200