Share this article:

If you are a Korean cloud or AI engineer designing inference cost for the repeated work inside an unattended agent loop, this article is for you. A task the loop has succeeded at once is handled one of two ways next time: plan again from scratch, or compile the validated success record into a skill routine and replay it. This ThakiCloud paper prices that choice as a closed-form cost-quality model on self-hosted GPU. It is a break-even law that answers at which recurrence frequency and at which drift the compilation cost is amortized, and at which point amortization breaks down. In one line: it is the rule for getting the inference cost of repeated work back along the time axis.

In Plain Terms: A Recipe for a Dish That Succeeded Once

Picture a kitchen. There is a chef who looks at the ingredients and thinks up the cooking order from scratch for every single order. There is also a kitchen that has distilled its successful dishes into recipes. A recipe is a template. It has blanks for quantities and ingredients. If a value does not fit the recipe’s domain, a guard trips. When a regular order comes back, you fill in only the blanks and cook. If the regular’s taste has drifted outside the recipe, you cook from scratch again. If the kitchen is remodeled, the recipe has to be rewritten. The paper’s question is when this recipe pays for itself. If the regular comes once a day and their taste does not change much, the recipe is worth it. If orders change often or the kitchen is touched often, cooking from scratch is cheaper. In this analogy, the recipe is the agent routine, and the chef’s improvisation is the clean-slate interpretation in which the model plans from scratch every time.

The Problem: The Inference Cost Paid From Scratch on Every Repeated Task

An unattended agent loop is a system that runs 24 hours without a person standing on the critical path. Its capacity comes from inference capacity the operator owns directly, namely self-hosted GPU. In a self-hosted environment, tokens are time. Every decoded token consumes the machine’s wall-clock time, so one second not spent decoding one task is one second available to the next.

When tasks of the same family arrive repeatedly with new parameter values, the loop plans from scratch and interprets every time. That amounts to paying the full per-task inference cost on every repetition. These two execution regimes carry a long pedigree. Interpretation is the default mode of operation for modern LLM agents, as set by ReAct. Compilation is the growing skill library of Voyager. Voyager established that a skill library works. It did not ask when it pays for itself. There is no cost model, no drift accounting, and no break-even condition against replanning.

Skill-harness research in 2026 measures the gate for skills. It includes routing, retrieval, composition, maintenance, and the evolution of the harness itself. It measures the gate but does not price the behavior itself. The axis that prices replay and replanning in tokens and quality, as a function of how often a family recurs and how fast it drifts, is empty. The ThakiCloud platform has already priced the neighboring levers in the same self-hosted setup. There is the safe-skip frontier that caches the routing decision itself, Cached Router. A cached decision ages on two independent clocks: registry churn and paraphrase drift. There is the token economics of the fixed prefix that an agent loop resends every turn, Manifest Order. There is the price of one percentage point of tool-call quality under a free deterministic validator, Validator’s Dividend. These levers change how much the loop spends. They do not change what the loop does. Every instance of a recurring family is still planned from scratch. The last lever this paper prices is exactly that: compiling successful traces into reusable routines, and the rule for deciding which task family to compile.

The Core Contribution: A Closed-Form Model That Prices Replay and Replanning

The paper’s first contribution is a closed-form cost-quality model of compiled and interpreted execution on self-hosted GPU serving.

Interpretation cost is written as the sum of a shared resend block and a planning block. The shared resend block consists of the system prompt, tool schemas, and the skill manifest. Both execution regimes resend it every turn. The planning block carries the inference inflation of deliberative planning. Replay stacks a much smaller planning block on top of the same shared block. A known scaffold needs no deliberation, so the inflation factor is much smaller than in interpretation. Compilation is a one-time cost. It is the cost of distilling a parameterized template, per-slot guards, and validator gates from a single validated success trace.

Three things attach to it. The first clock is slot drift. When an instance value steps outside the surface form of the template, a guard trips. With m slots and per-slot mismatch probability ρ, the per-run guard-trip probability is 1-(1-ρ)^m. When a guard trips, the fallback fires. A fraction α of the replay prefix already in use, together with the full interpretation cost, leaves at once. The second clock is churn. When the scaffold changes in front of the routine, the routine is invalidated and recompiled with per-run probability p_c = μ/λ. The third is silent failure. A replay that passed the guard can still be semantically wrong: the schema passes and only the meaning is off. A free deterministic validator, with schema-level and AST-level checkers, catches the fraction δ of silent failures. Caught failures are repaired locally at cost η·C_I. The missed fraction (1-δ)·p_s^raw remains as quality loss with no token cost.

Compiled execution path of an unattended agent loop Conceptual diagram of the compiled execution path of an unattended agent loop. A validated success trace is compiled into a parameterized routine and replayed on every recurrence. When a slot guard detects drift, it falls back to from-scratch interpretation. When scaffold churn invalidates the routine, it is recompiled. (Analytical model, not measurements.)

The expected per-run cost is written as E[C_C] = C_I - S(p_m) + p_c·C_comp. S(p_m) is the per-run savings: net savings after subtracting repair cost and the replay prefix discarded on mismatch, linear and monotonically decreasing in the drift probability p_m.

Four results build on this savings. First, the break-even of compilation. Compilation is strictly worth it when S(p_m) > 0 and the family’s recurrence rate λ is higher than λ* = μ·C_comp / S(p_m). In the (churn, recurrence) plane, the frontier is a straight line with slope C_comp / S(p_m). As drift eats into the savings, the line moves up. The reverse reading of the same rule is the churn budget μ* = λ·S(p_m) / C_comp. With recurrence fixed, it is the maximum rate at which the scaffold can change per day and compilation still pays for itself, read as the operating budget for registry maintenance.

Second, the drift ceiling of the SLO. For the SLO to be met, p_m must be at or below p_m. But in the category where the SLO is at or below interpretation quality q_I, p_m exceeds 1. The ceiling remains inactive.

Third, the economics-first corollary. The fallback pins quality at or above the interpretation level at every drift level. So when the SLO is below the quality interpretation already delivers, tokens set the deployment boundary. Quality falls smoothly toward the interpretation level, linear in p_m. Expected cost rises much more steeply. Under realistic parameters, the cost boundary arrives first. For quality to threaten the SLO, drift would have to be one or two orders of magnitude larger.

Expected per-execution token cost: compiled replay versus clean-slate interpretation Schematic of expected per-run token cost. When drift is low, compiled replay is cheaper than interpretation. As drift eats into the savings, the replay cost crosses above interpretation. The crossover is S(p_m) = 0. To its left, the signal that stops compilation is tokens. (Analytical model, not measurements.)

Fourth, the minimum amortization window W* under a finite horizon. The first compilation cost is amortized only after W* = C_comp / (S(p_m) - p_c·C_comp) runs have accumulated. As p_c approaches the break-even churn p_c, W diverges to infinity. A routine near break-even fails to amortize inside any finite window.

This rule composes with the other cost levers without re-derivation. When a quality-neutral shared cost lever subtracts the same amount κ from both regimes’ costs, the savings S, the break-even λ, and the window W all stay unchanged. κ = φ·c_0 from the prefix cache hit rate φ and κ = s·C_route from the routing decision skip rate s are exactly that case. Only absolute prices change. The model tier and quantization levers scale the planning block by (1-ψ), moving S to (1-ψ)·S, λ* to λ/(1-ψ), and W to W*/(1-ψ). The shape of the frontier stays the same. The cheaper the tier, the lower the relative appeal of compilation. The validator composes by a different path. Strengthening the gate raises the caught fraction δ. The repair cost term grows. The silent floor (1-δ)·p_s^raw drops. Since the ceiling is inactive in the economics-first category, what sets the gate’s level is the silent floor risk.

The 14B Example on H200: Families That Arrive About Once a Day Are Compilation Targets

This is an interpretive paper. The numbers in this section are not measurements but an example instantiation that fixes the scale of the frontier. The setup serves a 14B-class dense model on a single H200 (141 GB HBM3e). Decode per loop is around 400 tokens/s for 14B FP8 with paged KV. The shared resend block is written as 2.0k tokens, interpretation as 12.4k, and replay as 2.9k tokens. The gross savings Δ_0 is set to 9.5k. There are 4 slots per family, and the compilation cost C_comp is 14k tokens, about 1.1 times interpretation. The silent failure probability is 0.03, the validator detection fraction δ is 0.85, and the repair cost coefficient η is 0.25. The churn rate is 0.5 per day, the family arrival rate is 2 per day, the interpretation first-shot success rate is 0.88, the SLO is 0.85, and the operating window is 30 days.

Here are the medium-drift numbers. With per-slot mismatch probability ρ at 0.06, each of the 4 slots deviates from the template with roughly a 6% chance per instance. The guard-trip probability p_m is about 0.22, and the per-run savings S is 6.7k tokens. The break-even recurrence λ* is 1.04 per day at a churn rate of 0.5 per day. A repeated task that arrives about once a day is a compilation target. The net per-run savings S - p_c·C_comp leaves about 3.2k tokens, 8 seconds of decode. The 30-day window savings rate is 24% against an interpretation budget of 744k tokens. At ρ = 0.01, the savings rate rises to 42%.

Drift pays back its own cost monotonically. As ρ rises from 0.01 to 0.15, λ* grows from 0.78 to 1.98, and the minimum amortization window W* grows from 2.6 to 443. So in this setup, with 2 recurrences per day and a 30-day window, the constraint hit first is the amortization window. It is hit earlier than the steady-state break-even. The cutoff of the window constraint is around ρ ≈ 0.14, well inside the point where cost degrades, ρ^S ≈ 0.30.

Break-even recurrence of routine compilation versus per-slot drift Schematic of the break-even recurrence. Compilation is worth it only when the arrival rate of the task family is higher than λ. λ* rises monotonically as per-slot drift and the scaffold churn rate go up. (Analytical model, not measurements.)*

The quality side is inactive from start to finish. Since the SLO of 0.85 is below interpretation quality of 0.88, the ceiling p_m* is 1.26, above 1. q_C holds at or above the interpretation level of 0.88 at every drift. At ρ = 0.06, the first-shot success rate is 0.97, so quality actually goes up while spending 24% less. In this category, what sets the boundary of compilation is tokens, not quality. Because the fallback option lets quality fall smoothly toward the interpretation level, quality breaking first does not happen until drift is far short of the cost degradation point.

Under the premise that tokens are time, each compiled run at ρ ≈ 0.06 gives back about 8 seconds of decode. If the fleet execution share φ_c belongs to the families that pass the rule, the ρ 0.03 to 0.10 band, the total token budget shrinks by that share times 13 to 25%. At φ_c = 0.3, that is a 4 to 7% savings of the total token budget with no new hardware. There are side benefits that are not a budget line. The compiled path makes per-task cost almost deterministic. On the replay-dominated path, only fallback and recompilation vary. It suppresses the run-to-run variance in which the same task consumes an order of magnitude or more in tokens between executions, and capacity planning returns to an average-rate problem.

So What Changes: Company, Society, Science

What remains for the company is a compilation policy for the unattended loop. The overnight autonomous loop of the ThakiCloud AI platform now carries a quantitative break-even rule that judges which repeated task to compile into a skill routine. The token and time spend of recurring tasks is amortized along the time axis. It can be measured directly with the Metis token counter. Every input of the frontier, recurrence λ, churn μ, per-slot drift ρ, and operating window τ, is read from the logs the production harness already uses. The measurement protocol stands these inputs on BFCL-style parameterized function-call families and deterministic success checks.

What remains for society is a structural change in marginal cost. The structure of paying the inference cost from scratch for every recurring task becomes the structure of amortizing it with a once-validated routine. It greatly lowers the marginal cost of fully autonomous task automation. The door opens even to small organizations that could not adopt on-station agent automation with a single self-hosted GPU.

What remains for science is the first cost-quality amortization frontier between the two execution regimes, compiled and interpreted. It extends the two-clock structure of routing decision caching, Cached Router, one step further into trace-level behavioral memory. It lifts Voyager’s qualitative skill library into a quantitative break-even law over the recurrence-frequency and drift axes. If Voyager established that a skill library works, this paper gives, as a law, when it pays for itself.

What Cannot Be Trusted: An Interpretive Paper With Four Limitations

First, this is an interpretive paper. The frontier is a closed-form rule. The numbers in the H200 section are an example instantiation of that model, not measurements. The experimental design under which each quantity is measured is spelled out in the measurement protocol section: ten or more parameterized function-call families, a drift sweep over ρ 0.01 to 0.25, churn schedules of 0.2/0.5/1.0 per day, and interleaved execution of the two arms on the same serving instance. The success criteria against a 30-day window are also fixed. Savings must be measured within 90 to 110%, break-even recurrence within a factor of 1.3, and quality within 0.02.

Second, the channel independence assumption. Slot drift and scaffold churn are modeled as conditionally independent. In a self-evolving registry, churn can key on the observed task representation. If the two channels couple, this frontier should be read as a lower bound on actual cost increase. The rule must be applied with a shrunken window.

Third, the structural idealization of the cost model. Guard detection at α = 1, the late-detection case, is a conservative choice. A finer guard is a trade that exchanges guard tokens for early repair. C_comp is held as a constant per family. In practice it grows in proportion to scaffold length, slot count, and trace novelty. A routine distilled from a single trace can overfit to that trace’s surface form and raise the silent failure rate. For families whose compilation cost greatly exceeds interpretation, they stay above λ* under any realistic recurrence. The rule bounds this automatically. Churn detection is assumed immediate. If it is delayed by L runs, the quality risk is bounded by the silent floor and the cost risk is bounded by L·p_m·α·C_R. The validator is the shield.

Fourth, the scope of the accounting. Only the first-attempt token cost is priced. Downstream incident costs of silent failures are excluded. That risk is controlled by the validator level δ. It assumes a single self-hosted instance. A cloud API pricing regime changes absolute prices but not the rule. Multiple routine families and routine sharing across families are open extensions.

A repeated task that succeeded once is an asset. Planning from scratch is a cost every time. The law of this paper prices the gap between them in units of tokens. It is the rule for deciding which family to compile, when to recompile, and which family to leave to interpretation.

Paper and data: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-10-06-compiled-routine-agent-loops-amortization

Share this article:

Tags: amortized-inference-cost, cost-quality-frontier, fallback-to-interpretation, h200-serving, parameter-substitution, procedural-memory, routine-compilation, skill-ecosystem, trace-replay, unattended-agent-loops

Categories:

Updated: