Share this article:

Read this if you run, or plan to run, an operation where an oncall agent handles incidents every night, rewrites its own lessons from failures, and the version that looks improved by the next morning gets promoted to the cluster. This paper measures the question that operation inevitably raises: whether the improvement the agent itself claims is real, and exactly how far it extends. ThakiCloud’s paper “The Night Oncall: Measuring Overnight Self-Evolution of K8s Incident-Remediation Skills Against Frozen Fault-Injected Holdouts” measures the nightly postmortem-driven self-evolution loop of a Kubernetes oncall agent over 4 nights. A deterministic oncall agent handles 42 seed-deterministic fault-injected incidents per night, and improvement is scored not by routing or diagnosis recall but by ground-truth incident outcomes, namely safe remediation with zero collateral damage.

The Problem: What Improved Overnight Is Real, and How Far Does It Reach

Cloud reliability is ultimately priced in human time. Every page interrupts an operator, and the longer a human stays in the loop, the more each incident costs. In that context, LLM agents with cluster access are increasingly proposed as autonomous L1 oncall responders that diagnose, remediate, or escalate incidents. Our study of a static K8s oncall agent (2026-09-09) showed that the binding constraint for a static agent is not token cost but the collateral damage cap and capability. That front line tells you how far a version V0 with frozen runbook knowledge can be deployed. But production operation is not static.

Agents are increasingly left unattended overnight and entrusted with fixing their own skills in postmortems. Self-evolving agent harness research in 2026 scores the improved harness on the same benchmark that guided its edits, which is exactly the setting Goodhart’s law warns about. We formalized this risk as Goodhart shift (2026-08-29): the gap between the improvement a self-evolving agent gains on visible cases and the improvement it gains on a sealed holdout it has never observed, widening every night.

The oncall deployment question then splits into two. Does the nightly postmortem-driven self-evolution loop actually raise the safe remediation rate on real incident outcomes. And if it does, does that improvement survive faults the loop has never seen, or is it merely a Goodhart artifact of the very incidents it edited? The first design decision of this paper is the scoring criterion. An incident is “resolved” only when the command set the agent executed exactly matches the world’s gold command set, and “safe” only when it touched nothing in the collateral damage set. Routing recall and diagnosis recall are not the score.

Experimental Design: 42 Per Night, 30 of Which the Loop Structurally Cannot See

The incident world is a seed-deterministic multi-namespace Kubernetes fault-injection simulator. Each world, keyed by a (fault class, seed) pair, carries the injected fault, the minimal gold command set that resolves it, and the collateral command set that causes damage when executed. Ground truth is structural: an incident is resolved only when the executed command set equals the gold set, and safe only when not a single collateral command was executed. There is no live cluster and no learned oracle, and every outcome is exactly decidable.

There are five unique fault classes: crash-loop backoff, OOM kill, quota starvation, stuck leader lease, and node drain, the same taxonomy used in the static front-line study. There are four unexposed classes: expired TLS certificate, label/selector mismatch, stuck horizontal pod autoscaler, and node disk pressure. The frozen splits partition the 42 worlds of every night along two axes: whether the fault class has ever been exposed to the loop, and whether the seed has.

The training split is the 15 incidents from the 5 unique classes, seeds 0-2. It is visible and forms the postmortem and the gate. The sealed holdout is the 15 incidents from the same classes with unseen seeds 3-5, testing seed transfer. The sealed kind is the 12 incidents from the 4 unexposed classes, seeds 0-2, deciding whether the loop generalizes across classes. Over the 4 nights, 168 incidents are scored, of which 108 sit in the sealed splits. Every configuration reproduces from the (class, seed) pair. The evolution loop is structurally blind to the sealed splits: an external scorer evaluates them daily, and sealed outcomes reach neither the postmortem nor the gate.

The reported arm is effectively a deterministic skill interpreter. The agent carries a list of lessons, each a (label, fix template) pair. On an incident it performs a fixed 3-4 read-only observations (get pods, get events, get nodes, describe), picks the first lesson whose label matches the observed symptom, executes its fix template, and stops. If no lesson matches, it escalates without acting. LLM calls are zero, so the measured average cost per incident is structurally $0.00. The incumbent V0 is the static runbook knowledge measured in 2026-09-09, with 4 lessons: crash-loop is rollout undo, OOM is delete pod, quota starvation is delete pod, node drain is drain {node} --force, and there is no stuck leader lease entry.

A night is structured like this. In each night k = 0, 1, 2, 3, the active version first remediates all 42 worlds across the three splits. The postmortem runs only on training failures, under a mechanical template: a failing lesson is replaced in place with a hindsight-correct class fix, and a failing class with no lesson gets one added. The result is the candidate version. The promotion gate is deferred and train-only. The candidate is activated at night k+1, and if that night’s training safe rate is lower than the previous night’s, it is rolled back. The gate’s threshold set sits entirely outside the loop’s writable surface, and the sealed splits provide no information to any edit or decision.

Nightly self-evolution loop with deferred train-only promotion gate This is the structure of a single nightly cycle. The active version is scored on all 42 fault-injected incidents, the mechanical postmortem edits only the lessons for training failures, and the deferred train-only gate approves the candidate version for the next night as long as training does not regress. The sealed holdout and kind split are scored daily by the external scorer but never feed back into any edit or decision. (Measured. Measured in a CPU-only container.)*

The protocol also fixes a metered LLM arm that is not reported in this paper. The same interface is provided by a Claude Haiku 4.5 (claude-haiku-4-5-20251001) tool loop, observations are capped at 2 per incident, and all API calls are metered at $0.80/$4.00 per million input/output tokens. Postmortem lessons are free text, capped at 5 per night. The LLM arm’s cost-quality trajectory is the subject of follow-up research.

Results: A +0.60 in One Night, and a Ceiling Just Above the Class Boundary

On night 0, the incumbent V0 is safe on exactly two of the 5 unique classes, crash-loop and node drain. The OOM and quota lessons (delete pod) are harmless but ineffective: the agent acts, no damage occurs, and the fault recurs. Incidents with no stuck leader lease lesson are escalated without gaining any information. The night 0 safe rate is 0.40 on training and holdout, and 0.00 on kind, where every incident is escalated.

Night 1 is the only editing night of the campaign. The 9 training failures of night 0, 3 each for OOM, quota starvation, and stuck leader, trigger 9 postmortem actions. Of these, 3 are effective edits and 6 are no-ops, because the mechanical template applies at most once per lesson per night. The OOM lesson is replaced in place, with delete pod becoming set resources carrying a 1Gi limit and a 512Mi request. The quota starvation lesson is replaced, with delete pod becoming delete job against the batch job nightly-batch that consumes the quota. The stuck leader lease lesson is added, a delete lease against the scheduler’s lease kube-system/leader-scheduler. The two base lessons (crash-loop, node drain) survive untouched. V1 is V0 plus 3 class-level lessons: two replacements of ineffective fixes, and one addition that V0 did not have.

Under V1, safe remediation jumps to 1.00 on both training and holdout. Both splits gain +0.60, and kind stays at 0.00. Diagnosis accuracy moves from 0.80 to 1.00 and the escalation rate from 0.20 to 0.00, in step with training.

Measured safe-remediation rate by night and split (deterministic arm) The safe remediation rate over the full 4-night campaign for the visible training split, the sealed same-class holdout, and the sealed unseen-class kind split. Training and holdout agree every night, kind stays flat at 0.00, and Night 1 is the only editing night. (Measured. Measured in a CPU-only container.)

What happens on the sealed holdout separates transfer from memorization. The 9 sealed holdout incidents that failed on night 0, the unseen-seed counterparts of OOM, quota starvation, and stuck leader, all flip to safe on night 1. 9 case flips, all silent, all in the improving direction. No edit of the loop targeted any of them; the same three class-level lessons simply applied to the unseen seeds and flipped them. After that, flips are zero. The holdout improvement is transferred, not memorized at the incident level.

The four unexposed classes never flip. All 12 kind incidents are escalated every night, diagnosis accuracy holds at 0.00, and the escalation rate holds at 1.00. The maximum train-kind divergence is 0.60, exactly the entire training gain of the campaign. The kind transfer ratio is 0.0. The verdict we record for the campaign is a memorization ceiling: the loop generalizes across seeds within the observed fault classes and draws a hard line at the class boundary it has never seen.

Nights 2 and 3 are the fixed point. There are no training failures, no postmortem edits, no candidates, and V1 stays active throughout. In this campaign, the loop’s compounding ends after one night. The promotion gate fires once: V1 is approved on night 1 (training safe rate 1.00 ≥ 0.40) and is never rolled back. There is exactly one proposed version in total. Damage is 0.00 across all 168 scored incidents, and the average cost per incident stays $0.00 on every split.

This contrast is the point. If the loop had been scored on training alone, night 1 would have read as a clean +0.60 success with no visible regression, and across every split the loop can see, it still would be. The sealed splits replace the implicit generalization claim with a measured ratio that no self-report could produce: 0.0 on kind.

Nightly divergence of training safe rate from sealed splits While train-holdout divergence stays at 0.00 every night, train-kind divergence opens to 0.60 on Night 1 and holds that level, equal to the entire training gain of the campaign. The measured gap is not benchmark overfitting but the class boundary. (Measured. Measured in a CPU-only container.)

Implications: Where the Ceiling Is, What the Gate Cannot See, and What Remains for the Company, Society, and Science

The ceiling is not benchmark overfitting. Classical Goodhart overfitting shows up as a positive train-holdout divergence: the loop gaming the specific incidents it can see. That does not happen here. Train-holdout divergence is 0.00 every night, because the editable surface is class-level. Each postmortem edit rewrites a lesson indexed by fault class, and such a rewrite either transfers to all seeds of that class or not at all. The measured divergence is instead on the kind split, which isolates a different failure: not overfitting to incidents, but failing to reach the class boundary. The loop cannot create lessons for faults it has never seen, so improvement is capped exactly by the sum of the observed classes. The two sealed splits are complementary measurement instruments: the same-class holdout guarantees that the improvement is not incident-level memorization, and the kind split guarantees that it is not class-level generalization either. The self-reporting loop sees neither; the sealed scorer sees both.

The train-only gate is a deliberately minimal Goodhart defense: it compares the training safe rate night to night and rolls back regressions. It is structurally blind to the sealed splits. Because the threshold set is designed to sit outside the loop’s writable surface, it is exactly invisible to the regressions that Goodhart shift defines as silent case flips. In this campaign, that blind spot is harmless, because the mechanical postmortem only writes hindsight-correct class fixes, so all 9 silent flips are in the improving direction. The design lesson is a division of labor: the gate’s non-regression test protects the loop from churn, and silent sealed-split regressions must be caught by the external sealed scoring itself. We install the frozen holdout and kind, the two that are scored every day and never fed back, as the promotion gate before evolved skills meet a live cluster. The moment the editable surface widens to free-text LLM lessons or control parameters, the silent-regression special case becomes live, and deterministic guardrails outside the self-editing surface become the minimum safe configuration.

For ThakiCloud, what this measurement leaves behind is concrete. It answers by measurement whether unattended nightly evolution compounds into real improvements of incident outcomes on RCA and remediation skills, and it installs a frozen incident holdout and a kind holdout as the promotion gate before evolved skills meet a live cluster. It turns the static oncall agent measured in 2026-09-09 into a continuously improving agent with an explicit cost-per-incident budget. V0 was the incumbent of the static front line, so it was safe on only two of the 5 unique classes. The overnight loop moves the front line along the capability axis: on night 1, all 5 unique classes become safe at zero incremental cost, and the measured ratio pins down where the front line binds next: the unexposed classes. There, the agent’s safe action is escalation, and what can save the remaining incident mass is a generalization-capable arm or a human. The loop is exactly cheapest on the classes that static analysis calls runbook-grade, and the measured ceiling signals that the remaining incident mass belongs to the LLM arm.

Beyond the company, for teams running safety-critical autonomous cloud operations, this study shows a way to verify that self-improvement from failure proceeds without silently degrading. The train-vs-holdout Goodhart discipline applied to K8s remediation reduces both the human on-call burden and the compute waste of incident loops that fail repeatedly. Scientifically, to the best of our knowledge, this is the first measurement of unattended overnight skill evolution scored on ground-truth task outcomes, namely binary remediation success on fault-injected K8s incidents, rather than routing or diagnosis recall. It characterizes train-vs-kind-holdout divergence and the compounding cost-per-incident trajectory of self-editing ops skills in a reproduction-grade incident world where every outcome is exactly decidable from the (class, seed) pair.

Limitations: A Ceiling of This Loop, and Only One Corner of the Front Line

The horizon is short and the class coverage is narrow: 4 nights, 5 unique classes and 4 unexposed classes, seeds 0-5. The fixed-point claim holds only within this horizon. The kind classes were our choice, and some, like expired TLS certificate, are runbook-adjacent, so a richer loop could transfer even to them. The measured ceiling is this loop’s ceiling.

There is a teacher signal in the postmortem. Because the mechanical template rewrites failing lessons with hindsight-correct class fixes, this paper measures the teacher-signal corner. Purely self-generated lessons, the free-text postmortems of the LLM arm, might generalize further, or silently degrade the sealed splits in ways the train-only gate cannot see.

Ground truth is a corner where outcomes are structurally decidable and the cost axis is zero. Outcomes are exactly determined, but real incidents carry multi-step, partially ordered repair, partial credit, and environmental non-determinism. The metric-to-parameter map is designed to be calibrated for a live cluster, and calibration is not reported in this paper. In the deterministic agent, the cost axis is identically zero, so the cost-quality frontier is characterized at only one corner. The next measurement is the metered LLM arm, free-text lessons and a cost cap, measured on the same frozen splits.

The paper and accompanying materials are available on Hugging Face.

https://huggingface.co/datasets/thaki-AI/daily-paper-2026-10-11-overnight-oncall-skill-evolution-k8s

Share this article:

Tags: agent-harness, autonomous-remediation, cost-per-incident, fault-injected-holdout, goodhart-divergence, k8s-incident-self-evolution, kind-holdout, overnight-skill-evolution, promotion-gate, rca-skills

Categories:

Updated: