The Silent Reroute: When One Line of a Skill Description Changes, the Routing of Untouched Tasks Changes Too
ThakiCloud’s production agent harness rewrites its own skill descriptions every night. Those descriptions are the only routing signal for the hybrid router that sends each natural-language turn to one of roughly 2,200 skills. If you are a cloud/AI engineer accountable for the cost of unattended runs, whether the system is a self-healing agent registry, an MCP tool registry, a plugin store, or a prompt library, you should read this post. The question is not whether the edited skill still routes well, but what the edit did to the tasks it never touched. That change is invisible to the current regression gate. This paper measures the silent reroute as a law and gives a gate that catches it before merge.
The Problem: The Gate Watches Only the Edited Skill, and the Rest Stays Silent
The registry grew from ~1,600 skills in July to ~2,200 in October, and the curator sweep that ran every night in between had already rewritten 11 SKILL.md descriptions. It ran unattended, and it was invisible to the gate.
The router the fleet relies on is hybrid. It fuses a BM25 lexical lane and an embedding (dense) lane at the rank level with reciprocal rank fusion (RRF). Both lanes read the description, and this is exactly where the registry-side mechanism lives: the lexical lane is corpus-coupled. Every term’s idf is a function of term frequency across all skills, so rewriting one description moves the lexical score of every bystander skill that shares the term, across all queries at once. The dense lane moves only the embedding of the edited skill. Either kind of movement can flip the top-1 decision on a task that has never touched the gold skill. The paper calls this collateral drift.
The current regression gate (SRA) re-scores the updated registry against a fixed 63-case golden suite, but all it checks is whether the updated skill still routes. Collateral reroutes of untouched gold skills are outside the gate’s line of sight. The sweep merges, the reroute ships, and the fleet’s behavior quietly changes. This paper is an explicit one-step upgrade over the earlier “Quantizing the Gatekeeper” post. That work showed that top-1 rank safety for a single query can be bounded only from perturbation of the dense component. This time the perturbation axis moves from the model side to the registry side, and the victims change from one edited query to every remaining task. Both studies spend the same risk currency: the displacement of the dense lane, measured against the margin to second place.
Core Contribution: A Drift Law as a Function of Edit Distance, Edit Type, and Registry Size
The center of the paper is the first-order collateral drift law. It writes the expected collateral top-1 flip rate as a function of the semantic distance of the edit, the edit type (paraphrase, semantic shift, truncation), and the registry size N. The law is calibrated to the recorded production point: N=2,029, top-1 0.489, Recall@5 0.867, survival curve P_top1(N)=0.867·e^(-2.83×10⁻⁴·(N-1)). Every number in the paper is either this recorded anchor, first-order arithmetic on top of it, or a value explicitly labeled as a model prediction.
The scaling law comes out of the hazard-margin model. Diffuse dense-channel drift is proportional to N·e^(-p(N-1)), and that function is maximized at N*≈3,535. At the recorded scale, g(2,200)/g(1,600)≈1.16. The same single sweep now causes ~16% more diffuse collateral drift than it did in July, and the registry sits on the rising leg of the drift curve. We flag it as a model prediction.
Drift law prediction: diffuse collateral top-1 drift rises with registry size N, peaks near N≈3,535, and declines beyond it. The current ~2,200-skill registry is on the rising leg. (Interpretive model, not measurement)*
The type law gives a severity ordering. At the same dense displacement, semantic shift ≥ truncation ≥ paraphrase. Paraphrase is idf-neutral up to first order and sits at the harmless noise floor. Truncation, i.e., deletion without replacement, raises the idf of a rare term by ≈1/(n_τ-0.5). For n_τ between 5 and 2, that is 0.22~0.67. Because the spike concentrates on the n_τ-1 bystanders that share the term and on the queries that contain it, truncation is the sharpest lexical lever. Semantic shift adds, on top of a large dense displacement, a targeted idf term of ≈7.29 per additional term not previously in the corpus (at N=2,200). If the injected term hooks into a suite query, the flip rate is maximal. This is the same rare-term lever that was the attack scenario in the earlier registry hijack paper; for the curator, it is a quiet risk.
At the same dense displacement, the expected collateral top-1 flip is larger for semantic shift than for truncation, and larger for truncation than for paraphrase; paraphrase sits at the harmless noise floor. (Interpretive model, not measurement)
The third readout is the silence itself. Applying the same hazard argument at the top-5 boundary gives ΔR@5 ≈ 0.25·F. A sweep that flips collateral top-1 accuracy by 3 points moves Recall@5 by only ~0.75 points. Recall-centric monitoring under-reads registry edit drift by ~4x. The reroute is silent at the recall level.
SRA-CI: The Gate That Catches Routing Regressions Before Merge
Where the drift law translates into engineering is SRA-CI, the pre-merge shadow-routing regression gate. The arithmetic behind it is pure binomial arithmetic. For a suite of effective size n with a single-case trigger, gate power is π(n, p_f)=1-(1-p_f)ⁿ. At the recorded suite sizes (63 nominal, 42 effective as of September), power is 0.469/0.344 at p_f=1%, 0.720/0.572 at 2%, and 0.960/0.884 at 5%. The current gate catches sweeps that harm 5% or more of collateral tasks well, but sweeps in the silent band, 1~2% flips per case that move Recall@5 by less than 1 point, are missed with probability 28~66%. Reaching 95% power at p_f=1% takes n≈299 cases, and the error-bound arithmetic of the rate trigger lands on the same ~292-case design point. The gate’s suite is ~300 cases.
| The design is concrete down to the executable level. For each sweep with edit set E, build a candidate index: re-embed only the | E | edited descriptions, update lexical statistics incrementally, no full reindex. Re-score the SRA suite and the collateral suite (tasks whose gold is outside E) under both the baseline and candidate indices, and compute F_top1, ΔR5, and the severity-weighted flip mass in cost units as drift readouts. Block if F_top1≥τ_block, or if a single flip exceeds the severity cap c_max. A single costly steal, a flip onto a long-running skill or a high-privilege skill, is blocked regardless of rate. Otherwise merge, with a routing-impact statement in the changelog. A nightly suite validity audit re-confirms the presence and optimality of each gold, retires and rebuilds invalid cases, and tracks the effective suite size n_eff in telemetry. Validity is part of the statistics: at n_eff=42, the 63-case gate loses another 13 points of power at p_f=1%. The hybrid router over a fixed index is deterministic, so what the gate sees is the true drift of the sweep, not a sample. The cost is light: an | E | =11 sweep with a 300-case suite is ~14% of the full re-embedding dense work, and amortized over a ~52-day churn cycle, under 0.3% of the nightly re-embedding budget. |
SRA-CI: each night’s curator sweep is re-scored against the SRA suite and the collateral suite on the candidate index before merge, and harmful description edits are blocked with an impact report instead of shipping quietly. (Conceptual illustration: a gate-structure diagram with no quantitative claims)
What It Leaves Behind for Company, Society, and Science
For the company, a pre-merge checkpoint for the self-evolution loop now exists. Every curator sweep is re-scored against the candidate registry, and edits that push the collateral top-1 flip rate or the Recall@5 drop past a calibrated threshold are blocked. The description churn that is blind today becomes CI-checked skill routing in a 2,200-skill harness. For society, fleets of unattended agents that heal their own registry need an auditable bound on quiet behavior change. The drift measurement methodology and the minimal regression suite transfer to retrieval-based skill and tool ecosystems, MCP tool registries, plugin stores, and prompt libraries. The case for transfer rests on three structural facts: natural-language descriptions are the substance of the routing signal; the lexical lane couples every document’s score to the frequency of every term, so even a harmless edit produces bystander effects; and the validity of the fixed golden suite that tests the ecosystem itself degrades. For science, this is the first controlled measurement of registry-side routing instability. It complements the query-side paraphrase instability and index staleness results from the same lineage, and the drift law (top-1 flip rate against edit distance, type, and size) plus the required suite size make “routing regression”, changes in the router’s decision on unchanged inputs caused by changed corpus artifacts, a named reliability dimension. It also composes with the earlier quantization study: edit displacement and the quantization radius spend the same margin budget additively.
Limits: Interpretive, First-Order, and Not Yet Measured
The drift law theorem is first-order in dense displacement. Batched sweeps have second-order collision terms, so it overestimates for harmless edits and underestimates for clustered injection (the hijack regime). Neighborhood-mass scaling, the assumption that each skill holds a constant share of the workspace, is the only assumption the recorded point cannot pin down; the constant α is calibrated telemetrically. The hazard-margin model inherits the constant-hazard fit of the lineage, so structural change in the skill population can move p. The query population is fixed, and input-side paraphrase flips compose multiplicatively against the same margin. The gate cost arithmetic is all computation, not measurement. The paper pre-registers the falsification criteria for H1~H5: the type ordering; 10~25% more diffuse flips at N≈2,200 versus N≈1,600 (model value ×1.16) and stabilization after passing 3,500; median ΔR@5 at no more than 0.4x median F_top1; the per-flip-rate numbers for suite power; and deletion of n_τ≤3 terms producing a lexical spike that crosses the median marginal-difference gap. Once the gate is live, nightly sweep telemetry turns this interpretation into measurement.
Paper and data: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-10-03-silent-reroute-skill-router-drift