Is 4-Bit Safe for Tool Calls: Pricing the Precision Cliff of Agentic Structured Output
Serving at 4-bit reliably brings costs down. But the structured calls an unattended agent emits can break faster than free prose at the same low precision. If you self-host unattended agents, or if you answer for that serving bill as a cloud engineer, the question this post answers is today’s work. The paper defines whether 4-bit is safe for structured calls with a single number and aligns the location of that cliff with serving costs.
The background stays short. The same model checkpoint is re-encoded at three precision tiers and served on a single H200. The requests come in two kinds: tool calls, scored against a calling-structure contract, and prose, scored on factual correctness.
This post has one question. Where is the precision point at which tool-call fragments break faster than prose, and is that loss more expensive than the bandwidth saved? The answer ends in three operating rules.
A visual metaphor for the article’s key idea.
In Plain Terms
When you send money, you write the account number and the purpose field. The account number is a short string, but a single wrong digit and the transfer fails outright. Write the purpose as a gift but misspell it, and the money still arrives just fine. The outputs of an unattended agent look the same way.
Free prose is the purpose field, and structured calls are the account number. Calls whose function name, argument values, and call count are consumed verbatim by downstream code stop the whole workflow when a single token goes off. A typo in prose is absorbed inside that sentence, but an error in a call comes back as a retry and a user-visible failure. At the same precision, the two output types can break at different rates.
Lowering the quantization precision makes token-level judgments wobble. But the wobble is not uniform. The tokens that hold the structure up, namely the call boundaries, function names, and argument value positions, sit exactly where the margin between the right answer and the wrong one is thinnest. So when precision drops from 8-bit to 4-bit, prose typos stay typos, while a single digit flip in the account number becomes a failed transfer. The paper calls that point the precision cliff.
Infographic generated by NotebookLM from the sources.
Why There Is No Evidence Yet for Answering Whether 4-Bit Is Safe
Precision is the first lever for cutting serving costs. Unattended agents run in the batch-1 regime, one request at a time. In this regime, emitting a single token requires reading the entire weights once, so GPU-seconds per case scale almost proportionally to the number of weight bytes. Precision is the lever an engineer pulls before buying another GPU.
Yet the evidence available to answer this question is still only perplexity (the model’s sentence prediction error) and prose accuracy. Compression benchmarks leave agentic capability out of scope, and the closest external measurements show only that standard success scores stay flat across 16/8/4-bit. If the overall score does not move, it looks lossless, but which requests broke is left unanswered. Tool-call requests and free prose have not been measured separately, and there are reports that low precision can hide a hidden cost, inference token bloat, inside the accuracy column.
Three things are missing. How do tool-call fragments break as precision is lowered, at which boundary do they break, and is that loss more expensive than the bandwidth savings? No controlled experiment answers these three questions yet. This paper collects the three questions under a single definition.
Two Definitions That Measure the Cliff, and the Amplification Mechanism
The paper’s first contribution is two definitions that measure the cliff. The conditions are pinned to one point. The same dense checkpoint is re-encoded at three precision tiers and served on a single H200 with thinking mode off and temperature 0. The model scale is 27B dense.
The first definition is the accuracy tax. With 16-bit as the baseline, it measures how much each precision tier cuts the per-category pass rate. The second definition is the cliff ratio. It is the tool-call tax divided by the free-prose tax. If the cliff ratio exceeds 1, tool-call fragments break faster than prose at the same precision. This is where the claim that a precision cliff exists gets its number.
The tax concentrates on the structured-call side because of the structured decode stages. Tool calls have stages that must be obeyed: call boundaries, function names, call count, argument value positions. The core of the organizing premise is that exactly those stages are where the margin between the right and the wrong answer is thinnest. When quantization noise wobbles token-level logits, judgments flip starting from where the margin is thinnest, and a flip at a single structured stage fails the whole case.
So the case failure probability amplifies in proportion to the number of structured stages. The amplification is largest in the parallel-call categories, where multiple calls appear. The formalization leaves two predictions here.
The first prediction means a change in the per-category failure composition. Simple calls are dominated by argument value errors, and parallel calls show a marked rise in call-count errors. The second prediction concerns the location of the cliff. Because the noise variance quadruples with each step down in precision, the cliff stands at the 4-bit tier, and the tax jump from 8 to 4 is much larger than the jumps above it.
Concept diagram of the predicted shape of the precision cliff. The tool-call curve drops steeply from 8-bit to 4-bit while the free-prose curve declines gently. An example from the analysis model, not a measurement.
The formalization does not end alone. One measured point already in hand is used to check the prediction. This side’s 4-bit (NVFP4) checkpoint has a tool-call evaluation of 800 cases, with 89.6% overall accuracy.
The per-category pass rates and failure composition are in the table below. At this single point, the predicted failure composition is visible.
| Category | Cases | Pass Rate | Dominant Failure Mode |
|---|---|---|---|
| Simple | 400 | 90.8% | Argument value mismatch (35 of 37 failures) |
| Multi | 200 | 89.0% | Mixed (16 argument value, 4 call count, 2 function name) |
| Parallel | 200 | 88.0% | Argument value (14) + call count (9) |
| Overall | 800 | 89.6% |
The failure composition matches the prediction too. Most of the simple-category failures are argument value mismatches, and the parallel-category failures carry a heavy share of call-count errors. The fact that call-count errors, rare in the simple category, increase ninefold in the parallel category is the shape of the composition change the prediction spoke about.
The remaining work is completing this single point into a curve. Once 16-bit and 8-bit, and the free-prose control, are measured under the same conditions, the cliff-ratio curve is complete. The next section looks at what operating rules that curve leaves behind when it meets cost.
The Price of the Cliff: Break-Even and Operating Rules
Saying a cliff exists is different from saying the cliff has a price. This paper converts the cliff into money. It goes from the cost axis.
Batch-1 decoding is bound by weight reads. So GPU-seconds per case can scale with the bit count, ideally. 8-bit against 16-bit is half, and 4-bit is a quarter.
Concept diagram of GPU-seconds per case scaling with bit count in the batch-1 regime bound by decoding. The ideal costs of the NVFP4, FP8, and BF16 re-encodings come out to 1:2:4. An example from the analysis model, not a measurement.
That ratio is an ideal upper bound. Kernel overhead and the dequantization path narrow it in practice, and measured values go into the decision. Because the value is asymmetric. A wrong call stops a workflow and leaves retries and user-visible failures, while an error in one sentence of prose is absorbed inside that sentence.
So the break-even condition shrinks to one line. If the bandwidth savings from running calls at low precision exceed the accuracy tax with value attached, unifying the whole fleet at low precision is optimal; if not, the 2-tier policy that raises only calls to one higher precision tier wins. Thanks to this inequality, the standing decision shrinks to two measured values: the 4-bit call tax and the measured cost ratio.
Is 7 out of 10 requests a call in this traffic? The table below compares the three policies under this cost ratio. Even with the cliff, the floor is not bad, and without the cliff the precision debate disappears from the bill. In plain terms: in no case is it necessary to run everything at 16-bit.
| Policy | Calls | Prose | Savings vs 16-bit |
|---|---|---|---|
| Unified 4-bit | 4-bit | 4-bit | 75% |
| 2-tier | 8-bit | 4-bit | 56% |
| Floor | 16-bit | 4-bit | 19% |
The 2-tier policy is, so to speak, writing only the account number precisely and keeping the purpose field rough. The operating rules are three. If the 4-bit call tax measures within 2pp of the prose tax, unify the whole fleet at 4-bit. The simplest answer, needing neither a classifier nor routing.
When a cliff is confirmed, route by task type. The needs_tools flag already in the request is the coordinate. Any rule wraps a safety net around call fragments. Calls are validated against the contract, and on failure, retried exactly once at a higher precision tier. Since failures sit around one in ten, this retry costs almost nothing.
The experiment that arbitrates the rules is defined along with them. A four-arm design that pairs 16, 8, and 4-bit calls with the free-prose control on the same cases. Each arm’s falsification criterion, R1 through R4, is fixed in advance. Four things: the presence of the cliff, the monotonicity of the tax, the failure composition, and the break-even verdict. A structure where hitting a criterion collapses the conclusion with it: that is the design of this paper.
What this post leaves the company is the standing decision of the token factory. Whether to run agentic workloads at 4-bit, or to apply an accuracy tax to call requests and route 2-tier, is answered with two measured values. Socially, it cuts the compute and energy cost of autonomous agents running 24 hours a day.
Identifying which workload types are safe at low precision makes always-on agents affordable even for small teams. Scientifically, it leaves the first attempt to measure the interaction between task type and precision at the same model scale in a controlled experiment. It publishes the call-cliff curve that difficulty-based precision routing missed and completes the price of the three axes, the thinking-token axis, the retriever axis, and this one, in the same token factory.
Infographic generated by NotebookLM from the sources.
The Parts You Should Not Trust
This paper is an analysis paper, and it admits it. The only measured number entering the body is the single 4-bit baseline row. The cliff curve and the superlinear jump are still predictions, and the four-arm design arbitrates whether they hold.
First, the model is one family only, 27B dense. Curves crossing size and model family remain future work. Second, the 4-bit on H200 is a storage and bandwidth story.
The weights are read at 4-bit, but computation falls back to 16-bit. So the measured cost becomes a conservative upper bound for native 4-bit hardware. The accuracy criterion is independent of the compute path, but the break-even margin must be measured again on Blackwell.
Third, the grader is this side’s contract checker, not an official benchmark checker. Relative comparisons across tiers are valid, and leaderboard absolute-score claims are out of scope. Fourth, only deterministic decoding is covered.
Whether the cliff survives stochastic decoding is an open question. Fifth, the 70% tool-call share is this traffic’s statistic. Fleets with a different share must recompute the break-even.
Sixth, every tier here is post-training quantization (PTQ). 4-bit QAT, which lowers precision during training, moves the cliff rather than erasing it. The batch-1 regime is also fixed, and batching effects and prefill sharing change the cost.
The paper’s detail page is available here: The Tool-Call Cliff: Measuring the Accuracy Decay of Agentic Structured Output Under Low-Bit Quantization in Self-Hosted H200 Serving
Both figures in this post are concept diagrams from the analysis model, not measurements. The only numbers entering as measurements are the 4-bit baseline (800 cases, 89.6% overall).