🎧 ▶ Listen to the 5-minute briefing
▶ Play audiobook (Google Drive)
Locally synthesized AI audiobook (Qwen3-TTS)

In late August, at the Dell Technologies Forum 2026 on stage at COEX in Gangnam, Seoul, a warning worth noting was delivered. “It encourages AI usage, but it is never sustainable.” That was the line a Dell presenter used to frame the problem of enterprise AI cost. Every company has a different reason for using AI, but the point was that the bill on the books always seems to slip out of control no matter what that reason is.

The unit price of AI drops every day. Cheaper means more usage, more usage means bigger volume, and bigger volume means a thicker bill. The cycle looks obvious, yet nobody has a real answer for it. And this story is not new. Economic history dealt with it once already, 160 years ago, and even gave it a name.

Image visualizing the concept of the day tokens get cheaper and bills get bigger A visualization of the article’s core concept.

The 1860s Paradox of Coal

In mid-19th century Britain, there was a widespread expectation that improving the efficiency of steam engines would save coal. William Stanley Jevons argued the opposite. As engines used coal more efficiently, the real cost of running steam went down. As the cost went down, more factories started running engines, and as more factories ran them, total coal consumption actually increased. Jevons called this the “paradox of economical improvement.” As a resource gets cheaper and more efficient to use, total consumption does not shrink. It grows.

This paradox has resurfaced again and again beyond coal, through energy and computing, and now AI. The pattern is the same each time: the moment price sends a signal, consumption reacts faster than the price itself. Today’s coal is the token, and today’s steam engine is inference. The cycle Jevons described is playing out again, this time on the AI cost ledgers of today’s enterprises.

Unit Prices Fall, Bills Rise

The diagnosis Dell laid out at the forum was concrete. Token unit prices are falling, but the bill is rising because total usage is growing. That’s why token flow needs to be built into AI architecture design from the start. The presenters proposed a hybrid setup: let the cloud handle the latest models, demand spikes, and large-scale training, while bringing repetitive inference, personalization work, and sensitive data down to the enterprise desk side.

Cost and control are the reasons for this split. Repetitive, predictable inference is work that needs to run cheap and stable, and pulling it out of the data center is what makes the bill predictable. Sensitive data is simply safer never leaving the building at all. According to Byline Network, 2026 is being called the year on-premises AI inference demand surges as reliance on the cloud declines. The same logic is behind manufacturers like Dell, ASUS, HP, and Lenovo rushing into inference environments outside the data center, centered on NVIDIA GB10-based workstations.

The pace at which prices are falling is faster than expected. AIVE, a Korean GPU cloud startup selected for TechCrunch’s 2026 Startup Battlefield 200, combines RTX-class GPUs with idle data-center-grade GPUs into a single operating environment to provide inference infrastructure. It already runs more than 30 paying customer sites and claims inference costs 40 to 80% cheaper than public cloud. Even if only half of that claim holds, a meaningful amount of “economical improvement” in inference has already happened. The split landscape of the inference infrastructure market points the same way. One side is neocloud built around large data centers, the other is distributed aggregation that turns idle GPUs into a service. The directions differ, but the premise is the same: token unit prices keep falling, and the more they fall, the more workloads pile onto inference. The premise of the Jevons paradox is complete. What’s left is the consequence: total consumption swelling.

Where Cheap Tokens Go

Companies don’t let cheap tokens sit idle. They hand them to agents.

SK AX and SAP signed a business agreement at SAP’s headquarters in Walldorf, Germany, to build an “AI-native enterprise.” The goal is to redesign integrated ERP operations, including accounting, HR, procurement, and inventory and sales, around agentic AI. It’s a turning point past the era when software merely recorded information, toward an operating structure where agents judge and execute. SK AX plans to register its own AXgenticWire ERP Suite with SAP’s AI agent hub to supply it to global customers, and will also join an early-adopter program that applies the product to real business environments before its official launch. Finance, manufacturing, and energy are the priority industries.

What this signals is significant. AI adoption is moving from chatbots and one-off automation into the interior of core business systems. ERP is the system that ties together a company’s money, inventory, trading partners, and HR data. The moment an agent that judges and executes sits on top of that, AI becomes part of the operating structure. And once a tool becomes part of the operating structure, the question shifts to who is using how much, and what they did with it. The more an agent takes over judgment and execution inside ERP, the more cost and risk grow together if there’s no mechanism to answer that question.

According to a global consulting firm’s forecast, companies that restructure operations around agentic AI see roughly double the revenue growth and more than 40% in cost savings compared to companies that don’t. If those numbers hold, cheap tokens become the fuel that doubles revenue. And the cheaper the fuel, the more of it gets burned.

The national level is heading the same way. At the ETRI Conference marking its 50th anniversary, ETRI put forward “AI Co-Scientist for Everyone” as its core concept: a multimodal foundation model that supports the full cycle of hypothesis, data analysis, experimentation, and verification. The Ministry of Science and ICT, ETRI, and KISTI are putting in 150 billion won with the goal of commercializing a Korean AlphaFold equivalent by 2031. The backdrop is a national goal to invest more than 1,000 trillion won in AI data centers and expand GPU capacity 15-fold by 2035. The flood of cheap tokens does not stop at a single chatbot. It’s being pulled into ERP systems, into research labs, and into national infrastructure plans.

The Flood Has a Temperature

The flood of tokens carries a physical signature too: heat. As agents multiply, inference grows, and as inference grows, power consumption grows with it. LG CNS is adopting direct-to-chip (DTC) cooling at its Samsong data center in Goyang, Gyeonggi, running coolant through cooling plates attached directly to GPU chips to absorb heat at the source. Traditional air cooling maxes out at 20 to 30 kW per rack, while next-generation GPU racks are designed for power densities above 200 kW.

Together with Naver Cloud, LG CNS is building an 80MW AI Factory under a design-build-operate integrated model, with plans to complete demonstration by 2027, premised on operating NVIDIA’s Vera Rubin, which offers roughly 3.3x the inference performance of Blackwell Ultra. When performance goes up 3.3x, the same workload can run more inference on fewer chips. And that brings us back to Jevons: as unit cost drops, total volume rises. Per MarketsandMarkets, the global data center liquid cooling market is projected to grow roughly 6.8x, from $4.07 billion in 2026 to $27.65 billion by 2033.

Alongside coal, Jevons also mentioned iron and limestone. Better machines meant not just coal but other resources saw rising consumption too. The heat from token bills is already showing up on other ledgers, in the form of electricity bills and cooling budgets.

The Missing Variable Isn’t Price, It’s Control

So what’s the real problem? It isn’t price. It’s that while usage keeps growing, the means to control it aren’t growing along with it.

This is exactly where Dell’s warning lands: a lack of means to monitor and centrally control token usage. An agent making judgments and taking action inside ERP. A Co-Scientist running experiments overnight. Desk-side inference handling sensitive data. Each of these is a reasonable choice on its own. But when thousands of reasonable choices pile up without a meter, without a gate, without a record, a company only learns the facts when the bill arrives at the end of the month.

The reason this is dangerous is simple. Uncontrolled usage isn’t just a cost problem, it spills into data leaks, audit gaps, and regulatory risk. If tokens are cheap and there’s no control mechanism, they get used more, more broadly, and faster, in direct proportion to how cheap they are. The Jevons paradox isn’t a problem you solve by making tokens expensive again. You can only escape it by placing a point of control wherever tokens flow. Ultimately, the question of “how much are we using” needs to become the question of “who gets to use it, how far can they go, and how do we verify it.”

Who Holds the Bill

ThakiCloud’s Paxis is a formal product built as an answer to exactly this question. Paxis is an agent-native cloud that manages skills, tools, policies, and audit logs as first-class resources. Each agent operates at an autonomy level from L0 to L3, risky actions must pass through a policy gate, and every execution is recorded in an audit log. It runs in isolated sandboxes, connects to external systems through MCP connectors, and also runs on an enterprise’s own K8s. It even has a CostRouter that picks the appropriate model for each task.

The pain points that surfaced in today’s news map one-to-one onto this structure. What Dell called the “absence of a central control mechanism” is exactly what Paxis’s combination of policy and audit logs addresses. The suggestion to “build token flow into architecture design” becomes a concrete setting through CostRouter. The move to bring sensitive data down to the desk side overlaps with Paxis’s on-premises path, which provides isolated execution on a company’s own K8s. The ERP agents SK AX is pushing and the Co-Scientist ETRI is preparing are exactly the shape of agent that Paxis’s autonomy governance and sandboxed execution already assume.

In the end, there are two choices. Put a point of control in place while tokens are still cheap, or fit control mechanisms in later, after the bill arrives. The former turns Jevons’ consequence into your own advantage. The latter makes you a victim of the paradox. In an era where token unit prices keep falling, the value of a company’s point of control keeps rising right alongside it.

The Next Ledger

The day AI got cheaper. That’s also the day the bill got bigger. Jevons looked at coal and said usage would increase. Today, we looked at tokens and confirmed the same sentence still holds.

But there’s one difference in texture from back then. The steam engine had a throttle. No matter how good efficiency got, if someone closed the valve, no more steam went out beyond that point. Agents still lack that kind of throttle. So today’s warning really comes down to a matter of control. The companies that attach meters and gates to the pipeline where tokens flow first will see a different number on next month’s bill.

References

This article was written by synthesizing the following news sources.

Tags: agent-infrastructure, enterprise-ai, gpu-cloud, jevons-paradox, paxis, token-economics

Categories:

Updated: