The $300 Billion Behind the Profit Turning Point
This morning, the person selling the chips said that AI hit its profit turning point this year. The side using those chips put out a cash burn forecast of $278 billion through 2030. The ‘profit’ of this industry and the ‘money’ that supports that profit are still standing in different places. Follow the gap and you can see where the next bill for enterprise AI will be written.
An image visualizing the core concept of the article.
The Surface of the Profit
In an interview on CBS Sunday Morning, Jensen Huang said AI became actually useful this year and that token generation was ‘staggeringly profitable’ through 2026. He went as far as the phrase profit turning point. Read the phrase precisely and the first thing you see is who the subject of the profit is. For the chip seller, more sales mean more profit. The more models generate tokens, the more of the chips and data centers that generate them get sold. So the claim that ‘token generation is profitable’ is already a statement of fact for the seller.
The logic hidden behind ‘actually became useful’ is the same. Tokens have to be useful to sell, and chips sell only when tokens sell. Expanding usefulness is expanding consumption, and expanding consumption is expanding chip revenue. A closed circle is at work here, in which the seller’s revenue and the buyer’s usage grow together. ‘Staggeringly profitable’ is not about the size of the revenue. It points to margin: what is left after subtracting the cost of generating tokens. In the chip business, that margin is protected by structure, because there are few who can make the top-tier chips, so price does not come down easily even as volume grows. The profit the seller is talking about is that margin.
One more thing to pin down here is who is speaking. The person declaring that ‘the industry has reached its profit turning point’ is not a neutral observer. It is the head of the chip-selling company himself. It is hard to read this ‘profit turning point’ as a measurement of the industry as a whole. It is closer to an assessment the seller makes from his own seat. The fact that the subject of the declaration and the subject of the profit are the same person has to be read together with the $278 billion the buyer put on the table.
The problem is the side listening: the companies that use tokens to produce actual work. For them, tokens are a cost. And that cost, as shown next, was moving in two directions at once that same morning.
An infographic generated by NotebookLM from a synthesis of the sources.
The Money That Moved Off the Balance Sheet
The money that sits in a different place from the word ‘profit’ is fairly clear today. OpenAI’s internal financial outlook projects cash burn reaching $278 billion by 2030. It is also reported that fundraising on the order of $1.2 trillion is under discussion. The plan is to put $856 billion into compute and infrastructure for training and inference over the coming years. That means the side growing the biggest models is the side burning cash the fastest. It is a structure in which the size of the bet on infrastructure determines the size of the cash burn.
And the way that money is being propped up has changed. The same morning brought reports that big tech companies issued $300 billion in AI guarantees off the balance sheet, off the income statement, to protect their credit ratings. The phrase ‘to protect their credit ratings’ is the key. If the money going into AI infrastructure were posted on the financial statements as is, the ratings would wobble, so the point is that it is pushed out in the form of guarantees while the ratings are kept. Among them is Nvidia guaranteeing up to $105 billion for the OpenAI data center project in Ohio. The chip-selling company has stood as guarantor on the construction of its largest customer, the one building data centers with those chips. The person who spoke of the ‘profit turning point’ and the person who wrote the off-balance-sheet guarantees are the same person.
Read structurally, this is a circle. The buyer burns cash to build data centers. The seller sells more chips as a result. The seller adds guarantees to the same projects. As the circle turns, more chips sell and more guarantees pile up. The ‘profit’ the seller talks about is supported by this circle, and the risk goes off the balance sheet in the form of guarantees. The profit is on the surface, and the risk is off it.
What is interesting here is the direction. In the past, the money for infrastructure went inside the balance sheet under the name of assets. Now it is being pushed out in the form of guarantees. This is a signal that the industry is maximizing two things at once: pull in as much capital as possible for the expansion, while not wanting to stake the credit rating on the confidence of that expansion. A guarantee is the financial technique of putting exactly that ‘confidence’ off the balance sheet.
Why ‘Profitable’ Started to Sound Like a Fact
There is a reason the word ‘profitable’ has started to sound like a fact. Token prices are falling. China’s StepFun announced its flagship agentic model Step 5 Preview, saying it provides performance comparable to Kimi K3 at 65% lower cost. It carries a 1 million token context and a vision feature. List API pricing is $1.00 per million input tokens and $2.70 per million output tokens. The open weights are scheduled for release on October 15.
The frontier model side is competing on price too. Anthropic is preparing Opus 5.5 against GPT-6 Astra, with a Tuesday launch as the target, and the input cost is set at $4 per million tokens. The same company is targeting a $2 trillion valuation in its IPO, with annualized revenue expected to pass $120 billion by the end of 2026. Yet its customer retention rate one year out shows at 22.5%. As prices fall, revenue grows and customers change often. The reason ‘token generation is profitable’ is true is that a single token has become so cheap that many are sold. And that ‘many’ is passed straight through to the buyer side as cost.
The number 22.5% deserves a separate look. It means that after a year, roughly 1 in 4 customers is still there. At the model layer, this is structure. Prices keep falling and alternatives keep coming out, so staying with one model is a matter of brief price comparison. The reason ‘profitable’ sounds like a fact is that the model layer is this unstable. For the seller, unstable usage is profit. For the buyer, unstable usage is the source of risk.
The Price Ladder That Is Widening
Even within the same morning, the price ladder is widening. One end is Step 5 Preview at $1.00 per million input tokens, and the other end is Opus 5.5 at $4 per million input tokens. The input price gap between the two models is 4x. And below that, the open weights released on October 15 are another variable. Because you can host them yourself, price becomes a variable that steps outside the frame of list API pricing from the start.
This means the cost of a single task is no longer fixed by ‘which model.’ Expensive models go only to the steps that need heavy judgment, and cheaper models go to the execution and summarization steps. Allocated this way, the average cost per task falls below the list price of any single model. The axis of cost management moves from ‘picking one model’ to ‘allocating across many models.’
For the enterprise buyer, this means purchasing is no longer a one-time choice. It is no longer the era of buying a model and using it for years. On a moving price ladder, you have to keep re-allocating the work. And the precondition of re-allocation is visibility. If you cannot see which task went through which model and what it cost, you cannot move the allocation. The problem of cost management is, in the end, a problem of visibility.
Budgeting changes with this shift too. In the past, budgets were set by model license or API quota. Now they have to be set by unit of task execution. Even with the same budget, you find savings only by looking step by step at where and how much is spent. The moment the unit of the budget moves from model to task, the enterprise has no choice but to record the execution of each task.
The Moment the Record Becomes the Question
When allocation becomes ongoing work, the ‘record’ of execution becomes a precondition. A record is a bundle of three answers: which task was executed, by which model, at what cost. Without that bundle, you cannot prove the effect of the savings, and you cannot explain what happened when something went wrong.
As agents grow from ‘answering’ to ‘doing,’ the weight of this record grows. That is because the unit of execution is not a single query. It is a chain of steps passing through multiple tools and models. As steps increase, cost accumulates step by step, and an error propagates from one step to the next. In such a chain, asking ‘who, with what authority, did what at which step’ moves beyond the domain of audit and becomes a precondition of operations.
The Moment the Axis of Profit Comes to Execution
That is the bill today’s news leaves behind. The profit is on the seller’s side. The money is supported by guarantees off the balance sheet. Tokens keep getting cheaper, and customers change often. In a market where all four of these are true at once, the question the enterprise should look at first is not ‘which frontier model to buy.’ It is ‘which model, at what cost, and with what record, is each task executed with.’
ThakiCloud’s Paxis is the agent-native cloud that handles exactly that unit of execution, and it is a full product (v1.1 GA). It responds to the pain the news revealed as first-class resources. To the pain of ‘tokens keep getting cheaper and models keep changing,’ CostRouter answers by allocating a model to each task. To the pain of ‘agents execute,’ the autonomy level from L0 to L3 and policy gates decide how far an agent may go on its own. An isolated sandbox contains the execution. Audit Logs leave each execution as a record that can be verified. Skills, Tools, and Policies are managed as first-class resources. MCP connectors and the skill marketplace connect external tools. Sovereign or on-premise, it can be placed on top of K8s.
When the axis of profit moves to execution, the question changes. The enterprise now asks what cost and what record each execution went through. Without that record, it cannot even negotiate the next bill. With the record, the cost of each execution becomes an asset that can be compared and improved.
If this reading is right, the 2027 enterprise AI market will stand on a different standard. The frontier model remains an asset of the seller, and the record of execution remains an asset of the buyer. Token prices keep falling. The value of the record that explains each execution keeps rising. The companies that start building that record now are the companies that will turn the next price drop into their own margin.
An infographic generated by NotebookLM from a synthesis of the sources.
References
This article was written by synthesizing the news below.
- HuggingNews, OpenAI Forecasts $278 Billion Cash Burn Through 2030 in $1.2 Trillion Funding Talks
- HuggingNews, Anthropic Targets $2 Trillion IPO Valuation With 22.5% Customer Retention
- HuggingNews, StepFun Launches 600B Step 5 Preview AI Model With Oct 15 Open Weights
- HuggingNews, Anthropic Targets Tuesday Launch for Opus 5.5 to Fight GPT 6 Astra
- HuggingNews, StepFun Step 5 Preview Matches Kimi K3 at 65% Lower Cost and Opens Weights Oct 15
- HuggingNews, Nvidia CEO Jensen Huang Says AI Hit Profitable Turning Point This Year
- HuggingNews, Big Tech Issues $300 Billion in Off Balance Sheet AI Guarantees to Protect Credit Ratings