🎧 ▶ Listen: 5-minute briefing
▶ Play audiobook (Google Drive)
Locally synthesized AI audiobook (Qwen3-TTS)

Image visualizing the concept of 30B in the pocket and 15x in the office Visualizing the core concept of the post.

A smartphone spec sheet that includes the word “parameters”

In this industry, the spec sheet is read first. On September 22, local time, Qualcomm held its annual Snapdragon Summit. On the table were two top-tier platforms: the Snapdragon 8 Elite Extreme Gen 6 and the Snapdragon 8 Elite Gen 6. Both are built on a 2nm process, which stands out too. Since the first summit in 2016, Qualcomm had shipped exactly one top-tier 8-series chipset per year. In this 10th anniversary year, the top tier is, for the first time, two platforms, and Qualcomm calls the arrangement dual Elite. It explains the lineup as a response to form-factor diversification such as foldables and tri-folds. Different form factors mean different agent experiences, and it signals the end of the era in which a single chip covered every phone.

The more interesting part is not the count but what is written in the spec sheet. For the first time in mobile processor history, the 5GHz clock barrier was broken. The 8-core custom Oryon CPU is built from two prime and six performance cores, and a new Oryon Flex Cache based on an L2 shared pool was added. The Hexagon NPU supports up to a 32K-token context, on-chip shared memory was expanded by 50%, and prefill performance improved by up to 80%. Combined with UFS 5.0, a 30B-class MoE model runs directly on the device. Compared to 2023, when the Snapdragon 8 Gen 3 brought 10B-class on-device generative AI to smartphones, the class of models that fit in a pocket has tripled in three years.

Chew on the 32K-token figure and it is clear this spec sheet is for agents. An agent does not answer once; it reads documents, calls tools, and verifies results, accumulating context as it goes. Remembering long and producing the first sentence fast: those two have become the criteria of the spec at the same time.

Beneath that was a quieter change. The dual micro NPU and Qualcomm Sensing Hub are said to cut power by 20% while lifting performance by 85%. The Personal Script feature turns everyday conversation into an on-device knowledge graph to help personalize the agent. And the CPU itself is described as handling the orchestration, control flow, and task scheduling of agentic workflows. The word workflow entered the chip’s documentation. It means agent work has become a first-class workload of the chip, not an app. Battery, heat, and power budgets start to be occupied by the agent’s work directly.

Qualcomm’s own summary was blunt. It says the criterion of smartphone competition is moving from hardware specs to what an agent can execute on the user’s behalf. It also brought out a “my ecosystem” vision: the smartphone gathering the context of PC, XR, audio, and wearables as its center of gravity. The Xiaomi 18 Pro has confirmed world-first adoption, and the race for early adoption among major OEMs has already begun.

Key-concept summary infographic 1 Infographic generated by NotebookLM from the sources.

Same week, 15x in the office

In another room, a different number came out this week. According to materials Microsoft released on the 23rd, the number of agents in the MS365 ecosystem grew 15x in one year. Large enterprises recorded 18x. Sixty-six percent of surveyed AI users answered that AI use let them spend more time on high-value work.

OpenAI pinned this change to a single word: delegation. Its analysis is that enterprise AI use is moving from support to delegation. It published examples of applying agents to work such as onboarding, customer management, and developer support. In survey samples, 70.2% of Codex users delegate tasks that would take over an hour, and 25.6% delegate work worth more than 8 hours, to AI. Delegating eight hours of work: that phrase shows the level of today.

Google Cloud goes one step further. It points out that agent adoption is complete only when separated work tools are connected and data access rights, security, and cost are managed together. It sees the criterion of competition moving from model benchmark scores to how much error is reduced when complex work is delegated and how reproducible the results are made.

So the office-side story goes like this. What people hand to agents is no longer questions but work. Files to check, data to analyze, programs to run, tools to call. Composing and completing multiple steps on their own. And that volume is growing at 15x speed. For the near term, the industry consensus is that a division of labor will spread rather than large-scale job replacement: AI performs parts of the work, and people take on goal-setting, decision-making, exception handling, and result verification.

Two rivers flowing in opposite directions

Placed side by side, these two make a strange picture. Consumer-side agents are moving down into the pocket, while enterprise-side agents are moving up onto the platform. The same word, agent, is flowing in two directions at once.

Why? The reasons for the side coming into the pocket are clear. Latency must be short, it must run inside the battery, and the fact that the data of conversation does not leave is an advantage. When a 30B-class MoE comes down, light agent experiences like scheduling or booking agents are absorbed by the device. And there is an interesting side effect. When consumers first internalize the agent experience on their phone, the psychological barrier of enterprise customers lowers. It is bringing back to the office work already tried in the pocket.

The side going up has its reasons too. Enterprise data is not on one device, permissions differ by person, and results must be audited. The bigger the delegation, the more what is needed is not a smarter model but a more explainable structure. Looking at both flows together shows today’s competitive landscape. Both sell intelligence. Both sell execution. The smartphone expands the range that can be executed on the user’s behalf, and the office expands the range of work an agent can finish. If this movement needs a unit, it is work. In the pocket the unit of work is personal practice; in the office the unit of work is a business process. The size of the unit differs, but the measure is the same. The criterion moving from hardware specs to execution means the two news stories are, in the end, saying the same thing.

Yet looking only at the execution environment, the two appear to be entirely different things. In the pocket, the data is yours, the blast radius of a mistake is your daily life, and the price of failure is one canceled reservation. The office is different. An agent stands before ERP systems, customer data, and payment information, and if it gets it wrong there, the question is not “it made a mistake” but what was done, under whose authority, and under what record.

That is the line dividing today’s agent terrain. It is not model size. Nor is it a parameter count. Which data is touched, whether the process can be explained, and ultimately whose responsibility it becomes. Size determines speed and cost, but trust determines the seat.

The line is also drawn by a price list

That line is not drawn by trust alone. A price list draws it together. Inference hardware is getting more expensive. According to BizWatch analysis, HBM4 prices are observed to rise about 65% this year. Next year’s HBM demand growth rate is projected at 56%, outpacing the 50% supply growth. TrendForce forecast that Q3 commodity DRAM fixed transaction prices will rise 13-18% quarter over quarter.

The HBM landscape is shifting fast too. In Counterpoint Research tallies, Samsung’s share jumped from 21% in Q1 to 33% in Q2, and SK Hynix recorded 50%. Samsung’s HBM4 yield is reported to have risen from below 60% at the start of mass production to around 80% recently. The global HBM market this year is projected at about $54.6 billion, up 58% year over year. The more memory costs, the more inference that uses it costs. Nvidia also announced a target of doubling chip sales next year.

The more expensive hardware gets, the more “how much does running this task once cost” becomes an executive question. Akamai Korea’s Gang Sang-jin, executive officer, says the future of AI infrastructure rests on distributed environments. Its Korea head Lee Kyung-jun defines AI infrastructure as extending from core to edge, that is, to places close to users and data. The picture is one in which inference does not end in a central data center but spreads out toward users and data. That is the same point where Akamai names security as the core of practical AI operation: past deployment, scalability, and optimization, the deeper it goes into actual operation, the more security becomes the core.

For enterprises, the question changes from which model to use to which task to run where. What fits the pocket goes to the pocket; what must be explained goes to the platform with governance. If on-device absorbs light work, the calculation comes out here too: cloud inference cost structure must be designed together with on-device routing. What makes that judgment possible is governance that can route by task and leave an audit trail.

Enterprise agents: the question of the seat

So what shape is the enterprise seat? It is the place that defines which agent can do what. What was done, when, and under whose authority remains as a record. It isolates the execution environment and picks the model that fits each task. ThakiCloud’s Paxis is an agent-native cloud designed for this question, and it is already on the market as a formal product.

In Paxis, skills, tools, policies, and audit logs are treated as first-class resources. An agent’s autonomy receives governance step by step from L0 to L3. Policy gates verify each execution and audit logs leave a record. Execution happens in an isolated sandbox, and an enterprise’s own tools can be connected through MCP connectors and the skill marketplace. Sovereign and on-prem K8s deployments are also possible. CostRouter selects the model per task. So that work fitting the pocket does not needlessly climb to the cloud, and so that work that must go to the cloud is not unconditionally handed to a cheap model.

In organizations like finance, manufacturing, and public sectors where data does not leave the building, this question grows a notch larger. For them, the platform is not a matter of preference but of requirement, and on-prem and sovereign deployment become the starting point of the discussion.

The agent’s roadmap is being drawn. The pocket takes the light and fast work, and the platform takes the work that must be explained. What today’s spec sheet showed is that this division has already started at the hardware level. The speed at which 30B comes down and the speed at which 15x climbs up are opposite sides of the same agent era. What enterprises must decide now is which seat to give to which agent. The pocket-side spec sheet has already been rewritten, and the office-side seat is the next agenda.

Key-concept summary infographic 2 Infographic generated by NotebookLM from the sources.

References

This post is a synthesis of the news below.

Tags: agentops, enterprise-ai, paxis, thakicloud

Categories:

Updated: