Proofs Out, Model in the Vault
Today, let me propose we look at the homework before the lesson. The release order in the AI industry is quietly flipping. The first crack in that flip starts with the repository of 722 math proofs that OpenAI posted to GitHub.
There is no model in the repository. Hundreds of new mathematical discoveries are grouped into 372 families, and the size is 722 on a draft basis. The side that produced this output is an OpenAI internal model that has not yet had an official release. The output arrived in the public domain before the product. The proofs are already under someone’s verification, and only the model that wrote them remains in the vault.
The grouping into 372 families is itself interesting. It is not one or two discoveries. It means there are enough to list by field. And the place of publication is GitHub. It is not the form of a press release bragging about a model’s benchmark score. It took a platform where code and output are verified as its stage. When the stage of an announcement changes, so does the kind of verification the announcement receives.
The position of the word “unreleased” has changed. In press releases it was always in the future tense. It was a modifier meaning something coming soon, but today it is a present-tense noun standing in the middle of the headline. The output arrives before the product, and verification runs ahead of release. Proofs are the kind of output that most clearly demand verification. That is because a verdict is possible whether they hold or not. When that kind of output accumulates 722 at a time and is published, the question the industry asks shifts from how smart the model is to who checks this output and how.
For the past few years, the announcement order of frontier labs has been fairly fixed. First they introduced the model, showed its capability in a demo, and then the paper and output followed. The model was the protagonist and the output was evidence in the background. Today that order ran the other way. The evidence came out first and the protagonist does not have a name yet. When the order of announcement changes, the burden of cost changes too. The cost of verifying capability originally sat on the side preparing the release. Once output starts circulating first, that cost moves to the side accepting the output. The remaining stories in today’s digest are the scene of exactly that move.
Let me go one step more concrete. Two things come attached to the output of 722 proofs. A standard for checking the proofs, and a place to leave that check. If the side that made the output bears both at the same time, it is not a big problem. But once the output starts circulating toward the user side first, the checking cost moves with it. At that point, the ledger becomes the subject.
The flip, as a diagram:
flowchart LR
subgraph OLD["Before, model first"]
M1["Model announcement"] --> M2["Capability demo"] --> M3["Paper and output"]
end
subgraph NEW["Now, output first"]
O1["Output arrives first<br/>722 proofs"] --> O2["Verification runs ahead of release"] --> O3["Verification cost<br/>moves to the receiving side"]
end
O3 --> L["Ledger<br/>one line per output"]
L --> L1["Which model"]
L --> L2["Which authority"]
L --> L3["Which cost"]
L --> L4["Which output"]
Proofs out, model in the vault. The core concept of this post, visualized.
The Morning the Outputs Poured In
If you read today’s digest from start to finish, the output side is moving simultaneously all morning long.
Perplexity’s new multimodal decision model pplx-decider-v1.1 reached first place on the Hugging Face Decision Index 0.3 benchmark. The cost is 2 cents per million tokens, $0.02. In an agent workflow, the decision step was originally the most expensive. Now the price has made decisions a consumable. When a decision costs 2 cents, an agent runs decisions more often. When the number of decisions grows, the record of decisions grows too, and the size of that record becomes the operating cost.
Google’s Nano Banana 2.1 is the latest image generation and editing tool. It outperforms the Pro model at 4x lower cost, and is available in AI Studio and Google Cloud, and in the Gemini app. The official rollout is expanding. Quality jumps one step and price drops two, this morning. Images are different from text. Errors are less visible and the surface they are used on is wide. When that kind of output becomes better at 4x lower price, the ledger on the blocking side grows before the ledger on the using side.
Google and Google DeepMind released EmbeddingGemma 2. It is introduced as the first open multimodal embedding model. It can process 58 frames of video, or 5.5 minutes of audio, in a single pass, and targets local search and retrieval workloads. Embeddings are no longer the exclusive product of a remote API. They have become something that can be placed locally. The “local” the receiving side talks about is not a question of performance. What it points to is the boundary. It means multimodal search can run in an environment where data does not go outside.
Mistral Large 4 recorded 38 points on the Artificial Analysis Intelligence Index. The highest score among AI systems built outside the US and China. The headline calls this model “Le Chonk”. The fact that a non-US, non-China model is first on the index is a brake on the unification of supply. It is a morning when institutions thinking about where their data should stay gained one more menu item.
Line the four stories up and a common point emerges. All of them are mornings when the output arrived before the distribution mechanism. The decision became 2 cents, but where to run the decision is not set. The image got better, but the surface to block got wider. The embedding can come down to local, but which policy to run multimodal search under is still each company’s share. The sovereign top-ranked model came onto the menu, and the standard for choosing from the menu is not yet on the menu. The arrival of output is an act of exposing the gaps in the receiving procedure one by one.
The Standards Drawn on the Receiving Side
The receiving side is moving too. But the form is different. The receiving side does not stack up output. What the receiving side draws is rules.
Meta and Sierra announced an open standard for personal AI agent interaction. Walmart and Shopify joined the consortium, and the goal is to unify how digital assistants interact with enterprise systems. Read in reverse, this story is exactly the other half. An agent’s output is already knocking on the company’s door, and the standard is building a handle on that door. The knock comes first and the handle later.
This agent is personal. When an agent moving on behalf of a person reaches out to an enterprise system, the enterprise must take two questions at once. On whose behalf, and how far to allow. The standard ultimately gathers the answers to these two questions in one place. For a company already running agents, this is not a distant ecosystem story. It is a question about the door that will lead to its own system tomorrow. Distribution and commerce platforms being together also means the standard’s target is not developer documentation. What the standard points to is the actual business system.
The currency of receiving is also price. ChatGPT’s US paid subscriber count is 3x that of Claude and Gemini. But less than 5% of US adults pay for an AI assistant. Adoption is wide and wallets are narrow. The 3x gap speaks to market structure. A configuration where one output leads and the rest follow. But if less than 5% of the whole pays, the leading side’s gap is not yet fixed by price. The gap in output was large, and receiving is still at the starting stage.
The third currency is trust. The US Department of Defense parted ways with Anthropic, the developer of Claude, over a clash of safeguards around autonomous weapons and mass surveillance. It put Google and OpenAI in instead. A case where the world’s largest buyer replaced a supplier because the safeguard standard did not fit. It cannot be explained by score. Because it is a question of the control threshold. The standard of the contract shifted to how far to allow that model’s output. It is highly likely the same order will repeat in enterprise agent adoption.
What happens when the three currencies are late at the same time. Output comes in and there is no door. In enterprise agent adoption, this scene is already familiar. Without a standard, each team builds its own connector. Without price, the budget sheet looks like it has no AI cost. Without trust, control items are born only after an accident. The morning the receiving side draws standards is also the morning those three scenes start at once.
The Ledger of a Company Where Output Arrives First
What ledger should a company hold in a world where output arrives first? The question is not the purchase of a model. The core of the problem is whether you can record, block, and route. A line is added to the ledger every time a new output comes in. What is written on the line is fixed. Which model, with which authority, at which cost, produced which output. One more model to manage is born, one more authority to approve grows, and one more cost to track attaches.
An organization without a ledger endures that speed as is. Authority is scattered inside scripts, and cost is first learned from the bill every time. What was executed is the part you find only by digging through terminal history. The more output grows, the larger these three become by multiplication. If the models to manage double, the authorities to approve and the costs to track can grow by that multiplication. Not having a ledger does not mean not controlling. The cost does not disappear. It just comes in all at once later.
Paxis is a product that has made this ledger into infrastructure. It is ThakiCloud’s Agent-Native Cloud and an official product at v1.1 GA. Skills, Tools, Policies, and Audit Logs are first-class resources. First-class resources means the skills and tools, policies and logs that an agent uses are registered and managed as objects of the platform. They are not notes or files. They exist as resources with authority and version. Autonomy from L0 to L3 governs the range of an agent’s movement, and policy gates and audit logs leave behind what was done, with what authority, before and after execution. Execution proceeds in an isolated sandbox, and MCP connectors and the skill market absorb the flow of standards. CostRouter takes model selection per task, and the sovereign on-premises K8s deployment ai-platform keeps data inside the company.
Today’s digest corresponds to this ledger one by one. The 722 proofs that arrived before the model are the reason output demands audit itself. The policy gate fixes before execution which model can perform which task at which autonomy level, and the audit log answers the question of what was executed. A 2-cent decision model, a 4x cheaper image model, open multimodal embeddings, a sovereign top-ranked model. As the model menu widens every week, per-task selection becomes operation rather than judgment. CostRouter mechanizes that operation. The Meta and Sierra standard and the MCP ecosystem are a question of interface, and the connector and skill market absorb that threshold. The moment an agent on behalf of a person and an agent on behalf of a company meet in front of one system, the boundary of authority must become a resource of the platform. The Department of Defense’s supplier replacement shows that trust is switching cost. Replacing a supplier following a model’s score and replacing one yet still having output audited by the same standard are different things. Isolated sandbox and on-premises deployment make the latter possible.
Proofs out, model in the vault. The morning this state becomes a rule rather than an exception. In a world where output arrives first, the preparation left to a company is two things. A ledger that knows which output came out, with which authority, from which model. And a structure designed so that the cost of writing that ledger is smaller than the benefit of the output. In an industry where the release order is flipping, the only thing you can prepare ahead of is the ledger.
References
This article was written by synthesizing the news below.
- HuggingNews, OpenAI Releases 722 Math Proofs from Unreleased Internal Model
- HuggingNews, Mistral Le Chonk Tops Intelligence Index for Non-US and China AI
- HuggingNews, Google Nano Banana 2.1 Beats Pro Model at 4x Lower Cost
- HuggingNews, Google Launches EmbeddingGemma 2 as First Open Multimodal Embedding Model
- HuggingNews, Meta and Sierra Launch Open Standard for Personal AI Agent Interaction
- HuggingNews, ChatGPT’s U.S. Paid Subscribers Triple Those of Claude and Gemini
- HuggingNews, Perplexity New v1.1 Decision Model Tops Hugging Face Index at $0.02 per Million Tokens
- HuggingNews, Pentagon Hires Google and OpenAI to Replace Anthropic AI