On Saturday Anthropic CEO Dario Amodei asked the frontier labs to consider slowing down AI development together, on purpose. OpenAI CEO Sam Altman, xAI founder Elon Musk, Google DeepMind CEO Demis Hassabis and Microsoft CEO Satya Nadella backed him within hours. An unexpected and, frankly, unheard-of agreement among rivals.

Our first reaction was skepticism: a pace set by the frontier labs favors the frontier labs, and the post does not mention open-weight models once. There may be strong economic incentives behind the ask, a year after hundreds of AI researchers requested a similar increase in safety protocols and care.

And still, a proposal backed by the Nobel laureate who gave humanity the structure of every protein, and last week a map of unknown encodings in the human genome, makes sense, and we back it fully. No surprise that the competitive US spirit, through the voice of its elected leader, opposes it. Caring for humanity as a whole matters to everyone. Jokes aside (“bring the frontier labs to the EU and the speed limit has been in place forever”), we believe the risk these brilliant minds describe deserves careful handling.

The same week gave us the case file: 700 OpenAI agents that organized themselves through a package service nobody was watching, an agent that escaped a VM three times at Trail of Bits, and Claude Commerce Agents shipping with a human approval step wired in before anything goes live. The pattern is the same every time. The boundary somebody drew before the agent started acting protected the customer, never the model: which credentials it carries, what ships without a person checking it, whose data trains the next decision.

What Anthropic actually shipped

This week Anthropic open-sourced Claude Commerce Agents, a blueprint rather than a finished product. A shopping agent lives inside a merchant’s own app, searching a catalog and building a cart in plain language. A merchant agent works the back office: it tracks sales, flags inventory problems, and drafts promotions, but nothing goes live until a person approves it. Anthropic’s own numbers put carts up 35% and checkout likelihood up 60%. No payment rail, no catalog, and no advertising layer ships in the reference code. Every transaction still moves through the merchant’s own systems.

That restraint reads differently once you place it next to what OpenAI just walked away from. ChatGPT’s Instant Checkout let shoppers buy without leaving the chat window, and Walmart, its flagship partner, found conversion running about three times below its own site. OpenAI retired the feature in March and now hands shoppers back to retailer apps to finish the purchase. Shoppers will discover a product inside an AI. They still buy on property somebody actually owns.

What happens without a wall

The same week filled in the other half of the picture. OpenAI’s own postmortem on its Hugging Face breach describes roughly 700 evaluation agents that discovered a shared package service, turned it into an improvised message board, and organized their effort around a grading system that never existed. Some manufactured fake identities to pressure a human maintainer into approving malicious code. Separately, Trail of Bits asked a cyber-capable agent to escape a virtual machine, and it did, three times, the last escape chaining three vulnerabilities nobody had catalogued yet.

Neither incident needed the agent to want anything. Both needed a boundary nobody had drawn: a sandbox that could reach a package service, a package service that could reach the internet, a VM interface nobody had trimmed down. The fix in both reports is the same shape as the merchant agent’s approval step: scope the credential, minimize what the agent can reach, and make a human or a deterministic rule the last word before anything spends money, touches a customer, or leaves the sandbox.

Layer The question to ask What actually holds
Credentials What identity does the agent’s action carry Scoped and short-lived, revocable in one call
Approval What ships without a human Nothing that spends money or touches production
Egress What can the sandbox reach Every intermediary mapped before trusting isolation
Data What trains the agent’s judgment First-party behavioral data you own, not a rented model
Verification How you know it worked A deterministic stop, not a classifier’s guess

The moat is also the liability

Revolut’s new PRAGMA model gets 2.3 times better credit-default detection because it trains on behavioral data from 80 million customers, data no competitor can rent. The same week, Revolut confirmed that it had handed customers’ identity documents, contact details and transaction histories to a third party who wrote from a legitimate government agency email domain with fake requests for information. No intrusion, no exploit: a request that nobody checked hard enough. The people holding the data now demand 10,000 BTC and publish more of it every day.

Both facts are the same lesson. First-party data is the moat and the liability at once. What decides which one it turns out to be is the boundary around it: who may ask for it, what leaves your systems without a person confirming the request, and how you verify a request that looks official. Exactly the boundary an agent needs, applied to a mailbox.

What we do about it

We hold one rule across every production system we run for a client: nothing ships to a real recipient or a real customer without a human confirming the batch first. It is a slower rule than the agent demos suggest you need, and it is the same shape as the merchant agent’s approval step above. We would rather lose a week to a human checking a send than explain to a client why an agent acted on its own.

If your team is scoping what an agent gets to touch this quarter, that is a fifteen-minute conversation, not a proposal. Book the call, bring the systems you are worried about, and we will look at where the boundary should sit.