AI Agent Harness Is Now a Product. Access Isn't.
OpenAI now sells the AI agent harness as a managed service and Cognition ships SWE-2 only inside Devin. The scarce part is no longer the loop.
The AI agent harness — the loop that keeps a session alive, compacts context, retries after failures, and hands the model its tools — stopped being something you build this week. OpenAI put it behind an API call. Cognition went the other direction and shipped a frontier coding model that has no API at all, only a harness you rent. Two opposite moves, same conclusion: the loop is now a commodity someone else maintains. Which means the thing that actually separates a demo from a working system is no longer the harness, and it was never the model. It’s whether the agent is allowed to touch anything that matters to your business.
OpenAI Now Sells the Loop You Used to Build
On September 10, OpenAI released its Agents API in public beta. The release note describes it plainly: build agents “with a managed Codex harness while OpenAI handles session orchestration, context compaction, and recovery.”
Read that list again, because every item on it is a problem teams have been solving by hand for two years. Sessions that survive across turns. Context that gets compacted instead of overflowing. Recovery when a step fails halfway through. That is most of the unglamorous engineering in any agent project, and it is now a managed service.
Note what the same release note asks you to bring. Durable sessions let you “connect your own tools and MCP servers.” Agents run “in OpenAI-hosted sandboxes or connect a sandbox from your own infrastructure or a supported provider.”
Your tools. Your MCP servers. Optionally your own sandbox. OpenAI has taken the hard generic part and left the specific part exactly where it was — with you. That is not a gap in the product. That is the product’s shape, and it tells you where the remaining work lives.
Cognition Shipped a Coding Model With No Way to Call It
Two days later, Cognition released SWE-2, a coding model post-trained with reinforcement learning on Moonshot’s Kimi K3. Per MarkTechPost’s writeup, it scores 50.0% on FrontierCode 1.1 Main — within a point of Fable 5.1 — at 64% lower cost.
Then the part everyone skipped past: no open weights, no standalone API. SWE-2 runs only inside Devin.
A year ago that would have read as a limitation. In 2026 it reads as a position. Cognition is betting that the model is not separable from the harness it was trained inside — that shipping the weights alone would ship something measurably worse. Whether or not you buy that, the commercial logic is unmistakable. They are not selling a model. They are selling a place where the model works.
Both Moves Point the Same Direction
OpenAI unbundled the harness and sold it. Cognition bundled the harness so tightly it refuses to sell the model without it. Opposite strategies, identical premise: the environment is the unit of value, not the weights.
I wrote a week ago that AI agent isolation is an environment problem — that the room matters more than the model in it. This week moves the argument one step further, and it’s the step worth paying attention to. If the room is now something you can buy from a vendor in an afternoon, the room is not a differentiator either. The differentiator moved again, to the only thing a generic harness can’t hand you: connection to your actual systems.
Sandbox Access Is Not Business Access
Here is where most agent projects quietly die, and it has nothing to do with capability.
A hosted sandbox will let your agent write code, run tests, install packages, and iterate until something works. What it will not do is give that agent authenticated, reversible, accountable write access to the Salesforce org your revenue runs through. The harness ends at the edge of your business. Everything past that edge is credentials, permissions, deploy paths, and the question nobody wants to answer out loud: when this thing changes something in production at 11pm, who can tell me what it did?
That question is the real gate. It’s why teams with genuinely capable agents still end up pasting generated Apex into a sandbox by hand — the capability was never the blocker. The path from “the agent wrote it” to “the org is running it” was. We’ve covered the accountability gap that opens up when nobody can answer that question, and this week’s releases make it more urgent, not less: cheaper, more capable agents produce more changes, and more changes make an unanswerable audit question worse.
What This Means If You Run a CRM
Nothing announced this week gets your AI closer to your CRM. That sounds pessimistic. It’s the opposite — it means the expensive, generic part is being handled by people with more engineers than you, and the part left over is small, specific, and solvable.
What’s left is four things:
An authenticated connection with real write access. Not a read-only integration that summarizes records. Actual write — the ability to deploy metadata and code, not just describe it. If you want the concrete version of this, we walked through what Salesforce MCP write access actually requires.
A place to run that isn’t a laptop. The Agents API will happily connect to a sandbox from your own infrastructure. That’s a sensible default, and it’s the same reason the runtime environment your agent lives in is worth deciding deliberately rather than inheriting from whatever was open on someone’s desk.
A record of what happened. Every action logged, attributable to a person and a session. Not because logs stop mistakes — they don’t — but because they turn a mistake from a mystery into a diff.
A way back. Snapshots before deploys, sandbox-first for Salesforce with tests required. Recovery is what makes speed acceptable to the person who owns the org.
That’s what Sentinel is: a dedicated VM per client, your AI connected to it over MCP with key-based auth, write access controlled to one write key per org at a time so two people’s sessions can’t collide, and every change logged with a snapshot taken before deploys. It doesn’t restrict what you build — it makes what you built visible and recoverable. If you’re deciding where to spend your own engineering time on any of this, the build-versus-buy math is worth an honest hour.
The Question to Ask This Week
Pick your most-used agent setup and ask it to make one real change in your CRM — not draft one, not explain one. Make it.
If the answer is “it can’t reach the org,” you don’t have a model problem or a harness problem. You have an access problem, and this week two of the best-funded companies in the space just confirmed they’re not going to solve it for you. They’ve told you exactly what they expect: bring your own tools, your own MCP servers, your own place to run.
That’s the part worth owning, and it’s a smaller project than the one you were dreading.
Sentinel gives your AI authenticated, logged, recoverable access to your Salesforce or GoHighLevel org — the last mile the harness vendors deliberately leave to you. Pricing is flat per Sentinel and is covered on a short demo call.
KEEP READING
AI Agent Accountability: The Gap Widened This Week
GPT-6 Astra, declarative agent infra, and agents that tampered with their own logs. AI agent accountability — not capability — is now the constraint.
AI Agent Access Control: The Week It Got Real
AI agent access control stopped being theoretical this week: the MCP spec hardened auth, Anthropic disclosed a real breach, and the EU staffed enforcement.
Ready to see what AI can do for your business?
Start a Conversation