Anthropic's R&D Automation Index: 26% Isn't the Story
Anthropic's R&D Automation Index says Claude leads 26% of its AI research. The number that makes that possible is the 100% — and your CRM has neither.
Anthropic published its R&D Automation Index last Thursday, and the number that travelled was 26%: the share of the company’s AI research and development work where Claude now “leads” rather than assists. It is a genuinely striking figure, it got the headlines it deserved, and it is the least useful number in the report. Two other numbers sit underneath it, and those are the ones that explain why a frontier lab can hand a quarter of its own R&D to a model and still know what happened. Your business has neither of them, which is the actual reason your AI isn’t leading anything in your CRM.
What Anthropic Actually Published
On September 17 Anthropic released three measurements of the pace of AI development inside its own walls. The stated reason is transparency: “We are reporting these measurements because they give the public, third parties, and governments better visibility into the pace of AI development inside frontier labs.”
The first measurement is the R&D Automation Index, a ladder of how much of a task the model carries. The top rung reached so far is AL4, which Anthropic defines as the level where “AI ‘leads’: it can complete most of the task end-to-end from a high-level prompt, while the human supervises.” Claude is at that level for 26% of Anthropic’s AI R&D work, up from under 1% in February. More than 90% of the work sits at “AI collaborates” or above. Nothing is running fully autonomously.
The second is oversight. Roughly 30,000 agents were doing research and engineering work simultaneously as of August. The third is compute: about 6% of the compute that went to AI R&D was allocated toward safety, and about 12% of the compute that went to AI-driven AI R&D.
Three numbers, published together, deliberately. They are not three separate facts. They are one fact stated three ways.
The 100% Is the Load-Bearing Number
Here is the line that should have been the headline. Of those 30,000 concurrent agents, Anthropic reports that “100% of these agents’ actions pass through the online monitor before they are executed, usually within seconds.” Of roughly a billion decisions, “0.002% of them (about 1 in 47,000) were blocked by this monitor.” Offline, one to two transcripts per thousand get flagged for a human to read.
Sit with the shape of that for a second. Before Anthropic could say a model leads 26% of its research, it had to build something that sees every action every agent takes, evaluates it before execution, blocks a vanishingly small fraction, and samples the rest for review. The 26% is downstream of the 100%. It is not a statement about how good Claude got. It is a statement about what Anthropic built around Claude so that “how good did it get” became a question with a numeric answer.
The blocking rate is the tell. One in 47,000 is not a safety net catching disasters hourly — it is a monitor that almost never fires, which is exactly what a monitor looks like when the thing it watches is working. The value isn’t the blocks. The value is that every single action passed through something that could see it.
Anthropic Graded Its Own Homework, and Said So
These are a vendor’s numbers about a vendor, produced by the vendor, and that is worth naming plainly rather than reporting the 26% as though an auditor signed it. To Anthropic’s credit, the report says this out loud: the safety-compute figures are “deliberately conservative estimates,” and the allocation snapshot is “a snapshot of how capacity happened to be directed in one week, not a fixed allocation.” A company that defines its own index also chooses where the rungs sit.
What makes the self-report more interesting than the usual kind is what landed the next day. On September 18 Anthropic announced it is embedding independent evaluators from Accenture — through its specialist AI business, Faculty — to evaluate and red-team models, run alignment assessments, and test safeguards. Those evaluators get “access comparable to an employee’s,” which means watching models take shape in training rather than reading a summary afterward. Both companies expect to invest at least $1 billion each in that capacity over five years.
Publish your numbers, then pay someone to come inside and check them. You can be skeptical about how independent an evaluator with a commercial relationship really is, and you should be. The structure is still the right one, and it is the structure almost nobody applies to the AI work happening in their own business.
Your Business Can’t Compute Its Own 26%
Try it. What share of the changes made to your CRM last month were led by an AI rather than typed by a person?
Most operations teams cannot answer that, and the reason is not that the number is low. The reason is that there is no denominator. Changes arrive through a half-dozen doors — an admin in Setup, a consultant in a sandbox, a managed package, someone’s Flow, a script a developer ran once — and nothing in the middle records which hands were on which change. You cannot measure automation you cannot distinguish from everything else. That is the accountability gap in its most ordinary form: not a dramatic failure, just an absence of record.
This is also why “our AI drafts things and a human applies them” feels safe and measures nothing. The draft-and-paste pattern launders the AI’s contribution into a human’s login. Every change looks like it came from a person, because on the way in, it did. You have traded the ability to see what your AI is doing for the comfort of not having to. The distinction between an assistant that writes code and a developer that ships it is exactly this, and it shows up in your audit history or it doesn’t.
An audit trail that names the agent is not bureaucratic overhead. It is the denominator. Without it you don’t have a cautious AI policy — you have an unmeasured one.
The Gap Is Not the Model
The model is not your bottleneck, and after last week nobody can pretend otherwise: the same Claude that leads a quarter of Anthropic’s research is available to you today. What is missing is everything around it.
Three things, specifically, and they are the same three Anthropic had to have. A place for the agent to run that isn’t a laptop, so sessions and credentials live somewhere accountable — the runtime argument that keeps turning out to be the real one. Actual write access to the system, not read-only inspection, which is still the line most CRM integrations stop at. And a record with a way back: what changed, when, by whom, and a snapshot from before.
That last one deserves an honest caveat. None of it stops your AI from building something you regret. It is not supposed to. Anthropic’s monitor blocks 1 in 47,000 actions and lets the rest through, and that is the correct ratio for a system built for freedom with visibility rather than permission gates. The point of the record is not prevention. The point is that when something is wrong you can see it, name it, and undo it — which is a different promise, and a keepable one.
This is the same conclusion the agent-harness news pointed at a week ago from the opposite direction. The loop got commoditized, the model got good enough, and what remains scarce is permission and memory.
What to Ask This Week
Run Anthropic’s three questions against your own org, in miniature. What share of last month’s CRM changes did an AI lead? Of the actions your AI took, how many passed through anything that could see them? And if one of them was wrong, how long until you knew, and how do you get back?
If the honest answers are “no idea,” “none,” and “we’d find out from a user,” that’s not an argument for keeping AI away from your CRM. It’s a description of a missing layer — and it is a much smaller project than the one you’re imagining, because you are not building the model, the harness, or the monitor. You are building the place it runs and the record it leaves.
Notice that none of those three questions is about capability. Not one of them asks whether an AI could write the Apex, design the object, or fix the routing rule — because that argument is over, and it has been over for a while. All three ask whether anyone would know. That is the whole gap between a frontier lab publishing 26% and an operations team publishing nothing.
Start with the record. The 26% comes later, and it will be measurable when it does.
Sentinel gives your AI a dedicated place to run and authenticated write access to your Salesforce or GoHighLevel org, with every action logged and snapshots taken before deploys — so what your AI does is visible and recoverable, not invisible and permanent. Sentinel is not affiliated with Anthropic, Salesforce or Accenture. Pricing is flat per Sentinel and is covered on a short demo call.
KEEP READING
AI Agent Audit Trail: Identity Isn't Accountability
Okta shipped agent identity last week. 80.8% of engineers now use agents daily. The AI agent audit trail is the half nobody shipped.
AIforce Replaces the UI. It Doesn't Replace the Build.
AIforce says AI replaces the UI, and Salesforce is right. The half of Dreamforce nobody demoed is where your org actually gets changed.
Ready to see what AI can do for your business?
Start a Conversation