AI Agent Pricing Moved Three Ways in One Week
AI agent pricing halved, went free, and roughly doubled — all in seven days. What that volatility should change about what you build on.
AI agent pricing did something strange this week: it went down by half, down to zero, and up by roughly double — inside seven days, across three of the largest model providers in the world. If you are trying to build a budget, a business case, or an architecture on top of a model, that is the most useful thing that happened all week, and almost nobody framed it that way.
Here is what actually shipped between August 10 and August 16, and the conclusion I think it forces.
Google halved the price of its agent workhorse
On Thursday, August 14, Google released Gemini 3.7 Flash, which it describes as its most intelligent workhorse model yet for coding and agents. The introductory price is $0.75 per million input tokens and $3.75 per million output tokens — per eWeek’s coverage, exactly half the original cost of Gemini 3.6 Flash.
The benchmark movement is real, not cosmetic. FrontierCode went from 34.4% to 43.6%; DeepSWE v1.1 from 48.6% to 65.3%. Google’s own framing is that the model “better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity,” with improvements in multistep planning and tool calls.
Note the word introductory. The published rates run through the end of 2026, then rise to $1.50 and $7.50 on January 1, 2027. The 50% cut is a promotion with a printed expiration date, which is a detail worth carrying into any spreadsheet you build this month.
Meta gave one away
The same day, Meta released Muse Glimmer, a 30-billion-parameter open-weights model under an Apache 2.0 license. InfoQ’s writeup describes it as optimized for local, on-device agentic execution — it “enables developers to run autonomous agents, complex tool invocation, local coding, and LLM-as-a-judge evaluations directly on consumer GPUs and workstations without depending on cloud APIs.”
The hardware bar is a workstation, not a data center: 24–32 GB of unified memory or VRAM, and roughly 17–20 GB after 4-bit quantization. A Mac with enough memory or a current consumer GPU clears it.
This is the number that should get your attention more than Google’s. One provider cut the metered price in half. Another removed the meter entirely for anyone willing to run the thing themselves.
DeepSeek went the other direction
Then on August 13, DeepSeek shipped V4-Pro-0813, with a launch price near $0.44 per million input tokens — and peak/off-peak billing taking effect August 16, today, that raises it. Artificial Analysis currently lists $1.32 per million input and $3.96 per million output, with a blended rate of about $0.69, and scores the model at 53 on its Intelligence Index — well above the median of 27 for open-weight models of similar size.
So a capable model roughly tripled its input price within 72 hours of shipping.
Three providers. One week. Down 50%, down to free, up ~3x. There is no trend line you can draw through that.
The lesson isn’t “AI is getting cheap”
The comfortable read is that agent intelligence is racing toward zero and you should wait for the price to bottom out. I don’t think that’s what this week showed. What it showed is that the model layer reprices faster than you can plan around it, in both directions, on promotional terms with expiration dates.
That has a practical consequence. If the core of your plan is “we picked a model and the economics work,” your plan has a shelf life measured in weeks. The vendor can halve the price, triple it, or make it moot by open-sourcing a competitor — and did all three this week, to each other.
So build the plan on the part that doesn’t reprice.
What doesn’t reprice
Notice what none of this week’s announcements changed: not one of them made it any easier for an AI to change something inside your business.
Gemini 3.7 Flash got better at multistep planning. Muse Glimmer will run agent loops on the machine under your desk. Neither one has permission to touch your CRM, a path to deploy into it, or any record of what it did. Those are properties of your environment, not of the model — and they cost the same today as they did last Sunday, because nobody is running a promotion on them.
This is the same seam I wrote about in why “shipped work” is the only AI metric that survives a budget review and, from the architecture side, in the difference between an AI coding assistant and an AI developer. Cheaper tokens make the drafting cheaper. The handoff — where a suggestion becomes a deployed change — is where the value has always leaked, and no price cut touches it.
There is a real cost comparison to make here, but it isn’t between two models. It’s between a token bill and what it costs to hire a Salesforce developer. That gap is the one that actually moves a business case, and it barely moved this week.
The bring-your-own-model position
This is the argument for treating the model as a swappable input rather than a foundation. Sentinel connects over MCP, which means the AI is yours and the choice is yours — if Gemini 3.7 Flash is the right economics in September and something else is right in January, that’s a swap, not a migration.
What stays put is the part you’d otherwise have to rebuild each time: a dedicated server your AI works from, key-based access with one write key at a time per org, sandbox-first deploys for Salesforce, a snapshot before each deploy, and a log of what changed and when. Those exist so a change is visible and recoverable — the durable question is what the AI can reach and what you can see afterward, not which model wrote the code.
Being straight about the trade: Sentinel does not prevent your AI from making a change you’ll regret. It is not a gate. It makes the change legible and undoable, which is a different promise and, I’d argue, the honest one.
If you want the mechanics of pointing an AI at a real org, connecting Claude to Salesforce walks the setup, and the overview of what it means for your AI to become your CRM developer is the wider picture.
The one thing to take from this week
Three of the world’s largest AI providers repriced agent-grade intelligence in three incompatible directions inside a single week, and two of those prices carry expiration dates. Treat the model as weather. Build on the ground.
The question that survives all of it is unchanged: when your AI figures out what to do, what is it allowed to touch, and can you see and undo what it did?
Pricing is exactly this: $2,500 one-time onboarding on your first Sentinel, plus $500/month per Sentinel. You can own more than one.
KEEP READING
AI Decision Authority: Reversible Beats Smart
An AI manager recommended firing someone this week. The real test for AI decision authority isn't how smart the model is — it's whether you can undo it.
AI Governance as Code: This Week in AI
AI governance as code arrived this week — EU transparency rules took effect, Red Hat launched asago, and the Army put Agentforce in front of 9.2M people.
Ready to see what AI can do for your business?
Start a Conversation