On 8 July 2026, SpaceXAI — the renamed AI division that grew out of xAI — and Cursor shipped Grok 4.5: the first model the two companies trained together, built in part on Cursor’s own codebase-interaction data, and tuned for coding, agents, legal and financial work. It launched priced competitively against the field. It also launched unavailable in the EU, on day one, with no date given for when that changes. Three weeks earlier, on the other end of the spectrum, Zhipu released the weights of GLM-5.2 under an unrestricted MIT license — a model that independent benchmarks place fourth in the world, first among open-weight models, at roughly a sixth of the cost of comparable closed models. Two launches, three weeks apart, that quietly moved “which model do we use” out of the engineering-preference column and into the same column as any other vendor decision with cost, data-residency, and continuity risk attached.
What most teams still get wrong
Most engineering organizations pick one frontier model the way they once picked one cloud provider: by default, applied everywhere, revisited only when the bill arrives. That worked when the price gap between “good enough” and “best available” was small. It no longer is. A flagship reasoning model now costs roughly double what its own previous generation cost, and roughly ten times what a genuinely capable open-weight alternative costs for the same class of task. Paying flagship prices for routine work — formatting, boilerplate, straightforward refactors — is not a quality choice. It is an unexamined default, and defaults are exactly what a governance review exists to catch.
What the current landscape actually looks like
| Model | Positioning | Price (input/output per million tokens) | Worth knowing |
|---|---|---|---|
| Claude Fable 5 | Flagship reasoning & orchestration | $10 / $50 | Tops independent intelligence benchmarks; roughly double its own predecessor flagship’s cost — reserve it for the handful of tasks where judgment is genuinely scarce |
| Claude Opus-class | Strong general-purpose flagship | $5 / $25 | The mid-point between frontier capability and frontier price |
| GLM-5.2 (Zhipu) | Open-weight, MIT-licensed | $1.40 / $4.40 official; as low as $0.56 / $1.76 via third-party hosts | Ranks #1 among open-weight models and #4 overall on independent benchmarks; weights are downloadable — you can run it inside your own tenant or on-premises |
| Grok 4.5 (SpaceXAI + Cursor) | Coding, agents, legal & finance tasks | $2 / $6 standard tier; $4 / $18 faster tier | Trained in part on Cursor’s own usage data; not available in the EU at launch |
Figures as of 9 July 2026, sourced from vendor pricing pages and independent benchmark trackers — verify before relying on them; this table will be stale within weeks.
Five things worth doing about it
Route by task, not by loyalty. Reserve the most expensive model for the step where judgment is genuinely scarce — architecture decisions, ambiguous debugging, anything you’d hand to your most senior engineer — and push mechanical work to a cheaper tier deliberately, not by accident. We apply this to our own back-office agents: a cost-efficient model by default, a premium model only for the handful of client-facing or judgment-heavy tasks that warrant it, and every run’s cost logged before anyone has to ask what it cost.
Ask where the weights live before you ask how smart the model is. An open-weight model you can self-host changes the data conversation completely — nothing leaves your infrastructure — but running a frontier-scale open model at full precision is a real capital decision, not a laptop hobby: a model in GLM-5.2’s weight class needs on the order of an eight-GPU, high-memory server to run in full. That makes self-hosting a genuine option for a mid-market or enterprise buyer with a real infrastructure decision to make, not a shortcut.
Check availability before you commit a roadmap to a vendor. A model can be well-priced and well-reviewed and still be unusable for you. Grok 4.5 launching without EU availability is this month’s example; it will not be the last. If you operate in a regulated EU sector, “when does this reach our region” is a procurement question now, not a footnote you check after the fact.
Ask what trained the model before you hand it your codebase. A coding model trained in part on a code editor’s aggregated usage data is a genuine capability edge — and a genuine provenance question. Ask any vendor, open or closed, what went into training their model and what your own usage feeds into next. It is a fair question and a answerable one; vendors that can’t or won’t answer it are telling you something.
Treat “unlimited” tiers as a pricing decision, not a feature. Fallback and “auto” modes that once felt free are being metered at flat rates across the industry as usage scales. Budget for the fallback tier explicitly — don’t discover its cost in the invoice.
The honest limitation
Everything in the table above will be partially wrong by the time enough people read this — that is the nature of a market moving this fast. We are not making a case for any one model. We are making a case for treating model selection with the same discipline you would apply to any vendor decision that touches cost, data location, and business continuity: written down, reviewed, and revisited on a schedule — not defaulted to once and forgotten.
That discipline — routing, logging, and reviewing which model touches what, and why — is close to what we help clients build for their own AI and agent deployments, just applied one level up from the compliance workflows we usually write about. If your organization is choosing model vendors the way it chose cloud vendors a decade ago, that gap is worth a look before it becomes a bigger one.
We help regulated and mid-market organizations design governed agent deployments and the AI governance practices that sit around them — including the unglamorous parts, like knowing which model is allowed to touch which data, and why. If that question doesn’t have a confident answer inside your organization yet, a Discovery Workshop is a two-week way to get one.
How many models does your engineering organization actually use today — and who decided that?
