LLM economics
The 119× price spread, and what it means for builders
Per million tokens, the most expensive top-tier model (Claude Fable 5.1 at $11.90 blended) now costs 119 times more than the cheapest (Meta's Muse Spark 1.3 at roughly $0.10). That's not a rounding difference between vendors — that's two different products wearing the same unit label. The September pricing ledger is worth scrolling in full.
My take: stop shopping by the token
A per-token price tells you what text costs. It doesn't tell you what work costs. The number that matters is dollars per completed task, and the ranking completely reshuffles once you compute it:
- A cheap model that needs 40 tool calls and 3 retries loses to an expensive one that gets it done in 8. Meta claims Spark 1.3 cuts ~20% of tool calls and ~25% of tokens versus 1.2 — that's the right axis of improvement, even at the bottom of the market.
- Cache economics can dominate everything. Anthropic's 75% cache-read cut means long agent sessions on Claude got 25–45% cheaper without the model changing at all.
- Off-peak windows are real products now: DeepSeek's off-peak pricing was the value play of the summer, until GPT-6 Luna undercut it at $0.10/$0.50 flat.
The uncomfortable corollary
When the price floor reaches $0.10 per million tokens, "which model is cheapest" stops being a differentiator for anyone. Your moat is the harness: routing, caching, retries, evals, and knowing which 5% of your workload actually needs the $11.90 model. The labs commoditized the model; the margin moved to the plumbing — which is exactly where it always goes.
Source: Local AI Zone — September 2026 AI Model Updates (price ledger)