Skip to content

NewsroomTechnology briefing

AI inference prices hit a 2026 low. Your contracts should notice.

AI inference prices hit a 2026 low of $1.16–1.18 per million tokens as OpenAI cuts and Anthropic locks Sonnet 5 at $2/$10. What operators should renegotiate now.

The TailorAI teamAugust 12, 2026 · 4 min read

Enterprise AI inference prices hit their lowest level of 2026 in early August. Average prices ran US$1.16–1.18 per million tokens from August 6–8, per Jefferies research citing Silicon Data — down from $2.04 on May 31 and $1.45 in late July, as reported by the South China Morning Post. The decline reflects a heated global price war and surging adoption of low-cost Chinese open-source models such as DeepSeek's, and OpenAI has cut prices on its GPT-5.6 series by up to 80%. Within days of the report, Anthropic made Claude Sonnet 5's introductory pricing permanent and cancelled a scheduled September increase. For operators, the message is plain: the cost side of every AI business case you priced this spring is now wrong — in your favor.

Key takeaways

  • Prices fell roughly 43% in ten weeks. Average enterprise inference cost US$1.16–1.18 per million tokens on August 6–8, down from $2.04 on May 31, per Jefferies citing Silicon Data.
  • Anthropic cancelled a planned increase. Claude Sonnet 5's $2/$10 introductory rate is now the standard price; the scheduled September 1 rise to $3/$15 will not happen.
  • Rejected business cases deserve a re-run. Automation projects that missed cost thresholds at spring rates may clear them at August rates.
  • Contracts should pass declines through. Committed-spend deals locked at last quarter's prices are likely overpriced; negotiate pass-through and keep workloads portable.

How fast prices are falling

The data describes a curve, not a dip. Average unit costs stood at $2.04 per million tokens on May 31, $1.45 in late July, and $1.16–1.18 in the first week of August, per Jefferies citing Silicon Data — the lowest level recorded this year. Two forces are driving it. Providers are cutting: OpenAI reduced prices on its GPT-5.6 model series by up to 80%. And low-cost Chinese open-source models are pulling the floor down — DeepSeek's V4-Flash-0731 is priced around US$0.03 per task, per the same report.

Neither force looks spent. Price competition rewards whoever cuts next, and open-source alternatives cap what any vendor can charge for routine work. The practical takeaway is not that prices will always fall. It is that the price you signed last quarter is no longer the market price.

From May 31 to August 8, average enterprise inference prices fell from $2.04 to $1.16–1.18 per million tokens — roughly 43% in ten weeks, per Jefferies research citing Silicon Data.

What Anthropic's reversal tells you

Claude Sonnet 5 launched on June 30 at $2 per million input tokens and $10 per million output tokens, billed as introductory pricing through August 31, with an increase to $3/$15 scheduled for September 1. Around August 10, Anthropic made the introductory rate the standard price. The increase will not happen. Coverage attributed the reversal to competitive pressure, with OpenAI's GPT-5.6 Sol available at comparable pricing.

Read this two ways. Tactically: if your FY2026 budget carried the September increase, delete it — Sonnet 5 workloads stay 33% cheaper than planned. Structurally: that increase was scheduled, and it died only because a competitor priced against it. Introductory rates carry expiry dates. A falling market does not automatically lower your bill — your contract terms and your ability to switch models do.

The token price is a market rate now. Any contract that treats it as a fixed cost is quietly overpaying.

What cheaper inference changes for automation business cases

Cost per task is falling on a curve, and business cases age with it. A proposal your team rejected in February was priced against winter or spring rates. The same workload runs roughly 40% cheaper today on market averages alone — before provider-specific cuts are counted. Document processing, classification, triage, drafting: work that sat just under the ROI line deserves a re-run at August prices.

One caution as you re-price. Unit cost is no longer the main driver of the bill — volume is. Agentic workloads consume far more tokens per task than single prompts, so the bill grows with how much you automate, not with the token rate. Budget on tasks completed and cost per task, and treat the token price as a market input you revisit quarterly.

What to do about your contracts

Four moves for this quarter.

  1. 01Re-price the rejected backlog. Pull every automation proposal that failed on cost in the past year and re-run the numbers at current rates.
  2. 02Audit committed spend. Any agreement that fixes unit price for more than a quarter is a bet against the market. Ask for a pass-through clause or a re-open trigger tied to published rates.
  3. 03Negotiate pass-through, not a discount. A discount off last quarter's price can still trail the market. The contract should track declines, not memorialize old prices.
  4. 04Keep workloads portable. Price floors are not promised, and introductory rates expire. Systems built so a model can be swapped without a rebuild are the ones that actually capture market declines.

Falling prices are the second half of a story we covered when OpenAI cut GPT-5.6 prices in July. And a re-run business case is only as good as its success measures — our note on acceptance criteria, not demos covers how to build one that survives contact with production. If you want a second pair of eyes on an AI contract or a re-priced automation backlog, book a consult.

Filed underOpenAIGoogleAnthropicDeepSeekdata
Share

Where this lands in our work

Reading is free. So is the first call.

Wondering what this means for your workflow? That's a thirty-minute conversation, not a research project.