Skip to content

NewsroomAI advancements

OpenAI cut GPT-5.6 prices up to 80% in three weeks. Read the signal.

OpenAI cut GPT-5.6 Luna prices 80% and Terra 20% three weeks after launch, undercutting Anthropic and nearing DeepSeek. What falling inference prices mean for AI budgets.

The TailorAI teamJuly 31, 2026 · 4 min read

OpenAI cut API prices on GPT-5.6 Luna by roughly 80% on July 30 — three weeks after the model launched. Terra dropped about 20% in the same move. The cut puts Luna well below Anthropic's Haiku 4.5 on price and lands it near open models from DeepSeek. Coverage called it the largest price move since the GPT-5 launch and evidence of a frontier price war. Both readings are true, and both miss the operator story: inference pricing now moves mid-quarter, and every AI budget built on launch-day numbers is already out of date.

Key takeaways

  • The numbers: Luna fell from $1/$6 to $0.20/$1.20 per 1M input/output tokens (~80%); Terra from $2.50/$15 to $2/$12 (~20%); flagship Sol held at $5/$30.
  • OpenAI's framing: the company attributes the cut to rewritten GPU kernels and a redesigned speculative-decoding model — an "efficiency dividend", per OpenAI, not a loss leader.
  • Market position: post-cut, Luna undercuts Anthropic's Haiku 4.5 by roughly 5x on input and 4x on output, and sits near open-model pricing from DeepSeek V4 and GLM-5.2.
  • The signal: inference pricing is now a competitive weapon that moves mid-quarter. An ROI model priced at contract time goes stale fast.
  • The move: design for routing flexibility and push usage-based terms, so falling prices reach your unit economics instead of stopping at a vendor's margin.

What OpenAI changed on July 30

Effective July 30, GPT-5.6 Luna — the small, high-volume tier — dropped from $1 per million input tokens and $6 per million output tokens to $0.20 and $1.20. Terra, the mid tier, went from $2.50/$15 to $2/$12. Sol, the flagship, did not move.

OpenAI says the cut is engineering, not subsidy: rewritten GPU production kernels the company credits with roughly 20% lower serving costs, and a redesigned speculative-decoding draft model it says improved token generation by 15% or more. The "efficiency dividend" framing is OpenAI's own — treat the specific numbers as vendor claims. But the direction fits the pattern of the year: serving costs fall, competition passes the savings through, and the price sheet you negotiated against stops being real.

How Luna now compares with Anthropic and DeepSeek

The new numbers leave Luna at roughly one-fifth of Anthropic's Haiku 4.5 price on input and about a quarter on output, and put it within range of open models like DeepSeek V4 and GLM-5.2. When a frontier lab prices its workhorse tier against open-weight alternatives, the message is plain: capability at that tier is becoming a commodity, and price is the lever labs are pulling to hold volume.

For operators this is mostly good news. High-volume workloads — document intake, classification, extraction, routing — run on exactly this tier, and they just got dramatically cheaper. The caveat: a cut of this size, three weeks after launch, means list price carries very little information. What a workload costs in month one tells you almost nothing about month six.

When a frontier lab reprices 80% in three weeks, any ROI model carved in stone at contract time is stale by the first invoice.

Why your AI budget is already stale

If you scoped an automation project against Luna pricing in early July, the same workload now costs about one-fifth as much. That moves three calculations at once:

  • Build-vs-buy math. Workloads that failed cost-benefit analysis at launch pricing may clear it now. Projects shelved on cost grounds deserve a second pass.
  • Vendor margins. SaaS tools running on these APIs just saw their input costs drop. If their per-seat price did not move, the spread widened in their favor — which is negotiating room in yours.
  • Model selection. Any cost comparison made before July 30 is void. A model that lost on price four weeks ago may win today, and the ranking can flip again next quarter.
A high-volume Luna workload budgeted on launch-day pricing costs roughly one-fifth as much today. If your vendor's price did not move, the difference landed in their margin, not yours.

What operators should do now

  1. 01Reprice on every move, not every renewal. Put model pricing on a watch list and re-run cost models when it changes. A quarterly budget cycle is too slow for an input that repriced 80% in three weeks.
  2. 02Design for routing flexibility. Systems built behind a thin abstraction layer can shift workloads between models and tiers as prices move. Systems welded to one endpoint cannot capture a single cut without an engineering project.
  3. 03Push usage-based and pass-through terms. When negotiating with AI vendors, ask how underlying inference cost declines flow to you. Fixed per-seat pricing locked to launch-day economics converts every future cut into their margin.
  4. 04Re-run the kill list. Pull the automation candidates that died on cost in the last two quarters and score them against current prices. Some are now viable.

We covered the launch tiering in GPT-5.6 is here: what Luna, Terra, and Sol change for buyers, and the same pricing pressure from the other side of the market in Claude Opus 5: near-frontier performance at half the price. If your build-vs-buy assumptions were set before July 30, our build-vs-buy analysis is a good place to restart — or book a consult and we will reprice it with you.

Filed underOpenAIDeepSeekAnthropicmodels
Share

Where this lands in our work

Reading is free. So is the first call.

Wondering what this means for your workflow? That's a thirty-minute conversation, not a research project.