The cheapest way to run a frontier-lab model through OpenAI's API is now 20 cents per million input tokens. It was one dollar until 30 July 2026, when OpenAI cut GPT-5.6 Luna by 80 percent in a single move and, in the same post, quietly told the market where it now expects to make money instead.

The numbers are precise and published. In its developer post titled Advancing the price-performance frontier with GPT-5.6, OpenAI set Luna, its fastest and lowest-cost tier, at 0.20 dollars per million input tokens and 1.20 dollars per million output tokens, down from 1.00 and 6.00 — a clean 80 percent off both sides of the meter. Cached input falls to 2 cents per million. GPT-5.6 Terra, the mid tier, dropped 20 percent to 2 dollars and 12 dollars. GPT-5.6 Sol, the flagship, did not move at all: still 5 dollars in, 30 dollars out. Sam Altman announced the change on X as, in his words, major price cuts today.

One correction to the record circulating in aggregator write-ups: this was a July action, not an August one. OpenAI dated the post 30 July 2026, eWeek reported it on 31 July, and the company said the new rates would begin rolling out on AWS later the same day. The 80 percent figure that has been drifting through August news roundups is real; the date attached to it often is not.

OpenAI's stated reason is worth reading closely, because it is an unusual claim. The company says GPT-5.6 Sol, working inside a human-led engineering process, autonomously rewrote and optimised production kernels and designed and ran hundreds of token-generation experiments. That work, OpenAI says, cut the end-to-end cost of serving the model by 20 percent and raised token-generation efficiency by more than 15 percent.

The customer testimonials OpenAI published alongside the cut are vendor-selected and should be read as such, but they are on the record and they are specific. Michele Catasta, President and Head of AI at Replit, said: "GPT-5.6 Luna is the closest we've come to intelligence too cheap to meter. I've never seen a model this affordable be this powerful — it's unlocking use cases for Replit we didn't expect to build for a long time." Sid Pardeshi, co-founder and CTO of Blitzy, gave harder numbers: "Luna moved us from a single structured-output call to a full tool-calling agent loop, increasing prompt-cache reuse from 24% to 90%. Across thousands of production calls, Luna handles 2.2× more context with 8.5× fewer output tokens—at 87% lower cost than GPT-5.4 mini."

The floor is falling faster than the middle

Set Luna against the rest of the board and the shape of the market is clear. DeepSeek's V4-Flash has been serving at roughly 14 cents in and 28 cents out, with the 0731 build moving to about 22 cents and 66 cents under a peak and off-peak schedule from 16 August. Google's Gemini 3.7 Flash carries introductory pricing of 75 cents and 3.75 dollars through the end of 2026 — a cut this newsletter has covered, as it has ByteDance Seed 2.1 Turbo's rates and Databricks routing work. Alibaba is pushing Qwen 3.8-Max out with open weights, which sets a self-hosting price that no API can undercut. Luna at 20 cents does not lead that pack. It does something more consequential: it puts a US frontier lab inside striking distance of Chinese commodity pricing, which is where the pressure was always going to land.

Two things about the arithmetic do not add up in the obvious way, and both matter. First, OpenAI disclosed a 20 percent reduction in serving cost and then cut price by 80 percent. Efficiency gains of that size do not fund a cut of that size. The remainder is a strategic decision — margin surrendered to hold the high-volume tier — not a pass-through of physics. Second, Anthropic is moving the other direction. Claude Opus 5 launched on 24 July at 5 dollars and 25 dollars, unchanged from Opus 4.8's rate, and Claude Sonnet 5 is scheduled to rise from 2 and 10 dollars to 3 and 15 dollars after 31 August 2026. The floor is collapsing while the reasoning tier firms up. That is a barbell, not a general deflation.

For the application layer, cheap tokens are less of a gift than the headline suggests. Falling per-token prices arrive at exactly the moment agentic architectures start consuming far more tokens per task, and the two effects fight each other. Blitzy's own numbers show it: 8.5 times fewer output tokens, but 2.2 times more context per call. Bills fall, then the loop gets longer and they climb back. The durable consequence is not lower spend but elastic scope — document analysis, interaction classification and routine code implementation become economical to run across an entire corpus rather than a sample. That is real. It is also a thin moat, because every competitor's costs fall at the same time.

The margin question is being answered by relocation. OpenAI held Sol's price flat, cut everything beneath it, and launched Fast mode charging double for 2.5 times the speed — then previewed Ultrafast mode on 13 August promising Sol at up to 14 times the rate. Intelligence is being commoditised toward zero. Latency is being priced as the premium good.

Three things to watch. Whether Anthropic's Sonnet 5 increase actually takes effect on 31 August, which would be the clearest test yet of whether reasoning-tier pricing has a floor of its own. Whether DeepSeek's peak and off-peak experiment spreads, turning inference into a utility with a demand curve. And whether OpenAI ever cuts Sol — because as long as the flagship holds at 30 dollars per million output tokens while Luna sits at 1.20, the price war is being fought entirely in the basement.

“GPT-5.6 Luna is the closest we've come to intelligence too cheap to meter. I've never seen a model this affordable be this powerful.”
— Michele Catasta, President and Head of AI, Replit
80%
Luna price cut
$0.20
Per million input tokens
20%
Actual serving-cost improvement
$0.14
DeepSeek V4-Flash, the floor