OpenAI Slashes GPT-5.6 Prices Up to 80%, Bringing Its Budget Model to 20 Cents

OpenAI on July 30 tore up the price list for two of its newest models, cutting the cost of its budget-tier GPT-5.6 Luna by 80% and its mid-tier Terra by 20% — a move that lands just three weeks after the GPT-5.6 family reached general availability and signals that the fiercest fight in artificial intelligence is no longer over raw capability, but over cost.

Luna, the family's cheapest option, now costs $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6. Terra falls to $2 per million input tokens and $12 per million output tokens, from $2.50 and $15. The flagship GPT-5.6 Sol was left untouched at $5 and $30, though OpenAI paired the announcement with a new "Fast" mode for Sol and a reduction in how many usage credits the cheaper models consume inside ChatGPT Work and Codex.

"Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna," the company wrote in a blog post announcing the change. Chief executive Sam Altman promoted the cuts on X as "major price cuts today," saying OpenAI wants to offer "the best price/intelligence tradeoff at every level."

Efficiency, not a loss leader

OpenAI framed the reductions as a dividend from engineering rather than a margin-sacrificing land grab. The savings, it said, flow from optimizations across its training and inference stack, including software and GPU infrastructure. In an unusual detail, the company said GPT-5.6 Sol was put to work optimizing the production GPU kernels that run its own AI workloads, trimming inference costs without degrading model performance. The goal, OpenAI said, is to deliver "more intelligence per dollar" for its most advanced family of models.

That efficiency story matters because it suggests the new prices are durable rather than promotional. It also lets OpenAI answer a complaint that has grown louder among its biggest customers. Altman told an OpenAI customer event in June that AI spending had gone, within months, from a subject that never came up to a dominant concern, with clients pressing the company to make its systems cheaper to run.

Undercutting the open-weight surge

The competitive backdrop is impossible to ignore. Cheap open-weight models from Chinese labs, led by DeepSeek and Moonshot's Kimi, have been eroding the pricing power that frontier labs once took for granted. By setting Luna's input price at 20 cents, OpenAI has moved to undercut DeepSeek on input tokens outright, even though it remains pricier on output. The timing was pointed on both sides: DeepSeek pushed out its V4 model the same day, touting stronger coding and debugging.

The pressure shows up in the usage data. Chinese-origin models have at times accounted for roughly 46% of US enterprise token consumption on the popular routing platform OpenRouter, occasionally edging past US models — a striking figure for an American company that once had the category largely to itself. Enterprises, meanwhile, have grown reluctant to greenlight open-ended AI budgets without a clear line of sight to returns, with some organizations burning through annual AI allocations in a matter of months.

Why it matters

For all the talk of cost cutting, analysts argue the more important effect will be to accelerate how far and fast enterprises deploy AI. Cheaper tokens change the math on which projects are worth building at all.

"For CIOs, the biggest impact is likely to be scaling AI adoption rather than simply cutting costs or lowering AI budgets," said Pareekh Jain, principal analyst at Pareekh Consulting. "Lower prices make it easier to move pilots into production, expand AI across more employees and business processes, and economically deploy more complex agentic workflows that require multiple model calls."

Jain expects most companies to reinvest the savings into heavier usage rather than shrink their budgets — a dynamic he likened to the Jevons paradox, in which efficiency gains tend to increase overall consumption of a resource rather than reduce it. Chandrika Dutt, research director at Avasant, echoed the point, noting that teams will likely use the new economics to build sophisticated agentic workflows that were previously hard to justify.

That reframing captures the strategic bet behind the cuts. Agentic systems, which chain together many model calls to complete a task, multiply token consumption with every step. Slashing per-token costs makes those workflows viable at scale — and every one of them that runs on Luna or Terra is one that does not run on a rival's model. Cheaper tokens are, in effect, a customer-retention strategy dressed as a discount.

There is a catch for the industry's economics. Cheaper tokens can lift volumes even as they compress the revenue per call, straining the finances of labs already spending heavily on compute. Both OpenAI and Anthropic filed confidential listing prospectuses in June, putting their unit economics under sharper investor scrutiny just as the price war intensifies.

What to watch next

The immediate question is whether rivals match the move. Analysts expect the trend to continue. "These price decreases are broadly sustainable in the long run," Jain said. "New chips, better software, and more efficient model designs will keep driving down the cost per token, and heavy competition makes big price hikes risky for any single provider and thus unlikely."

Watch, too, for how Anthropic, Google and Microsoft respond over the coming weeks, whether DeepSeek's V4 forces a second round of cuts, and whether OpenAI eventually extends discounts to the untouched Sol tier. For enterprise buyers, the takeaway is to build flexible architectures that can swap models as prices fall — because on the current trajectory, today's bargain is unlikely to be the last.

"Lower prices make it easier to move pilots into production, expand AI across more employees and business processes, and economically deploy more complex agentic workflows that require multiple model calls."
— Pareekh Jain, Principal Analyst, Pareekh Consulting
$0.20
Luna input / M tokens (from $1)
$1.20
Luna output / M tokens (from $6)
20%
Terra price cut, to $2 / $12
46%
US enterprise tokens from Chinese models