Z.ai's 50% launch discount on GLM-5.3-Flash expired at 16:00 UTC on September 9. Six days later, the model is being metered at list across most of the 22 providers now serving it, and the headline claim that sold it — a 17.2-point jump on the DeepSWE coding benchmark over GLM-5.2 — is still a number Z.ai produced, on a harness Z.ai chose, with no published third-party replication. That is the state of play for the first natively multimodal model in the GLM-5 family: widely deployed, cheap, open-weight, and graded by its own vendor.

GLM-5.3-Flash landed on August 26, 2026, ending a week of speculation about `ox-alpha`, an anonymous model that had been climbing OpenRouter's usage charts since August 20. Quartz reported that Z.ai "confirmed Wednesday that it was behind Ox Alpha," and that "every request to the model during its preview period was served on Chinese-made AI chips." The specs are unusual for the price tier. It is a mixture-of-experts model with 320 billion total parameters and 18 billion active per token, 45 layers, a one-million-token context window, and native text, image and video input. The weights are MIT-licensed on Hugging Face, where the repository has logged roughly 1.99 million downloads in the past month and spawned 109 quantizations and 16 finetunes.

The model card is direct about the pitch. "We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series," Z.ai writes. "With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks." The architectural claim is narrower and more checkable: "For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities." Z.ai pairs that with Manifold-Constrained Hyper-Connections and a 30-trillion-token multimodal pretraining corpus, and claims roughly 3.0x less attention compute and a 4.4x smaller KV cache than the text-only GLM-5.3.

The scorecard

Every comparative figure in the launch post is vendor-run. DeepSWE goes 63.4 versus GLM-5.2's 46.2. AutomationBench v1.0.6 goes 48.8 versus 26.2, a wider gap than the coding number that got the headlines. Toolathlon Verified is 78.4 versus 59.9. Below those, the deltas compress fast: Terminal-Bench 2.1 improves 3.3 points to 84.3, and Humanity's Last Exam with tools moves from 54.7 to 55.3, which is noise. Against Claude Opus 4.8, Z.ai loses NL2Repo outright, 56.3 to 69.7. Its own in-house Z.ai Code Bench v1.0 puts Flash at 29.0 against Opus 4.8's 29.5 — a number the company published and did not lead with.

The footnotes matter more than the bars. DeepSWE was run on the mini-swe-agent harness at temperature 0.95 with a six-hour timeout and 400K context. Terminal-Bench 2.1 was run inside Claude Code 2.1.207. Toolathlon is pass@1 averaged over three runs. These are defensible choices, individually; collectively they mean the comparison set is not reproducible without matching a stack of configuration decisions Z.ai made. Exactly one row in the table was scored by someone else: the model card's own footnote states that for GDPval-AA v2, "Models are evaluated by Artificial Analysis." That row reports 1773 Elo against GLM-5.2's 1504.

Independent coverage remains thin three weeks out. LLM Stats, which published the most detailed teardown of the launch, labeled the entire benchmark table "Not LLM Stats verified" and has not updated it. Price Per Token's model page, drawing on Artificial Analysis and Hugging Face leaderboard data, currently lists GLM-5.3-Flash at an Intelligence score of 41.9, Coding 71.5, and GPQA 91.2, with median throughput of 114 tokens per second and 2.11 seconds to first token. Those are the only non-Z.ai numbers on the board, and they do not test the DeepSWE claim.

Pricing, now that the clock ran out

List is $0.15 per million input tokens, $0.03 cached, $0.50 output. The promotional rate was half that. LLM Stats co-founder Sebastian Crossa flagged the expiry in the launch writeup: "Put a September 9 price check on the calendar. After the promo, list is the rate that matters." As of September 15, Z.AI's own endpoint bills $0.150/$0.500, as do Fireworks, Cloudflare, Together, BaseTen, OpenRouter and seven others. DeepInfra and FlexAI still show $0.075/$0.250. NextBit is the most expensive at $0.177/$0.590. Price Per Token also shows a $0.000 "cheapest provider" entry via Gonka and a derived claim that input pricing has fallen 99.7% in 90 days; that is an aggregation artifact of a free endpoint, not a price cut. The relevant comparison is internal: text-only GLM-5.3 is a different, 753-billion-parameter model at $1.40/$4.40, roughly 10x the input cost.

Why it matters

The release cadence is the story underneath the benchmark story. GLM-5.3 shipped August 18, GLM-5.3-Flash on August 26, DeepSeek V4.1-Flash on September 10, Sakana's Fugu Max and Fugu Ultra v2 on September 11. Open-weight frontier-adjacent models are now arriving faster than any independent evaluator can grade them, which means the vendor's own table is the only scorecard in existence during the window when procurement decisions actually get made. MIT licensing and a 1M context at $0.15 input are real, verifiable facts. A 17.2-point coding improvement measured by the party selling the model is not the same kind of fact, and the gap between those two categories is where the marketing lives.

What to watch

Whether Artificial Analysis or LLM Stats publishes replicated DeepSWE and AutomationBench runs — the GDPval-AA v2 row shows Z.ai will submit to outside scoring when it likes the result. Whether the promo-era $0.075 rate at DeepInfra and FlexAI holds or converges to list. Whether Chinese-accelerator serving, confirmed for the `ox-alpha` preview period, extends to sustained production traffic at a million-token context. And whether Z.ai ships a larger multimodal GLM-5.3, which would make Flash's 18B active parameters a floor rather than the pitch.

“With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.”
— GLM-5.3-Flash model card, Z.ai, published on Hugging Face
320B/18B
Total vs active parameters, 45-layer MoE
63.4 vs 46.2
Self-reported DeepSWE, 5.3-Flash vs 5.2
$0.15 / $0.50
List price per 1M input/output tokens
1M tokens
Context window, text, image and video