Shanghai-based StepFun unveiled Step 5 Preview this weekend, a 600-billion-parameter sparse mixture-of-experts model that activates just 27 billion parameters per token, ties Moonshot's Kimi K3 flagship on Artificial Analysis's Intelligence Index, and charges $1.00 per million input tokens and $2.70 per million output tokens. That output price is roughly a seventh of what OpenAI currently charges for GPT-5.6 Sol. The API and StepFun's Studio opened the same day; the weights, StepFun says, will follow on October 15.
StepFun's framing leaves little doubt about its target. “Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance,” StepFun said in its launch announcement, calling the release an effort at “Advancing the Pareto Frontier” between capability and cost.
That efficiency claim rests on an unusually sparse architecture. At 27B active out of 600B total, only about 4.5 percent of the model's parameters fire on any given token, which puts per-token serving cost in the band of a far smaller model while the full 600B footprint becomes largely a memory question. Pandaily reported that the model pairs that sparsity with a 92-layer narrow-deep stack, continuing a design lineage that runs from Step-3.5-Flash through Step-3.7-Flash. StepFun skipped a Step 4 line entirely.
Independent numbers back up at least the cost half of the pitch. Artificial Analysis scores Step 5 Preview at 44 on Intelligence Index v4.3.2, ranking it 24th of 200 models, level with Kimi K3 and one point behind Z.ai's GLM-5.3. On Terminal-Bench 4.0 it posts 33.3 percent, well ahead of Kimi K3's roughly 12.6 percent. Output speed measured 99.8 tokens per second, and the cost per Intelligence Index task came in at $0.71 against roughly $2 for Kimi K3. The benchmark firm summed up the model as “amongst the leading models in intelligence and well priced when comparing to other models of similar price. It's also faster than average, however very verbose.”
That caveat has a dollar figure. Step 5 Preview generated 160 million output tokens to complete the index, against a median of 92 million, and the full evaluation cost $922.84. Roughly $432 of that was spent on the model's own verbosity. StepFun's 95 percent cache discount, which drops cached input to about $0.05 per million tokens, is the lever that offsets this for agents that repeatedly reuse a system prompt, a codebase, or a stack of filings.
StepFun's self-reported results are more mixed than its marketing. On DeepSWE v1.1 the company reports 67.7, on its in-house StepCodeBench 49.0, and on ProgramBench 80.5, ahead of Kimi K3 and GLM-5.3 in StepFun's table but behind GPT-6 Astra and Claude Opus 5 on all three. On Terminal-Bench v4 StepFun itself reports 33.3 against 41.9 for GLM-5.3, 52.3 for Claude Opus 5, and 57.9 for GPT-6 Astra. Finance is where the model looks strongest: 66.4 on FrontierFinance versus GPT-6 Astra's 55, and 83.3 on DRACO versus 76.8, though Claude Opus 5 still leads both. StepFun also published a 24-hour autonomous run in which the model tuned an inference kernel on an Nvidia H100 to 508 TFLOPS, edging Claude Opus 5 at 493.
The context window deserves a direct answer, because coverage has been contradictory. A CoinDesk-sourced flash on KuCoin headlined the model at 100,000 tokens while its own body text said 1 million, and a third-party configuration in developer tooling listed 350,000 tokens with a 64,000-token output cap. StepFun's own developer documentation is unambiguous: a 1M-token context window, 1M maximum input, 1M maximum output, with max_tokens defaulting to unlimited. Artificial Analysis also lists 1M. The 100K figure appears to be a headline error.
What is not there yet is the model itself. The Hugging Face repository stepfun-ai/Step-5-Preview-BF16 was created at 03:52 UTC on September 20 and, as of this writing, contains exactly one file, .gitattributes, with no weights, license, model card, or config. As OrcaRouter's Gideon Frost put it, “the model is callable today and downloadable in about three and a half weeks, and those are two different things.”
Why It Matters
Step 5 Preview is the clearest expression yet of a strategy that has defined 2026's Chinese frontier labs: match the aggregate index of a Western flagship, undercut it by an order of magnitude on price, then open the weights. OpenAI's GPT-5.6 Sol lists at $4 and $20 per million tokens on promotional pricing through November 21, and $5 and $30 at list. Anthropic's Claude Opus 5 sits at $5 and $25. A model scoring 44 on the same index at $1 and $2.70 changes the arithmetic for any agent workload that burns hundreds of millions of tokens a month, especially in finance, where StepFun's numbers are strongest and where enterprise budgets are largest.
The company has the capital to sustain the posture. StepFun, founded in April 2023 by former Microsoft vice president Jiang Daxin, raised more than $717 million in a Series B+ in January and a reported $2.5 billion round in May, pushing its valuation to roughly $10 billion with Tencent, Qiming Venture Partners, and Shanghai state-linked funds on the cap table. Jiang's original conviction after ChatGPT launched, as he recalled in a 2025 profile, was “I can do it myself, maybe even better.” Step 5 Preview is his most ambitious test of that claim.
The risk is that aggregate parity is not parity everywhere. Early testers reported wrong tool-name selection and errors past roughly 340,000 tokens of context, and none of StepFun's agentic results have been independently reproduced.
What to Watch
October 15 is the date that settles most of this. Watch whether the BF16 repository actually fills, under what license, and whether a config file confirms the 600B and 27B figures. Watch whether StepFun holds $1 and $2.70 once the Preview suffix drops; Chinese flagship pricing has a history of moving when capacity tightens. And watch for the first independent tool-use reproductions, which will tell agent builders whether they are buying a cheaper frontier model or just a cheaper rate card.
"The model is callable today and downloadable in about three and a half weeks, and those are two different things."— Gideon Frost, OrcaRouter