TL;DR
- StepFun launched Step 5 Preview on September 20 — a 600B-parameter sparse mixture-of-experts model with 27B active per token, a 1M-token context window, and API pricing of $1.00 per million input tokens and $2.70 per million output tokens. - The 95% cache discount ($0.05 per million cached tokens) is the most enterprise-relevant number: agent loops that re-read the same codebase or document set pay almost nothing on repeated input. - On Finance benchmarks, Step 5 Preview beats both GPT-6 Astra and Claude Opus 5. On hard coding benchmarks, it trails both by six to seven points. - Artificial Analysis independently validates the Intelligence Index score of 44 and the pricing. Every coding benchmark in the launch table is self-reported on StepFun's own harnesses. There is no technical report. - The Hugging Face repo is currently unavailable. Open weights are promised for October 15 with no license named. - One fact the launch coverage missed: Xiaomi's MiMo-V2.6-Pro scores 46 on the same Artificial Analysis index at $0.435 input / $0.87 output per million tokens — cheaper and higher-scored than Step 5 Preview. - StepFun is one of the six Chinese AI firms named in the NSA/FBI/CISA advisory on AI model distillation, published September 8. Step 4, StepFun's previous model, is specifically named as a beneficiary of distilled US model outputs.
The pitch for Step 5 Preview is not complicated. StepFun is not claiming to beat GPT-6 Astra. It is claiming to sit on a better price-for-intelligence line than anything else in the cheaper tier — and for a specific class of enterprise workload, particularly agentic loops over large codebases and finance tasks, that claim is worth examining carefully.
The pricing, honestly
The headline numbers are $1.00 per million input tokens on a cache miss, $2.70 per million output tokens. On an Artificial Analysis index of 44, verified independently, that works out to $0.71 per weighted Intelligence Index task — which Artificial Analysis places among the better-priced models at this capability level.
The number that matters more for agent workloads is the cache discount: $0.05 per million cached input tokens. That is a 95% reduction on repeated input reads. An agent that loads a 1 million-token codebase ten times pays $1.00 for the first read and $0.45 for the next nine — $1.45 total in input, versus $10.00 uncached. For agentic coding loops that run repeatedly over the same context, this makes Step 5 Preview genuinely cheap in a way the headline price alone does not capture.
The output price is where the model's verbosity matters. Artificial Analysis measured Step 5 Preview generating 160 million output tokens during evaluation, against a median of 94 million for comparable reasoning models. The model reasons at length. Output tokens include the reasoning trace. On a workload where reasoning effort is set to High, the output bill can quietly eat through the savings created by the low input price. StepFun's reasoning_effort parameter (low/medium/high) lets you dial this down for simpler tasks — at the cost of reduced performance.
Straightforward token math against the "1/8 of Claude Opus 5's cost" marketing claim produces about 5.9x, not 8x. StepFun may be factoring in verbosity differences on specific tasks, but the methodology is undisclosed. Treat 1/8 as a marketing number.
What the benchmarks actually say
The comparison table StepFun published pits Step 5 Preview at High reasoning effort against GPT-6 Astra and Claude Opus 5 at Max mode. That is not a direct comparison. It is a reasonable commercial framing — High is what most production deployments would use — but it means the benchmark margins should be read with that asymmetry in mind.
On the numbers StepFun published:
DeepSWE v1.1: Step 5 Preview scores 67.7%, against 74.1% for GPT-6 Astra and 74.0% for Claude Opus 5. Roughly level with Kimi K3 (67.5%) and GLM-5.3 (66.9%).
Terminal-Bench v4: Step 5 Preview scores 33.3%. GPT-6 Astra scores 57.9%, Claude Opus 5 scores 52.3%. This is the widest gap in the table and the one that matters most for autonomous terminal agent work.
FrontierFinance: Step 5 Preview scores 66.4%, ahead of GPT-6 Astra (55.0%), Kimi K3 (62.6%), and GLM-5.3 (64.1%). Claude Opus 5 leads at 69.7%. Finance is Step 5 Preview's genuine standout performance, and StepFun has specifically positioned the model for professional finance work. If your primary use case is financial analysis and structured knowledge work, this benchmark differential is meaningful.
GPQA Diamond: Step 5 Preview scores 93.5%, matching Kimi K3 and within one point of Claude Opus 5 (93.2%), below GPT-6 Astra (96.1%).
The Agents' Last Exam (ALE-CLI) result — 29.5% — is interesting. It exceeds Claude Opus 5 (28.6%), Kimi K3 (27.6%), and GLM-5.3 (28.6%), while remaining below GPT-6 Astra (33.3%). This is an agentic evaluation, and Step 5 Preview's score here is the benchmark most directly relevant to its stated positioning as a model for "long-horizon agent tasks."
All of these numbers carry a caveat: they are self-reported on StepFun's own evaluation harnesses. StepCodeBench is a benchmark StepFun developed internally, covering 553 repositories across 33 languages. Models tend to score well on benchmarks their own team designed. There is no technical report, no independent replication, and no peer review. Artificial Analysis has independently validated the Intelligence Index score and pricing. Everything else is vendor-reported until external runs appear.
The number the launch coverage missed
Xiaomi's MiMo-V2.6-Pro is $0.435 per million input tokens and $0.87 per million output tokens. Artificial Analysis gives it an Intelligence Index of 46 — two points higher than Step 5 Preview's 44. It is the top-ranked open-weight model on the Artificial Analysis index and is less expensive than Step 5 Preview on both input and output.
Step 5 Preview's differentiation over MiMo-V2.6-Pro is not price-performance in the abstract. It is the finance positioning, the video input support, the day-one Claude Code integration via "Step Plan," and the 1M-token context window. MiMo-V2.6-Pro is also open-weight and downloadable today — Step 5 Preview's open weights are promised for October 15, license unnamed. If you are evaluating cheap capable API models this week, MiMo-V2.6-Pro is worth running in parallel with any Step 5 Preview pilot.
What is not yet real
The official Hugging Face repo — stepfun-ai/Step-5-Preview-BF16 — is currently unavailable. It was a placeholder at launch and has been removed or gated as of this week. A third-party 1.21TB shard repository exists with the same model name, tagged with an unknown license. Do not treat it as official.
Open weights are promised for October 15 in BF16 format. No license has been named. StepFun's Step-3.5 and Step-3.7-Flash both shipped Apache-2.0, which is a reasonable basis for expectation — but a more restrictive license would significantly change the self-hosting calculus. StepFun's IPO preparation (WSJ reports a Hong Kong listing at up to a $12B valuation) may influence licensing decisions for the flagship model in ways the Flash line was not subject to.
Maximum output tokens is 64,000 — not 1 million. The 1 million figure in the launch headlines refers to the input context window. This distinction matters for any workflow that expects long document generation.
The "Preview" label on pricing means the current rates may not survive the full release.
The one consideration the launch coverage uniformly skipped
StepFun is one of six Chinese AI companies named in a joint advisory from the NSA, FBI, and CISA published on September 8 — two weeks before the Step 5 launch. Advisory AA26-251A accuses StepFun, DeepSeek, Moonshot AI, Alibaba, MiniMax, and Z.AI of running industrial-scale distillation campaigns against US frontier AI models since late 2024. StepFun's Step 4 model is specifically named as a beneficiary of distilled model outputs.
This does not make Step 5 Preview's $1/$2.70 pricing illegitimate. Distillation is a standard technique in AI development, and the dispute is about the method and scale rather than the technique itself. But it does mean that enterprise teams evaluating Step 5 Preview for production deployment are operating in a supply chain that the US government has formally characterized as a national security concern. Vendor selection decisions that were previously engineering choices may now involve legal and compliance review.
The practical decision rule
The API is live, the pricing is independently verified, and the cache discount is real. Wire Step 5 Preview into a non-production pilot this week — particularly for finance or large-context agent workloads where the cache economics apply. The cost to evaluate is trivial.
Do not build a production dependency on it until October 15 brings actual weights with a named license, and until independent benchmarks confirm the self-reported coding performance. Check three things on October 15: whether the weights are live, what the license is, and whether the released config matches the announced 600B/27B sparse MoE specification.
If you are primarily optimizing for price-performance on the open market today, run MiMo-V2.6-Pro in parallel before committing to Step 5 Preview as your cost-optimization choice.


