StepFun has released its flagship model, Step 5 Preview, and said the full model weights will be made available on Oct. 15. The model uses a sparse mixture-of-experts, or MoE, architecture with 600 billion total parameters while activating 27 billion parameters per inference. It supports a 1 million-token context window and multimodal image-text input, while its API and Studio access are already fully open.
Artificial Analysis had completed its benchmark run ahead of the launch. In its Intelligence Index, Step 5 Preview scored 44 points, matching Kimi K3 Max. Artificial Analysis also measured the model at roughly 100 tokens per second, with a per-task cost of $0.71, compared with $2 for Kimi K3 Max.
StepFun highlighted the model’s performance on long-running tasks. In one experiment, Step 5 Preview autonomously optimized an H100 GPU kernel for as long as 24 hours by rewriting code, running tests, comparing results, and iterating. After about 22 hours, it reached 508 TFLOPS, versus 493 TFLOPS for Claude Opus 5 in the same test. In another 24-hour experiment, the model designed post-training data on its own and improved Qwen3-30B-A3B’s AIME24 accuracy from 53.3% to 60%, matching Claude Opus 5 while using fewer labeled tokens.
StepFun has formally released its flagship model, Step 5 Preview. The model uses a sparse mixture-of-experts architecture with 600 billion total parameters, while activating 27 billion parameters at a time. It supports a 1 million-token context window and image-text input.
Its API and Studio access are now fully open, and the company said the complete model weights will be released on Oct. 15.
Benchmark and cost data
Artificial Analysis had already completed its evaluation before the launch. Step 5 Preview scored 44 on the Intelligence Index, putting it level with Kimi K3 Max.
Artificial Analysis currently measures the model at about 100 tokens per second, with a cost of $0.71 per task. Kimi K3 Max, by comparison, comes in at $2.
Long-horizon task performance
StepFun put particular emphasis on the model’s ability to handle long-running tasks. In one demonstration, Step 5 Preview was allowed to autonomously optimize an H100 GPU kernel for up to 24 hours. The model rewrote code, ran tests, compared results, and then continued iterating.
After roughly 22 hours, it reached 508 TFLOPS. In the same experiment, Claude Opus 5 recorded 493 TFLOPS.
In another 24-hour test, Step 5 Preview designed post-training data by itself and raised Qwen3-30B-A3B’s accuracy on AIME24 from 53.3% to 60%. That result matched Claude Opus 5 while using fewer labeled tokens.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.