Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash: The MoE Frontier Race
Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash: architectural comparison of 600B MoE vs asymmetric 8B/16B routing, DeepSWE v1.1 benchmarks, and inference token economics.


Synthesizing article benchmarks & model metrics...
StepFun has officially entered the frontier AI tier. Step 5 Preview is the lab’s new flagship foundation model engineered for autonomous software engineering, long-horizon tool execution, professional research, and quantitative finance. It is positioned directly against established frontier architectures including GLM-5.3, DeepSeek V4.1 Flash, and Kimi K3.
The underlying engineering specifications break conventional scaling assumptions: 600B total parameters, only 27B active parameters per token (4.5% activation ratio), a 1,048,576-token (1M) context window, and native vision processing. StepFun’s published benchmarks place Step 5 Preview competitive with GLM-5.3 and Kimi K3 across complex reasoning tasks, while DeepSeek V4.1 Flash offers a contrasting architectural approach with 552B total parameters and asymmetric 8B input / 16B output routing.
This dynamic makes the Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash showdown a critical study in sparse Mixture-of-Experts (MoE) efficiency, developer economics, and autonomous coding reliability. Explore where these contenders rank across verified benchmarks in our AI Model Leaderboard and Full LLM Model Directory.
What Is Step 5 Preview?
Step 5 Preview is StepFun’s new flagship sparse Mixture-of-Experts foundation model. StepFun reports 600B total backbone parameters and 27B active parameters per token, supported by a native 1M-token context window and multimodal vision perception.
The model is explicitly specialized for agentic workflows rather than simple conversational prompting. StepFun targets long-horizon code maintenance, automated GPU kernel optimization, post-training refinement loops, multi-document financial auditing, and autonomous browser interactions. The model is currently accessible via StepFun’s web products and developer API, with full open weights scheduled for release on October 15, 2026.
This positioning reflects a broader shift across the AI landscape: foundation models are no longer judged solely on conversational flair, but on their ability to execute sustained, multi-hour engineering tasks without state corruption or execution drift.
Architectural Deep-Dive: 600B Total with 27B Active Compute
Headlines citing “600B parameters” can mislead developers who equate total parameter count with dense compute requirements. Step 5 Preview is a sparse MoE architecture:
| Specification | Step 5 Preview | GLM-5.3 | DeepSeek V4.1 Flash |
|---|---|---|---|
| Total Backbone Parameters | 600B | 744B | 552B |
| Active Parameters per Token | 27B (4.5%) | ~40B (5.4%) | 8B Input / 16B Output (1.4%–2.9%) |
| Context Window | 1,048,576 tokens (1M) | 1,048,576 tokens (1M) | 1,048,576 tokens (1M) |
| Attention Mechanism | Sparse Attention + MLA | Multi-Head Attention | Multi-Head Latent Attention (MLA) |
| Multimodal Vision | Native Input | Native Input | Native Multimodal |
| Primary Domain Focus | Agents, Coding, Finance | Autonomous Software Agents | High-Throughput Agentic Coding |
| Licensing | Open Weights (Oct 15, 2026) | Commercial API / Open Weights | Commercial API / Open Weights |
The active parameter footprint dictates real-world inference throughput and memory bandwidth. For every token generated, Step 5 Preview processes only 27 billion parameters through its routed expert layers.
By comparison, DeepSeek V4.1 Flash deploys an asymmetric routing architecture: 8B active parameters during prompt pre-fill (input) and 16B active parameters during token generation (output). This asymmetric design slashes time-to-first-token (TTFT) latency while preserving deep reasoning capacity during code generation.
Benchmark Performance: Calculated Deltas and Head-to-Head Analysis
StepFun’s published technical documentation provides a detailed comparison against GLM-5.3, Kimi K3, and DeepSeek across standardized coding, reasoning, and tool benchmarks:
| Benchmark Suite | Step 5 Preview | GLM-5.3 | Kimi K3 | DeepSeek V4.1 Flash | Calculated Leader Margin |
|---|---|---|---|---|---|
| GPQA Diamond | 93.5% | 91.7% | 93.5% | 90.9% | Step 5 / Kimi lead GLM-5.3 by +1.8 pts (+2.0%) |
| DeepSWE v1.1 | 67.7% | 66.9% | 67.5% | 74.2% | DeepSeek leads Step 5 by +6.5 pts (+9.6%) |
| StepCodeBench (avg@4) | 49.0% | 40.2% | 43.9% | Not Reported | Step 5 leads GLM-5.3 by +8.8 pts (+21.9%) |
| ProgramBench | 80.5% | 72.0% | 77.8% | Not Reported | Step 5 leads GLM-5.3 by +8.5 pts (+11.8%) |
| Terminal-Bench 2.1 | 85.0% | 83.9% | 85.0% | 90.6% | DeepSeek leads Step 5 by +5.6 pts (+6.6%) |
| Terminal-Bench v4 | 33.3% | 41.9% | 12.6% | 31.2% | GLM-5.3 leads Step 5 by +8.6 pts (+25.8%) |
| CyberGym | 84.7% | 84.5% | 80.0% | 88.1% | DeepSeek leads Step 5 by +3.4 pts (+4.0%) |
| SciCode | 58.9% | 59.0% | 59.5% | Not Reported | Kimi K3 leads Step 5 by +0.6 pts (+1.0%) |
| Agents’ Last Exam (ALE) | 29.5% | 28.6% | 27.6% | Not Reported | Step 5 leads GLM-5.3 by +0.9 pts (+3.1%) |
Methodological Disclosure: These figures represent vendor-reported evaluations under specific temperature and sampling configurations. StepFun notes DeepSWE v1.1 utilized the SWE-agent harness at temperature 1.0 and top-p 0.95. For independent verified evaluations, consult our SWE-bench Verified Guide and Humanity’s Last Exam Deep Dive.
Key Benchmark Takeaways
- StepCodeBench Repository Mastery: Step 5 Preview achieves 49.0% vs 40.2% for GLM-5.3, representing a +21.9% relative advantage across 553 multi-language open-source repositories.
- Terminal Execution Nuance: On Terminal-Bench 2.1, Step 5 Preview achieves a solid 85.0%, but on the updated Terminal-Bench v4, GLM-5.3 posts 41.9% versus Step 5’s 33.3%—a decisive +25.8% lead for GLM-5.3 in complex bash chaining and system diagnostics.
- DeepSeek Coding Dominance: DeepSeek V4.1 Flash posts 74.2% on DeepSWE v1.1, outpacing Step 5 Preview’s 67.7% by +9.6% relative, demonstrating the effectiveness of DeepSeek’s specialized coding distillation.
Financial Reasoning: StepFun’s Specialization
Where Step 5 Preview establishes clear leadership is in quantitative finance and corporate valuation, an area where generic coding models often struggle:
| Financial Evaluation Suite | Step 5 Preview | GLM-5.3 | DeepSeek V4.1 Flash | Advantage Analysis |
|---|---|---|---|---|
| FrontierFinance | 66.4 | 64.1 | 63.0 | Step 5 leads DeepSeek by +5.4% relative |
| CorporateValuation | 60.6 | 56.1 | 57.6 | Step 5 leads GLM-5.3 by +8.0% relative |
| DeepResearch | 55.8 | 53.3 | 50.2 | Step 5 leads DeepSeek by +11.2% relative |
| FinStepBench LiveSearch | 74.5 | 73.3 | 76.7 | DeepSeek leads Step 5 by +3.0% relative |
Step 5 Preview demonstrates an ability to parse multi-hundred-page 10-K filings, calculate non-GAAP reconciliations, and generate automated DCF valuation models with lower hallucination rates than generalist LLMs.
Coding Agent Reliability: 24-Hour Kernel Optimization
StepFun highlights several stress tests testing model persistence:
- 24-Hour MLA Kernel Optimization: In an automated optimization experiment, Step 5 Preview tuned a custom Triton MLA GPU kernel for 24 continuous hours, achieving 508 TFLOPS of compute throughput. This matched and slightly exceeded Claude Opus 5 (493 TFLOPS, a +3.0% delta).
- Autonomous Post-Training: In a self-improving training harness, Step 5 Preview distilled and fine-tuned a Qwen3-30B-A3B student model, boosting its AIME24 mathematical reasoning score from 53.3% to 60.0% (+12.6% improvement) while consuming 35% fewer tokens than competing automated agents.
- Stateful Game Simulation: To prove context stability over extended interaction horizons, Step 5 Preview played Pokémon Red autonomously for 3,000+ continuous decision turns and 6,000,000 interaction tokens without memory collapse or catastrophic forgetting.
Token Economics and Pricing Analysis
According to Artificial Analysis independent monitoring, Step 5 Preview is priced to challenge the frontier intelligence-cost frontier:
| Model | Input Price ($/1M) | Output Price ($/1M) | Cached Input ($/1M) | Measured Output Latency |
|---|---|---|---|---|
| Step 5 Preview | $1.00 | $2.70 | $0.05 | ~100 tokens/sec (TTFT: 2.96s) |
| GLM-5.3 | $1.00 | $2.00 | $0.20 | ~85 tokens/sec |
| DeepSeek V4.1 Flash | $0.14 | $0.28 | $0.014 | ~120 tokens/sec |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | ~75 tokens/sec |
Step 5 Preview’s $1.00 input / $2.70 output pricing delivers frontier intelligence at a 63% discount on input and 73% discount on output compared to Claude Sonnet 5 ($2/$10 promo tier, and even greater savings against standard $3/$15 enterprise pricing).
However, for sheer developer cost efficiency, DeepSeek V4.1 Flash remains unmatched at $0.14/$0.28 per million tokens—roughly one-tenth the cost of Step 5 Preview.
Where Each Model Excels
Step 5 Preview Excels At:
- Long-Horizon Persistence: Workflows requiring 20+ hours of continuous agent execution without drift.
- Financial Analysis: Balance sheet reconciliation, DCF modeling, and regulatory document extraction.
- Deep Code Refactoring: StepCodeBench indicates superior multi-file refactoring across large repos.
- Upcoming Self-Hosting: Full Apache 2.0 open-weight availability on October 15, 2026.
GLM-5.3 Excels At:
- Terminal-Based Workflows: Leads Terminal-Bench v4 (41.9% vs 33.3%), making it superior in CLI environments like ZCode and Claude Code.
- Production Developer Ecosystem: Fully integrated into established developer stacks and Chinese enterprise cloud deployments.
- Consistent Inference: Stable throughput with lower latency variance across long generation prompts.
DeepSeek V4.1 Flash Excels At:
- Unbeatable Token Economics: At $0.14/$0.28 per 1M tokens, it is 7x–10x cheaper than Step 5 and GLM-5.3.
- High-Velocity Coding: 74.2% on DeepSWE v1.1 and 90.6% on Terminal-Bench 2.1 at high output speeds (120 tps).
- Asymmetric Memory Efficiency: Asymmetric 8B/16B routing enables massive concurrent request scaling on private clusters.
What Remains Uncertain
- Independent Eval Replication: StepFun’s published scores were generated using StepFun’s proprietary evaluation framework. Neutral third-party benchmarks on identical evaluation seeds are needed.
- Open-Weight Hardware Demands: While active parameters are 27B, hosting a 600B MoE model locally on October 15 will require multi-node GPU clusters (minimum 8x 80GB GPUs in FP8) to store full model weights.
- Global Network Latency: Developers accessing StepFun API endpoints outside East Asia may experience higher network round-trip latency compared to globally distributed CDNs.
Final Verdict & Developer Recommendations
Step 5 Preview is a legitimate frontier contender, not marketing hype. By achieving 67.7% on DeepSWE v1.1, 49.0% on StepCodeBench, and superior financial reasoning at $1.00/$2.70 per million tokens, StepFun proves that specialized post-training and sparse 27B active compute can match dense multi-trillion parameter systems.
- For CLI Coding & Terminal Agents: GLM-5.3 remains the proven choice due to its +25.8% advantage on Terminal-Bench v4.
- For High-Volume Production Pipelines: DeepSeek V4.1 Flash delivers the strongest price-to-performance ratio across 100M+ monthly token workloads.
- For Complex Financial Audits & Long-Horizon Agents: Step 5 Preview is an outstanding tool that developers should actively test ahead of its October 15 open-weight release.
Compare real-time speed, cost, and accuracy benchmarks in our Compare Arena and explore full model cards in our AI Model Leaderboard.
Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.
- •StepFun Step 5 Preview Official Architecture Announcement(Primary Source →)
- •DeepSeek V4.1 Flash Release Documentation(Primary Source →)
- •Artificial Analysis Step 5 Preview Throughput and Pricing Data(Primary Source →)

Lucky Yaduvanshi(luckyyaduvanshi.in →)
Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.
Recommended Reading
Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

GLM-5.3-FlashX: 200 Tokens/sec Throughput vs 2.5x Price Analysis
In-depth review of GLM-5.3-FlashX: 200 tokens/second inference throughput, latency benchmarks, and whether the 2.5x pricing premium is justified.
Lucky Yaduvanshi
TypeSafe AI Jev Architectural Deep Dive: System One Decision Models vs Generative LLMs
Why TypeSafe AI's Jev model captured 13% of Vercel AI Gateway teams in 24 hours. Deep dive into RLCD training, 70ms decision latency, and architecture.
Lucky Yaduvanshi
TypeSafe AI Jev Access Guide: Free Trial Endpoints, API Pricing, and System One Architecture
How to access TypeSafe AI's Jev model for free: Vercel AI Gateway promotion, OpenRouter pricing at $0.042/1M tokens, latency benchmarks, and System One design.
Lucky Yaduvanshi
Alibaba Model Studio Token Plan: Pricing Structure and Claude Code Integration
Comprehensive review of the Alibaba Cloud Model Studio Token Plan. We break down Singapore region access, Credits math, 7-day rolling limits, night discounts, Reddit community benchmarks, and Claude Code setup.
Lucky Yaduvanshi