Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash: The MoE Frontier Race

Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash: architectural comparison of 600B MoE vs asymmetric 8B/16B routing, DeepSWE v1.1 benchmarks, and inference token economics.

Lucky Yaduvanshi
Lucky Yaduvanshi
Founder & AI Lead
Sep 20, 2026•Updated Sep 24, 2026•9 min read
Independent technical benchmark • Primary data & verified methodology cited below
Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash: The MoE Frontier Race

StepFun has officially entered the frontier AI tier. Step 5 Preview is the lab’s new flagship foundation model engineered for autonomous software engineering, long-horizon tool execution, professional research, and quantitative finance. It is positioned directly against established frontier architectures including GLM-5.3, DeepSeek V4.1 Flash, and Kimi K3.

The underlying engineering specifications break conventional scaling assumptions: 600B total parameters, only 27B active parameters per token (4.5% activation ratio), a 1,048,576-token (1M) context window, and native vision processing. StepFun’s published benchmarks place Step 5 Preview competitive with GLM-5.3 and Kimi K3 across complex reasoning tasks, while DeepSeek V4.1 Flash offers a contrasting architectural approach with 552B total parameters and asymmetric 8B input / 16B output routing.

This dynamic makes the Step 5 Preview vs GLM-5.3 vs DeepSeek V4.1 Flash showdown a critical study in sparse Mixture-of-Experts (MoE) efficiency, developer economics, and autonomous coding reliability. Explore where these contenders rank across verified benchmarks in our AI Model Leaderboard and Full LLM Model Directory.

What Is Step 5 Preview?

Step 5 Preview is StepFun’s new flagship sparse Mixture-of-Experts foundation model. StepFun reports 600B total backbone parameters and 27B active parameters per token, supported by a native 1M-token context window and multimodal vision perception.

The model is explicitly specialized for agentic workflows rather than simple conversational prompting. StepFun targets long-horizon code maintenance, automated GPU kernel optimization, post-training refinement loops, multi-document financial auditing, and autonomous browser interactions. The model is currently accessible via StepFun’s web products and developer API, with full open weights scheduled for release on October 15, 2026.

This positioning reflects a broader shift across the AI landscape: foundation models are no longer judged solely on conversational flair, but on their ability to execute sustained, multi-hour engineering tasks without state corruption or execution drift.

Architectural Deep-Dive: 600B Total with 27B Active Compute

Headlines citing “600B parameters” can mislead developers who equate total parameter count with dense compute requirements. Step 5 Preview is a sparse MoE architecture:

Specification Step 5 Preview GLM-5.3 DeepSeek V4.1 Flash
Total Backbone Parameters 600B 744B 552B
Active Parameters per Token 27B (4.5%) ~40B (5.4%) 8B Input / 16B Output (1.4%–2.9%)
Context Window 1,048,576 tokens (1M) 1,048,576 tokens (1M) 1,048,576 tokens (1M)
Attention Mechanism Sparse Attention + MLA Multi-Head Attention Multi-Head Latent Attention (MLA)
Multimodal Vision Native Input Native Input Native Multimodal
Primary Domain Focus Agents, Coding, Finance Autonomous Software Agents High-Throughput Agentic Coding
Licensing Open Weights (Oct 15, 2026) Commercial API / Open Weights Commercial API / Open Weights

The active parameter footprint dictates real-world inference throughput and memory bandwidth. For every token generated, Step 5 Preview processes only 27 billion parameters through its routed expert layers.

By comparison, DeepSeek V4.1 Flash deploys an asymmetric routing architecture: 8B active parameters during prompt pre-fill (input) and 16B active parameters during token generation (output). This asymmetric design slashes time-to-first-token (TTFT) latency while preserving deep reasoning capacity during code generation.

Benchmark Performance: Calculated Deltas and Head-to-Head Analysis

StepFun’s published technical documentation provides a detailed comparison against GLM-5.3, Kimi K3, and DeepSeek across standardized coding, reasoning, and tool benchmarks:

Benchmark Suite Step 5 Preview GLM-5.3 Kimi K3 DeepSeek V4.1 Flash Calculated Leader Margin
GPQA Diamond 93.5% 91.7% 93.5% 90.9% Step 5 / Kimi lead GLM-5.3 by +1.8 pts (+2.0%)
DeepSWE v1.1 67.7% 66.9% 67.5% 74.2% DeepSeek leads Step 5 by +6.5 pts (+9.6%)
StepCodeBench (avg@4) 49.0% 40.2% 43.9% Not Reported Step 5 leads GLM-5.3 by +8.8 pts (+21.9%)
ProgramBench 80.5% 72.0% 77.8% Not Reported Step 5 leads GLM-5.3 by +8.5 pts (+11.8%)
Terminal-Bench 2.1 85.0% 83.9% 85.0% 90.6% DeepSeek leads Step 5 by +5.6 pts (+6.6%)
Terminal-Bench v4 33.3% 41.9% 12.6% 31.2% GLM-5.3 leads Step 5 by +8.6 pts (+25.8%)
CyberGym 84.7% 84.5% 80.0% 88.1% DeepSeek leads Step 5 by +3.4 pts (+4.0%)
SciCode 58.9% 59.0% 59.5% Not Reported Kimi K3 leads Step 5 by +0.6 pts (+1.0%)
Agents’ Last Exam (ALE) 29.5% 28.6% 27.6% Not Reported Step 5 leads GLM-5.3 by +0.9 pts (+3.1%)

Methodological Disclosure: These figures represent vendor-reported evaluations under specific temperature and sampling configurations. StepFun notes DeepSWE v1.1 utilized the SWE-agent harness at temperature 1.0 and top-p 0.95. For independent verified evaluations, consult our SWE-bench Verified Guide and Humanity’s Last Exam Deep Dive.

Key Benchmark Takeaways

  1. StepCodeBench Repository Mastery: Step 5 Preview achieves 49.0% vs 40.2% for GLM-5.3, representing a +21.9% relative advantage across 553 multi-language open-source repositories.
  2. Terminal Execution Nuance: On Terminal-Bench 2.1, Step 5 Preview achieves a solid 85.0%, but on the updated Terminal-Bench v4, GLM-5.3 posts 41.9% versus Step 5’s 33.3%—a decisive +25.8% lead for GLM-5.3 in complex bash chaining and system diagnostics.
  3. DeepSeek Coding Dominance: DeepSeek V4.1 Flash posts 74.2% on DeepSWE v1.1, outpacing Step 5 Preview’s 67.7% by +9.6% relative, demonstrating the effectiveness of DeepSeek’s specialized coding distillation.

Financial Reasoning: StepFun’s Specialization

Where Step 5 Preview establishes clear leadership is in quantitative finance and corporate valuation, an area where generic coding models often struggle:

Financial Evaluation Suite Step 5 Preview GLM-5.3 DeepSeek V4.1 Flash Advantage Analysis
FrontierFinance 66.4 64.1 63.0 Step 5 leads DeepSeek by +5.4% relative
CorporateValuation 60.6 56.1 57.6 Step 5 leads GLM-5.3 by +8.0% relative
DeepResearch 55.8 53.3 50.2 Step 5 leads DeepSeek by +11.2% relative
FinStepBench LiveSearch 74.5 73.3 76.7 DeepSeek leads Step 5 by +3.0% relative

Step 5 Preview demonstrates an ability to parse multi-hundred-page 10-K filings, calculate non-GAAP reconciliations, and generate automated DCF valuation models with lower hallucination rates than generalist LLMs.

Coding Agent Reliability: 24-Hour Kernel Optimization

StepFun highlights several stress tests testing model persistence:

  • 24-Hour MLA Kernel Optimization: In an automated optimization experiment, Step 5 Preview tuned a custom Triton MLA GPU kernel for 24 continuous hours, achieving 508 TFLOPS of compute throughput. This matched and slightly exceeded Claude Opus 5 (493 TFLOPS, a +3.0% delta).
  • Autonomous Post-Training: In a self-improving training harness, Step 5 Preview distilled and fine-tuned a Qwen3-30B-A3B student model, boosting its AIME24 mathematical reasoning score from 53.3% to 60.0% (+12.6% improvement) while consuming 35% fewer tokens than competing automated agents.
  • Stateful Game Simulation: To prove context stability over extended interaction horizons, Step 5 Preview played Pokémon Red autonomously for 3,000+ continuous decision turns and 6,000,000 interaction tokens without memory collapse or catastrophic forgetting.

Token Economics and Pricing Analysis

According to Artificial Analysis independent monitoring, Step 5 Preview is priced to challenge the frontier intelligence-cost frontier:

Model Input Price ($/1M) Output Price ($/1M) Cached Input ($/1M) Measured Output Latency
Step 5 Preview $1.00 $2.70 $0.05 ~100 tokens/sec (TTFT: 2.96s)
GLM-5.3 $1.00 $2.00 $0.20 ~85 tokens/sec
DeepSeek V4.1 Flash $0.14 $0.28 $0.014 ~120 tokens/sec
Claude Sonnet 5 $2.00 $10.00 $0.20 ~75 tokens/sec

Step 5 Preview’s $1.00 input / $2.70 output pricing delivers frontier intelligence at a 63% discount on input and 73% discount on output compared to Claude Sonnet 5 ($2/$10 promo tier, and even greater savings against standard $3/$15 enterprise pricing).

However, for sheer developer cost efficiency, DeepSeek V4.1 Flash remains unmatched at $0.14/$0.28 per million tokens—roughly one-tenth the cost of Step 5 Preview.

Where Each Model Excels

Step 5 Preview Excels At:

  • Long-Horizon Persistence: Workflows requiring 20+ hours of continuous agent execution without drift.
  • Financial Analysis: Balance sheet reconciliation, DCF modeling, and regulatory document extraction.
  • Deep Code Refactoring: StepCodeBench indicates superior multi-file refactoring across large repos.
  • Upcoming Self-Hosting: Full Apache 2.0 open-weight availability on October 15, 2026.

GLM-5.3 Excels At:

  • Terminal-Based Workflows: Leads Terminal-Bench v4 (41.9% vs 33.3%), making it superior in CLI environments like ZCode and Claude Code.
  • Production Developer Ecosystem: Fully integrated into established developer stacks and Chinese enterprise cloud deployments.
  • Consistent Inference: Stable throughput with lower latency variance across long generation prompts.

DeepSeek V4.1 Flash Excels At:

  • Unbeatable Token Economics: At $0.14/$0.28 per 1M tokens, it is 7x–10x cheaper than Step 5 and GLM-5.3.
  • High-Velocity Coding: 74.2% on DeepSWE v1.1 and 90.6% on Terminal-Bench 2.1 at high output speeds (120 tps).
  • Asymmetric Memory Efficiency: Asymmetric 8B/16B routing enables massive concurrent request scaling on private clusters.

What Remains Uncertain

  1. Independent Eval Replication: StepFun’s published scores were generated using StepFun’s proprietary evaluation framework. Neutral third-party benchmarks on identical evaluation seeds are needed.
  2. Open-Weight Hardware Demands: While active parameters are 27B, hosting a 600B MoE model locally on October 15 will require multi-node GPU clusters (minimum 8x 80GB GPUs in FP8) to store full model weights.
  3. Global Network Latency: Developers accessing StepFun API endpoints outside East Asia may experience higher network round-trip latency compared to globally distributed CDNs.

Final Verdict & Developer Recommendations

Step 5 Preview is a legitimate frontier contender, not marketing hype. By achieving 67.7% on DeepSWE v1.1, 49.0% on StepCodeBench, and superior financial reasoning at $1.00/$2.70 per million tokens, StepFun proves that specialized post-training and sparse 27B active compute can match dense multi-trillion parameter systems.

  • For CLI Coding & Terminal Agents: GLM-5.3 remains the proven choice due to its +25.8% advantage on Terminal-Bench v4.
  • For High-Volume Production Pipelines: DeepSeek V4.1 Flash delivers the strongest price-to-performance ratio across 100M+ monthly token workloads.
  • For Complex Financial Audits & Long-Horizon Agents: Step 5 Preview is an outstanding tool that developers should actively test ahead of its October 15 open-weight release.

Compare real-time speed, cost, and accuracy benchmarks in our Compare Arena and explore full model cards in our AI Model Leaderboard.

Sources, Disclosures & Primary Benchmark Data

Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.

Share Article
Lucky Yaduvanshi

Lucky Yaduvanshi(luckyyaduvanshi.in →)

Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.

Curated Research

Recommended Reading

Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Explore All 38 Guides→