RankLLMs: AI Model Leaderboard & LLM Benchmarks
Pick between two models and see which one actually wins on coding, reasoning, speed, and cost - using dated numbers that link back to their source. No vendor marketing, no sold rankings.
AI Model Leaderboard
Scored on reasoning (GPQA/MATH), SWE-bench coding, agentic autonomy, inference throughput (tps), and API pricing - sourced from OpenRouter, Artificial Analysis, and models.dev.
| Rank | Model & Provider | Overall Score | Reasoning | Coding | Speed | Cost / 1M | Actions |
|---|---|---|---|---|---|---|---|
| #1 | GPT-6 AstraNEW OpenAI•Proprietary | 68.5 | 96.0 | 74.1 | 59 tps | $10.00 | |
| #2 | Anthropic•Proprietary | 66.8 | 88.5 | 86.2 | 42 tps | $5.00 | |
| #3 | Anthropic•Proprietary | 64.2 | 93.7 | 67.4 | 65 tps | $10.00 | |
| #4 | Anthropic•Proprietary | 62.0 | 67.5 | 84.5 | 110 tps | $10.00 | |
| #5 | GPT-6 SolNEW OpenAI•Proprietary | 60.5 | 78.5 | 84.8 | 126 tps | $2.00 | |
| #6 | Claude Mythos PreviewUNRELEASED Anthropic•Proprietary | 58.2 | 65.8 | 82.2 | 50 tps | $0.00 | |
| #7 | Anthropic•Proprietary | 57.5 | 64.5 | 81.8 | 68 tps | $5.00 | |
| #8 | Meta•Proprietary | 57.4 | 63.5 | 75.4 | 221 tps | $1.25 | |
| #9 | OpenAI•Proprietary | 57.2 | 65.5 | 81.2 | 89 tps | $4.00 | |
| #10 | Google•Proprietary | 56.8 | 62.5 | 73.7 | 297 tps | $0.75 | |
| #11 | Xiaomi•Proprietary | 56.5 | 65.2 | 78.4 | 55 tps | $0.44 | |
| #12 | Moonshot AI•Open Source | 56.0 | 64.2 | 80.5 | 52 tps | Free/Open Source | |
| #13 | Qwen / Alibaba•Proprietary | 55.8 | 62.8 | 79.4 | 39 tps | $2.00 | |
| #14 | Zhipu AI•Open Source | 55.8 | 60.5 | 78.2 | 61 tps | Free/Open Source | |
| #15 | xAI•Proprietary | 55.6 | 63.2 | 78.5 | 59 tps | $2.00 | |
| #16 | GLM-5.3NEW Zhipu AI•Open Source | 55.4 | 62.5 | 77.8 | 57 tps | $1.40 | |
| #17 | Hy4 PreviewNEW Tencent•Open Source | 55.2 | 58.5 | 77.0 | 60 tps | Free/Open Source | |
| #18 | OpenAI•Proprietary | 55.0 | 59.2 | 77.4 | 117 tps | $2.00 | |
| #19 | Anthropic•Proprietary | 54.8 | 59.5 | 76.5 | 57 tps | $5.00 | |
| #20 | Anthropic•Proprietary | 54.8 | 58.6 | 75.8 | 77 tps | $2.00 | |
| #21 | DeepSeek•Open Source | 54.4 | 62.5 | 81.0 | 176 tps | Free/Open Source | |
| #22 | DeepSeek•Open Source | 54.2 | 58.4 | 80.6 | 50 tps | $1.74 | |
| #23 | Google•Proprietary | 54.2 | 61.5 | 78.6 | 169 tps | $0.75 | |
| #24 | Anthropic•Proprietary | 54.0 | 57.2 | 72.4 | 35 tps | $5.00 | |
| #25 | OpenAI•Proprietary | 53.8 | 58.2 | 72.5 | 14 tps | $2.50 |
No models found
Try adjusting your search query or switching filter categories.
Head-to-head comparison
See exactly which model wins
Pick any two models for a side-by-side score comparison, accuracy delta, and per-million-token cost - updated from the same dataset as the leaderboard.
GPT-5.6 Sol
Claude Opus 5
Trending Head-to-Head LLM Showdowns
Visual tier showcase
Model tiers at a glance
Grouped from the verified overall index, the same composite behind the leaderboard. Every chip links to its scorecard. See methodology for the formula.
- S · Frontier
- 5 models
- A · Elite
- 13 models
- B · Strong
- 18 models
- C · Value
- 11 models
- D · Niche
- 38 models
Frontier
The clear break at the top. Judge everything else against these five.
Elite
Usable flagships, including the top open-weights models like Kimi K3.
- Claude Mythos PreviewAnthropicProprietary58.2#6
- Claude Opus 5AnthropicProprietary57.5#7
- Muse Spark 1.3MetaProprietary57.4#8
- GPT-5.6 SolOpenAIProprietary57.2#9
- Gemini 3.8 FlashGoogleProprietary56.8#10
- MiMo-V2.6-ProXiaomiProprietary56.5#11
- Kimi K3Moonshot AIOpen56.0#12
- Qwen3.8 MaxQwen / AlibabaProprietary55.8#13
- GLM-5.3-FlashZhipu AIOpen55.8#14
- Grok 4.6xAIProprietary55.6#15
- GLM-5.3Zhipu AIOpen55.4#16
- Hy4 PreviewTencentOpen55.2#17
- GPT-5.6 TerraOpenAIProprietary55.0#18
Strong
The bulk of capable flagships. Dense here, so never over-split mid-band.
- Claude Opus 4.8AnthropicProprietary54.8#19
- Claude Sonnet 5AnthropicProprietary54.8#20
- DeepSeek-V4-Pro-0813DeepSeekOpen54.4#21
- DeepSeek-V4-Pro-MaxDeepSeekOpen54.2#22
- Gemini 3.7 FlashGoogleProprietary54.2#23
- Claude Opus 4.6AnthropicProprietary54.0#24
- GPT-5.4OpenAIProprietary53.8#25
- Muse Spark 1.2MetaProprietary53.5#26
- Muse Spark 1.1MetaProprietary53.2#27
- Gemini 3.1 ProGoogleProprietary52.4#28
- MiMo-V2.6-FlashXiaomiOpen52.4#29
- Kimi K2.6Moonshot AIOpen52.2#30
- GPT-5.5OpenAIProprietary52.2#31
- DeepSeek-V4-Flash-MaxDeepSeekOpen52.0#32
- DeepSeek-V4-Flash-Vision-ExpDeepSeekOpen51.8#33
- GPT-6 LunaOpenAIProprietary51.2#34
- Grok 4.5xAIProprietary51.2#35
- Qwen3.8-Flash-NextQwen / AlibabaOpen50.5#36
Value
Mid band. Route by coding, agentic, or price rather than overall alone.
- Claude Sonnet 4.6AnthropicProprietary49.8#37
- GPT-5.6 LunaOpenAIProprietary49.8#38
- GLM-5.1Zhipu AIOpen49.8#39
- Gemini 3.6 FlashGoogleProprietary48.5#40
- Qwen3.7 MaxQwen / AlibabaProprietary48.4#41
- GLM-5.2Zhipu AIOpen48.2#42
- Seed 2.1 ProByteDanceProprietary47.0#43
- DeepSeek-V4-Flash-0731DeepSeekOpen46.1#44
- Qwen3.8-27BQwen / AlibabaOpen46.1#45
- GPT-5.5 ProOpenAIProprietary45.5#46
- Muse SparkMetaProprietary45.2#47
Niche
Legacy, budget, or specialist picks. Judge on the job, never on overall alone.
- Claude Opus 4.7AnthropicProprietary44.8#48
- Seed 2.1 TurboByteDanceProprietary44.0#49
- Gemini 3.5 FlashGoogleProprietary43.6#50
- GPT-5.2 ProOpenAIProprietary43.4#51
- Qwen3.7-PlusQwen / AlibabaProprietary43.3#52
- Sakana NamazuSakana AIProprietary43.2#53
- Hy3TencentOpen42.9#54
- GPT-5.2OpenAIProprietary42.0#55
- MiniMax M3MiniMaxOpen41.9#56
- Laguna S 2.1PoolsideOpen41.4#57
- Grok-4 HeavyxAIProprietary41.0#58
- Qwen3.6 PlusQwen / AlibabaProprietary39.9#59
Show all 26 niche models
- Kimi K2.7 CodeMoonshot AIOpen39.6#60
- Seed 2.0 ProByteDanceProprietary39.5#61
- Gemini 3.5 Flash CyberGoogleProprietary39.5#62
- Kimi K2.5Moonshot AIOpen39.4#63
- Inkling-SmallThinking Machines LabOpen39.3#64
- Qwen3.5-397B-A17BQwen / AlibabaOpen39.2#65
- Gemini 3 ProGoogleProprietary38.7#66
- Claude Opus 4.5AnthropicProprietary38.5#67
- MiniMax M2.5MiniMaxOpen37.4#68
- GLM-5Zhipu AIOpen37.4#69
- Step-3.5-FlashStepFunOpen37.3#70
- Gemini 3 FlashGoogleProprietary37.1#71
- GPT-5.3 CodexOpenAIProprietary37.1#72
- Nemotron 3 Ultra (550B A55B)NVIDIAOpen37.0#73
- GPT-5.1 ThinkingOpenAIProprietary36.9#74
- GPT-5.1 InstantOpenAIProprietary36.8#75
- GPT-5.1OpenAIProprietary36.7#76
- DeepSeek-V4-Flash-0423DeepSeekOpen36.6#77
- Kimi K2-Thinking-0905Moonshot AIOpen36.4#78
- MiniMax M2.7MiniMaxOpen36.1#79
- Qwen3.6-27BQwen / AlibabaOpen36.0#80
- Solar Pro 4Upstage AIProprietary36.0#81
- MiMo-V2-ProXiaomiProprietary36.0#82
- MiMo-V2.5XiaomiOpen35.9#83
- LongCat-Flash-Thinking-2601MeituanOpen35.8#84
- GPT-5.1 HighOpenAIProprietary35.7#85
Authoritative Research
Latest Benchmark Guides & Reviews
In-depth technical evaluations, CLI agent testing, and foundation model launch analysis.
Knowledge Base
Frequently Asked Questions About LLM Benchmarks
Clear answers to common technical questions about Large Language Model evaluation and API selection.
What is RankLLMs?
RankLLMs is an independent AI benchmark leaderboard and Large Language Model comparison platform. We aggregate data from OpenRouter, Artificial Analysis, and models.dev, normalize it into one dataset, and score it with a published formula - covering coding accuracy, mathematical reasoning, tokens-per-second speed, and real-world API inference costs. The engine is open source.
What is the highest-ranked AI model in 2026?
OpenAI's GPT-6 Astra currently holds the #1 overall position on RankLLMs with a composite score of 68.5, followed by Anthropic's Claude Fable 5.1 (64.2) and Claude Fable 5 (62.0). For open-source and open-weights models, Moonshot AI's Kimi K3 (56.0), Alibaba's Qwen3.8 Max (55.8), and Zhipu AI's GLM-5.3-Flash (55.8) lead the global rankings.
How are LLM benchmark scores measured on RankLLMs?
RankLLMs aggregates standardized evaluation frameworks including multi-file repository coding benchmarks, mathematical reasoning, Code Arena Elo rankings, and agentic tool-use capability, combined with verified inference speed (tokens/sec and Time-To-First-Token) and API token pricing per 1M tokens.
Which LLM is best for autonomous coding and software engineering?
OpenAI's GPT-6 Astra (88.5 Coding) and Anthropic's Claude Fable 5.1 (86.0 Coding) and Claude Fable 5 (84.5 Coding) rank highest among proprietary systems. For open-weights software development, Moonshot AI's Kimi K3 (80.5 Coding), Alibaba's Qwen3.8 Max (79.4 Coding), and Zhipu AI's GLM-5.3-Flash (78.2 Coding) provide near-commercial performance at a fraction of API token costs.
What are the fastest and most affordable open-weights models?
DeepSeek-V4-Flash-0731 ($0.14/M tokens at 176 tokens/sec), GLM-5.3-Flash ($0.19/M tokens), and Poolside Laguna S 2.1 ($0.11/M tokens) offer state-of-the-art inference efficiency for high-throughput enterprise pipelines.
How often is the RankLLMs leaderboard updated?
The RankLLMs leaderboard is continuously updated whenever foundation model providers (OpenAI, Anthropic, Google, DeepSeek, Zhipu AI, Moonshot AI, Alibaba Cloud / Qwen Team, Meta, xAI, ByteDance, MiniMax) release new model checkpoints, benchmark evaluations, or update their public API token pricing.
Explore RankLLMs
Transparent AI Benchmarks and Pricing
RankLLMs compares proprietary and open-weights models on coding (SWE-bench), reasoning (GPQA Diamond and MATH), agentic tool use, inference speed, and API price. Data is aggregated from OpenRouter, Artificial Analysis, and models.dev, normalized into one dataset, and scored with a published formula. We do not accept sponsored rankings or paid placements.
Pricing is in USD per million tokens and refreshed on every sync. The engine behind the dataset is open source, so you can reproduce and audit it yourself. Read the full scoring formula in our Evaluation Methodology or review our Editorial Standards.



