RankLLMs: AI Model Leaderboard & LLM Benchmarks
RankLLMs is the independent AI model leaderboard. Compare LLMs on real-world coding benchmarks, reasoning accuracy, tokens per second, and API inference pricing.
AI Model Leaderboard
Empirically evaluated on reasoning (GPQA/MATH), SWE-bench coding, agentic autonomy, inference throughput (tps), and API pricing.
| Rank | Model & Provider | Overall Score | Reasoning | Coding | Speed | Cost / 1M | Actions |
|---|---|---|---|---|---|---|---|
| #1 | GPT-6 AstraNEW OpenAI•Proprietary | 68.5 | 96.0 | 74.1 | 59 tps | $10.00 | |
| #2 | Anthropic•Proprietary | 66.8 | 88.5 | 86.2 | 42 tps | $5.00 | |
| #3 | Anthropic•Proprietary | 64.2 | 93.7 | 67.4 | 65 tps | $10.00 | |
| #4 | Anthropic•Proprietary | 62.0 | 67.5 | 84.5 | 110 tps | $10.00 | |
| #5 | GPT-6 SolNEW OpenAI•Proprietary | 60.5 | 78.5 | 84.8 | 126 tps | $2.00 | |
| #6 | Claude Mythos PreviewUNRELEASED Anthropic•Proprietary | 58.2 | 65.8 | 82.2 | 50 tps | $0.00 | |
| #7 | Anthropic•Proprietary | 57.5 | 64.5 | 81.8 | 68 tps | $5.00 | |
| #8 | Meta•Proprietary | 57.4 | 63.5 | 75.4 | 221 tps | $1.25 | |
| #9 | OpenAI•Proprietary | 57.2 | 65.5 | 81.2 | 89 tps | $4.00 | |
| #10 | Google•Proprietary | 56.8 | 62.5 | 73.7 | 297 tps | $0.75 | |
| #11 | Xiaomi•Proprietary | 56.5 | 65.2 | 78.4 | 55 tps | $0.44 | |
| #12 | Moonshot AI•Open Source | 56.0 | 64.2 | 80.5 | 52 tps | Free/Open Source | |
| #13 | Qwen / Alibaba•Proprietary | 55.8 | 62.8 | 79.4 | 39 tps | $2.00 | |
| #14 | Zhipu AI•Open Source | 55.8 | 60.5 | 78.2 | 61 tps | Free/Open Source | |
| #15 | xAI•Proprietary | 55.6 | 63.2 | 78.5 | 59 tps | $2.00 | |
| #16 | GLM-5.3NEW Zhipu AI•Open Source | 55.4 | 62.5 | 77.8 | 57 tps | $1.40 | |
| #17 | Hy4 PreviewNEW Tencent•Open Source | 55.2 | 58.5 | 77.0 | 60 tps | Free/Open Source | |
| #18 | OpenAI•Proprietary | 55.0 | 59.2 | 77.4 | 117 tps | $2.00 | |
| #19 | Anthropic•Proprietary | 54.8 | 59.5 | 76.5 | 57 tps | $5.00 | |
| #20 | Anthropic•Proprietary | 54.8 | 58.6 | 75.8 | 77 tps | $2.00 | |
| #21 | DeepSeek•Open Source | 54.4 | 62.5 | 81.0 | 176 tps | Free/Open Source | |
| #22 | DeepSeek•Open Source | 54.2 | 58.4 | 80.6 | 50 tps | $1.74 | |
| #23 | Google•Proprietary | 54.2 | 61.5 | 78.6 | 169 tps | $0.75 | |
| #24 | Anthropic•Proprietary | 54.0 | 57.2 | 72.4 | 35 tps | $5.00 | |
| #25 | OpenAI•Proprietary | 53.8 | 58.2 | 72.5 | 14 tps | $2.50 |
No models found
Try adjusting your search query or switching filter categories.
Compare Any Two AI Models
Select any two models from our verified dataset for an instant head-to-head spec duel, accuracy delta, and API cost calculation.
GPT-5.6 Sol
Claude Opus 5
RankLLMs Score
SWE-bench Verified
GPQA Diamond
Trending Head-to-Head LLM Showdowns
Latest Benchmark Guides & Reviews
In-depth technical evaluations, CLI agent testing, and foundation model launch analysis.
Frequently Asked Questions About LLM Benchmarks
Clear answers to common technical questions about Large Language Model evaluation and API selection.
What is RankLLMs?
RankLLMs is an independent AI benchmark leaderboard and Large Language Model comparison platform. We provide objective, reproducible evaluations of proprietary and open-weights models based on coding accuracy, mathematical reasoning, tokens-per-second speed, and real-world API inference costs.
What is the highest-ranked AI model in 2026?
OpenAI's GPT-6 Astra currently holds the #1 overall position on RankLLMs with a composite score of 68.5, followed by Anthropic's Claude Fable 5.1 (64.2) and Claude Fable 5 (62.0). For open-source and open-weights models, Moonshot AI's Kimi K3 (56.0), Alibaba's Qwen3.8 Max (55.8), and Zhipu AI's GLM-5.3-Flash (55.8) lead the global rankings.
How are LLM benchmark scores measured on RankLLMs?
RankLLMs aggregates standardized evaluation frameworks including multi-file repository coding benchmarks, mathematical reasoning, Code Arena Elo rankings, and agentic tool-use capability, combined with verified inference speed (tokens/sec and Time-To-First-Token) and API token pricing per 1M tokens.
Which LLM is best for autonomous coding and software engineering?
OpenAI's GPT-6 Astra (88.5 Coding) and Anthropic's Claude Fable 5.1 (86.0 Coding) and Claude Fable 5 (84.5 Coding) rank highest among proprietary systems. For open-weights software development, Moonshot AI's Kimi K3 (80.5 Coding), Alibaba's Qwen3.8 Max (79.4 Coding), and Zhipu AI's GLM-5.3-Flash (78.2 Coding) provide near-commercial performance at a fraction of API token costs.
What are the fastest and most affordable open-weights models?
DeepSeek-V4-Flash-0731 ($0.14/M tokens at 176 tokens/sec), GLM-5.3-Flash ($0.19/M tokens), and Poolside Laguna S 2.1 ($0.11/M tokens) offer state-of-the-art inference efficiency for high-throughput enterprise pipelines.
How often is the RankLLMs leaderboard updated?
The RankLLMs leaderboard is continuously updated whenever foundation model providers (OpenAI, Anthropic, Google, DeepSeek, Zhipu AI, Moonshot AI, Alibaba Cloud / Qwen Team, Meta, xAI, ByteDance, MiniMax) release new model checkpoints, benchmark evaluations, or update their public API token pricing.
RankLLMs - The Open AI Model Leaderboard & Comparison Platform
Welcome to RankLLMs (also searched as rankllm), the definitive independent platform for real-time AI model comparison, benchmark analysis, and Large Language Model performance tracking. Whether you are an AI engineer selecting the optimal LLM API for production, a researcher evaluating frontier reasoning accuracy, or a software developer searching for autonomous CLI coding agents, RankLLMs provides transparent, data-driven evaluations across proprietary and open-weights artificial intelligence models.
Our tracking covers the current frontier - systems like GPT-5.6 Sol, Claude Fable 5, Claude Opus 5, Gemini 3.7 Flash, and Grok 4.6 - alongside leading open-weights families such as DeepSeek-V4, Qwen3.8, Kimi K3, and GLM-5.3. The full catalog of 80+ model scorecards is updated continuously, so the leaderboard reflects the market as it is now, not a launch-day snapshot.
Software Engineering & Coding Accuracy
Evaluating models on multi-file codebase ingestion, terminal command execution, real GitHub issue resolution, and precise syntax-valid patch generation without hallucination. Frontier models such as GPT-6 Astra, Claude Fable 5.1, and Kimi K3 lead public coding benchmarks with exceptional resolve rates.
Reasoning & Formal Logic (Chain-of-Thought)
Measuring step-by-step chain-of-thought verification, mathematical proof formulation, and scientific reasoning depth across frontier reasoning systems including GPT-6 Astra (69.8 Reasoning), Claude Fable 5.1 (68.5 Reasoning), and Claude Fable 5 (67.5 Reasoning).
Throughput, Latency & Generation Speed (c/s)
Measuring real-world Time-To-First-Token (TTFT) and streaming character throughput (characters per second) across cloud endpoints to ensure interactive applications and automated agent pipelines maintain minimal latency - highlighted by Gemini 3.5 Flash (348 c/s) and DeepSeek-V4-Flash-0731 (176 c/s).
API Economics & Blended Token Pricing ($/M)
Tracking standardized blended token pricing per 1 million tokens across models from high-efficiency open weights like DeepSeek-V4-Flash ($0.10/M) and Poolside Laguna S 2.1 ($0.11/M) up to frontier systems like Claude Opus 5 ($7.22/M) and GPT-5.6 Sol ($7.78/M).



