RankLLMs: AI Model Leaderboard & LLM Benchmarks
RankLLMs is the independent AI model leaderboard. Compare LLMs on real-world coding benchmarks, reasoning accuracy, tokens per second, and API inference pricing.
Live AI Benchmark Charts
Visual performance across overall reasoning, coding capability, speed (tokens/sec), and API token pricing.
Best Overall (RankLLMs Benchmark)
Verified Artificial Analysis benchmark evaluations, latency throughput, and API inference pricing.
Comprehensive AI Leaderboard Index
Empirically evaluated on formal reasoning, multi-file coding, agentic autonomy, throughput (tps), and API token pricing.
| Rank | Model & Provider | License | RankLLMs | Reason | Code | Agent | Arena | Speed | Price | Context | Compare |
|---|---|---|---|---|---|---|---|---|---|---|---|
| #1 | GPT-6 AstraNEW OpenAI | Proprietary | 68.5 | 69.8 | 88.5 | 92.7 | 2,250 | 115 tps | $10.00 | 1M | Compare |
| #2 | Anthropic | Proprietary | 64.2 | 68.5 | 86.0 | 88.2 | 2,045 | 112 tps | $14.44 | 1M | Compare |
| #3 | Anthropic | Proprietary | 62.0 | 67.5 | 84.5 | 86.0 | 1,932 | 110 tps | $14.44 | 1M | Compare |
| #4 | Claude Mythos PreviewUNRELEASED Anthropic | Proprietary | 58.2 | 65.8 | 82.2 | 85.4 | 1,743 | 50 tps | $0.00 | 1M | Compare |
| #5 | Anthropic | Proprietary | 57.5 | 64.5 | 81.8 | 65.0 | 2,668 | 68 tps | $7.22 | 1M | Compare |
| #6 | Meta | Proprietary | 57.4 | 63.5 | 75.4 | 66.9 | 1,754 | 145 tps | $1.25 | 1M | Compare |
| #7 | OpenAI | Proprietary | 57.2 | 65.5 | 81.2 | 58.5 | 1,730 | 89 tps | $7.78 | 1.1M | Compare |
| #8 | Google | Proprietary | 56.8 | 62.5 | 78.8 | 61.4 | 1,545 | 175 tps | $0.75 | 1M | Compare |
| #9 | Moonshot AI | Open Source | 56.0 | 64.2 | 80.5 | 56.4 | 1,682 | 77 tps | $4.33 | 1M | Compare |
| #10 | Qwen / Alibaba | Open Source | 55.8 | 62.8 | 79.4 | 56.2 | 1,724 | 63 tps | $2.02 | 1M | Compare |
| #11 | Zhipu AI | Open Source | 55.8 | 60.5 | 78.2 | 52.5 | 1,773 | 50 tps | $0.19 | 1M | Compare |
| #12 | xAI | Proprietary | 55.6 | 63.2 | 78.5 | 58.0 | 1,753 | 57 tps | $2.44 | 500K | Compare |
| #13 | GLM-5.3NEW Zhipu AI | Proprietary | 55.4 | 62.5 | 77.8 | 51.0 | 1,769 | 50 tps | $1.73 | 1M | Compare |
| #14 | Hy4 PreviewNEW Tencent | Open Source | 55.2 | 58.5 | 77.0 | 49.0 | - | 60 tps | Free/Local | 1M | Compare |
| #15 | OpenAI | Proprietary | 55.0 | 59.2 | 77.4 | 52.0 | 1,523 | 117 tps | $2.00 | 1.1M | Compare |
| #16 | Anthropic | Proprietary | 54.8 | 59.5 | 76.5 | 83.4 | 1,890 | 57 tps | $7.22 | 1M | Compare |
| #17 | Anthropic | Proprietary | 54.8 | 58.6 | 75.8 | 81.2 | 1,618 | 77 tps | $2.89 | 1M | Compare |
| #18 | DeepSeek | Open Source | 54.4 | 62.5 | 81.0 | 55.0 | - | 176 tps | $1.74 | 1M | Compare |
| #19 | DeepSeek | Open Source | 54.2 | 58.4 | 80.6 | 51.8 | 1,554 | 50 tps | $1.74 | 1M | Compare |
| #20 | Google | Proprietary | 54.2 | 61.5 | 78.6 | 55.0 | 1,588 | 169 tps | $0.75 | 1M | Compare |
| #21 | Anthropic | Proprietary | 54.0 | 57.2 | 72.4 | 47.2 | 1,619 | 35 tps | $7.22 | 1M | Compare |
| #22 | OpenAI | Proprietary | 53.8 | 58.2 | 72.5 | 54.6 | 1,674 | 14 tps | $3.89 | 1M | Compare |
| #23 | Meta | Proprietary | 53.5 | 57.0 | 59.3 | 48.5 | 1,631 | 143 tps | $1.25 | 1M | Compare |
| #24 | Meta | Proprietary | 53.2 | 62.1 | 61.5 | 54.7 | 1,371 | 140 tps | $1.58 | 1M | Compare |
| #25 | Google | Proprietary | 52.4 | 58.6 | 74.2 | 48.8 | 1,314 | 151 tps | $3.89 | 1M | Compare |
No models found
Try adjusting your search query or switching filter categories.
Compare Any Two AI Models Side-by-Side
Directly evaluate SWE-bench verified coding accuracy, mathematical reasoning, latency, and calculate precise API inference cost savings for your engineering stack.
Popular Head-to-Head LLM Comparisons
In-depth benchmark showdowns comparing frontier closed models against leading open-weight architectures.
Latest Benchmark Guides & Reviews
In-depth technical evaluations, CLI agent testing, and foundation model launch analysis.
Frequently Asked Questions About LLM Benchmarks
Clear answers to common technical questions about Large Language Model evaluation and API selection.
What is RankLLMs?
RankLLMs is an independent AI benchmark leaderboard and Large Language Model comparison platform. We provide objective, reproducible evaluations of proprietary and open-weights models based on coding accuracy, mathematical reasoning, tokens-per-second speed, and real-world API inference costs.
What is the highest-ranked AI model in 2026?
OpenAI's GPT-5.6 Sol currently holds the #1 overall position on RankLLMs with a composite score of 57.2, closely followed by Anthropic's Claude Opus 5 (56.2) and Claude Mythos Preview (55.9). For open-source and open-weights models, Moonshot AI's Kimi K3 (54.7), Zhipu AI's GLM-5.3 (54.2), and DeepSeek's DeepSeek-V4-Pro-0813 (54.1) lead the global rankings.
How are LLM benchmark scores measured on RankLLMs?
RankLLMs aggregates standardized evaluation frameworks including multi-file repository coding benchmarks, mathematical reasoning, Code Arena Elo rankings, and agentic tool-use capability, combined with verified inference speed (tokens/sec and Time-To-First-Token) and API token pricing per 1M tokens.
Which LLM is best for autonomous coding and software engineering?
OpenAI's GPT-5.6 Sol (50.6 Coding) and Anthropic's Claude Fable 5 (48.7 Coding) and Claude Opus 5 (42.8 Coding) rank highest among proprietary systems. For open-weights software development, Zhipu AI's GLM-5.3 (45.4 Coding), Moonshot AI's Kimi K3 (45.8 Coding), and DeepSeek-V4-Pro-0813 (44.2 Coding) provide near-commercial performance at a fraction of API token costs.
What are the fastest and most affordable open-weights models?
DeepSeek-V4-Flash-0731 ($0.10/M tokens at 176c/s), GLM-5.3-Flash ($0.19/M tokens), Poolside Laguna S 2.1 ($0.11/M tokens), and Meta Muse Spark 1.2 ($0.11/M tokens at 143c/s) offer state-of-the-art inference efficiency for high-throughput enterprise pipelines.
How often is the RankLLMs leaderboard updated?
The RankLLMs leaderboard is continuously updated whenever foundation model providers (OpenAI, Anthropic, Google, DeepSeek, Zhipu AI, Moonshot AI, Alibaba Cloud / Qwen Team, Meta, xAI, ByteDance, MiniMax) release new model checkpoints, benchmark evaluations, or update their public API token pricing.
Subscribe to AI Benchmark Intel
Get weekly AI model benchmark evaluations, LLM speed/cost breakdowns, and exclusive free API credit alerts delivered to your inbox.
Explore RankLLMs
Model releases, benchmark changes, and pricing moves - every item dated.
What SWE-bench Pro, Humanity's Last Exam, and Terminal Bench actually measure.
Where every score comes from, how it is verified, and when it was last checked.
RankLLMs - The Open AI Model Leaderboard & Comparison Platform
Welcome to RankLLMs (also searched as rankllm), the definitive independent platform for real-time AI model comparison, benchmark analysis, and Large Language Model performance tracking. Whether you are an AI engineer selecting the optimal LLM API for production, a researcher evaluating frontier reasoning accuracy, or a software developer searching for autonomous CLI coding agents, RankLLMs provides transparent, data-driven evaluations across proprietary and open-weights artificial intelligence models.
Our tracking covers the current frontier - systems like GPT-5.6 Sol, Claude Fable 5, Claude Opus 5, Gemini 3.7 Flash, and Grok 4.6 - alongside leading open-weights families such as DeepSeek-V4, Qwen3.8, Kimi K3, and GLM-5.3. The full catalog of 80+ model scorecards is updated continuously, so the leaderboard reflects the market as it is now, not a launch-day snapshot.
Software Engineering & Coding Accuracy
Evaluating models on multi-file codebase ingestion, terminal command execution, real GitHub issue resolution, and precise syntax-valid patch generation without hallucination. Frontier models such as GPT-5.6 Sol, Claude Fable 5, and GLM-5.3 lead public coding benchmarks with exceptional resolve rates.
Reasoning & Formal Logic (Chain-of-Thought)
Measuring step-by-step chain-of-thought verification, mathematical proof formulation, and scientific reasoning depth across frontier reasoning systems including Claude Mythos Preview (56.8 Reasoning), GPT-5.6 Sol (56.6 Reasoning), and Kimi K3 (53.6 Reasoning).
Throughput, Latency & Generation Speed (c/s)
Measuring real-world Time-To-First-Token (TTFT) and streaming character throughput (characters per second) across cloud endpoints to ensure interactive applications and automated agent pipelines maintain minimal latency—highlighted by Gemini 3.5 Flash (348 c/s) and DeepSeek-V4-Flash-0731 (176 c/s).
API Economics & Blended Token Pricing ($/M)
Tracking standardized blended token pricing per 1 million tokens across models from high-efficiency open weights like DeepSeek-V4-Flash ($0.10/M) and Poolside Laguna S 2.1 ($0.11/M) up to frontier systems like Claude Opus 5 ($7.22/M) and GPT-5.6 Sol ($7.78/M).
