Verified AI Benchmarks • Updated September 2026

RankLLMs: AI Model Leaderboard & LLM Benchmarks

RankLLMs is the independent AI model leaderboard. Compare LLMs on real-world coding benchmarks, reasoning accuracy, tokens per second, and API inference pricing.

Live AI Benchmark Charts

Visual performance across overall reasoning, coding capability, speed (tokens/sec), and API token pricing.

Best Overall (RankLLMs Benchmark)

Verified Artificial Analysis benchmark evaluations, latency throughput, and API inference pricing.

Brand Palette:
Anthropic
OpenAI
Moonshot
Zhipu GLM
DeepSeek
Google
Qwen
Meta
Verified 2026 AI Benchmark Dataset

Comprehensive AI Leaderboard Index

Empirically evaluated on formal reasoning, multi-file coding, agentic autonomy, throughput (tps), and API token pricing.

RankModel & ProviderLicenseRankLLMsReasonCodeAgentArenaSpeedPriceContextCompare
#1
OpenAI official logo
OpenAI
Proprietary
68.5
69.888.592.72,250115 tps
$10.00
1MCompare
#2
Anthropic official logo
Anthropic
Proprietary
64.2
68.586.088.22,045112 tps
$14.44
1MCompare
#3Proprietary
62.0
67.584.586.01,932110 tps
$14.44
1MCompare
#4
Anthropic official logo
Anthropic
Proprietary
58.2
65.882.285.41,74350 tps
$0.00
1MCompare
#5
Anthropic official logo
Anthropic
Proprietary
57.5
64.581.865.02,66868 tps
$7.22
1MCompare
#6Proprietary
57.4
63.575.466.91,754145 tps
$1.25
1MCompare
#7Proprietary
57.2
65.581.258.51,73089 tps
$7.78
1.1MCompare
#8Proprietary
56.8
62.578.861.41,545175 tps
$0.75
1MCompare
#9
Moonshot AI official logo
Moonshot AI
Open Source
56.0
64.280.556.41,68277 tps
$4.33
1MCompare
#10
Qwen / Alibaba official logo
Qwen / Alibaba
Open Source
55.8
62.879.456.21,72463 tps
$2.02
1MCompare
#11
Zhipu AI official logo
Zhipu AI
Open Source
55.8
60.578.252.51,77350 tps
$0.19
1MCompare
#12Proprietary
55.6
63.278.558.01,75357 tps
$2.44
500KCompare
#13
Zhipu AI official logo
Zhipu AI
Proprietary
55.4
62.577.851.01,76950 tps
$1.73
1MCompare
#14
Tencent official logo
Tencent
Open Source
55.2
58.577.049.0-60 tps
Free/Local
1MCompare
#15Proprietary
55.0
59.277.452.01,523117 tps
$2.00
1.1MCompare
#16Proprietary
54.8
59.576.583.41,89057 tps
$7.22
1MCompare
#17Proprietary
54.8
58.675.881.21,61877 tps
$2.89
1MCompare
#18Open Source
54.4
62.581.055.0-176 tps
$1.74
1MCompare
#19Open Source
54.2
58.480.651.81,55450 tps
$1.74
1MCompare
#20Proprietary
54.2
61.578.655.01,588169 tps
$0.75
1MCompare
#21Proprietary
54.0
57.272.447.21,61935 tps
$7.22
1MCompare
#22
OpenAI official logo
OpenAI
Proprietary
53.8
58.272.554.61,67414 tps
$3.89
1MCompare
#23Proprietary
53.5
57.059.348.51,631143 tps
$1.25
1MCompare
#24Proprietary
53.2
62.161.554.71,371140 tps
$1.58
1MCompare
#25Proprietary
52.4
58.674.248.81,314151 tps
$3.89
1MCompare
Showing top 25 of 80 empirically verified AI models • Updated live dailyView All 50+ Models on AI Leaderboard
Model Evaluation Engine

Compare Any Two AI Models Side-by-Side

Directly evaluate SWE-bench verified coding accuracy, mathematical reasoning, latency, and calculate precise API inference cost savings for your engineering stack.

Launch Comparison Engine

Popular Head-to-Head LLM Comparisons

In-depth benchmark showdowns comparing frontier closed models against leading open-weight architectures.

Compare All

Latest Benchmark Guides & Reviews

In-depth technical evaluations, CLI agent testing, and foundation model launch analysis.

View All (33)
Knowledge Base

Frequently Asked Questions About LLM Benchmarks

Clear answers to common technical questions about Large Language Model evaluation and API selection.

What is RankLLMs?

RankLLMs is an independent AI benchmark leaderboard and Large Language Model comparison platform. We provide objective, reproducible evaluations of proprietary and open-weights models based on coding accuracy, mathematical reasoning, tokens-per-second speed, and real-world API inference costs.

What is the highest-ranked AI model in 2026?

OpenAI's GPT-5.6 Sol currently holds the #1 overall position on RankLLMs with a composite score of 57.2, closely followed by Anthropic's Claude Opus 5 (56.2) and Claude Mythos Preview (55.9). For open-source and open-weights models, Moonshot AI's Kimi K3 (54.7), Zhipu AI's GLM-5.3 (54.2), and DeepSeek's DeepSeek-V4-Pro-0813 (54.1) lead the global rankings.

How are LLM benchmark scores measured on RankLLMs?

RankLLMs aggregates standardized evaluation frameworks including multi-file repository coding benchmarks, mathematical reasoning, Code Arena Elo rankings, and agentic tool-use capability, combined with verified inference speed (tokens/sec and Time-To-First-Token) and API token pricing per 1M tokens.

Which LLM is best for autonomous coding and software engineering?

OpenAI's GPT-5.6 Sol (50.6 Coding) and Anthropic's Claude Fable 5 (48.7 Coding) and Claude Opus 5 (42.8 Coding) rank highest among proprietary systems. For open-weights software development, Zhipu AI's GLM-5.3 (45.4 Coding), Moonshot AI's Kimi K3 (45.8 Coding), and DeepSeek-V4-Pro-0813 (44.2 Coding) provide near-commercial performance at a fraction of API token costs.

What are the fastest and most affordable open-weights models?

DeepSeek-V4-Flash-0731 ($0.10/M tokens at 176c/s), GLM-5.3-Flash ($0.19/M tokens), Poolside Laguna S 2.1 ($0.11/M tokens), and Meta Muse Spark 1.2 ($0.11/M tokens at 143c/s) offer state-of-the-art inference efficiency for high-throughput enterprise pipelines.

How often is the RankLLMs leaderboard updated?

The RankLLMs leaderboard is continuously updated whenever foundation model providers (OpenAI, Anthropic, Google, DeepSeek, Zhipu AI, Moonshot AI, Alibaba Cloud / Qwen Team, Meta, xAI, ByteDance, MiniMax) release new model checkpoints, benchmark evaluations, or update their public API token pricing.

Subscribe to AI Benchmark Intel

Get weekly AI model benchmark evaluations, LLM speed/cost breakdowns, and exclusive free API credit alerts delivered to your inbox.

RankLLMs - The Open AI Model Leaderboard & Comparison Platform

Welcome to RankLLMs (also searched as rankllm), the definitive independent platform for real-time AI model comparison, benchmark analysis, and Large Language Model performance tracking. Whether you are an AI engineer selecting the optimal LLM API for production, a researcher evaluating frontier reasoning accuracy, or a software developer searching for autonomous CLI coding agents, RankLLMs provides transparent, data-driven evaluations across proprietary and open-weights artificial intelligence models.

Our tracking covers the current frontier - systems like GPT-5.6 Sol, Claude Fable 5, Claude Opus 5, Gemini 3.7 Flash, and Grok 4.6 - alongside leading open-weights families such as DeepSeek-V4, Qwen3.8, Kimi K3, and GLM-5.3. The full catalog of 80+ model scorecards is updated continuously, so the leaderboard reflects the market as it is now, not a launch-day snapshot.

Software Engineering & Coding Accuracy

Evaluating models on multi-file codebase ingestion, terminal command execution, real GitHub issue resolution, and precise syntax-valid patch generation without hallucination. Frontier models such as GPT-5.6 Sol, Claude Fable 5, and GLM-5.3 lead public coding benchmarks with exceptional resolve rates.

Reasoning & Formal Logic (Chain-of-Thought)

Measuring step-by-step chain-of-thought verification, mathematical proof formulation, and scientific reasoning depth across frontier reasoning systems including Claude Mythos Preview (56.8 Reasoning), GPT-5.6 Sol (56.6 Reasoning), and Kimi K3 (53.6 Reasoning).

Throughput, Latency & Generation Speed (c/s)

Measuring real-world Time-To-First-Token (TTFT) and streaming character throughput (characters per second) across cloud endpoints to ensure interactive applications and automated agent pipelines maintain minimal latency—highlighted by Gemini 3.5 Flash (348 c/s) and DeepSeek-V4-Flash-0731 (176 c/s).

API Economics & Blended Token Pricing ($/M)

Tracking standardized blended token pricing per 1 million tokens across models from high-efficiency open weights like DeepSeek-V4-Flash ($0.10/M) and Poolside Laguna S 2.1 ($0.11/M) up to frontier systems like Claude Opus 5 ($7.22/M) and GPT-5.6 Sol ($7.78/M).

Published and maintained by Lucky Yaduvanshi • Independent AI Benchmark Research.