Verified benchmark data85 models rankedOctober 2026

RankLLMs: AI Model Leaderboard & LLM Benchmarks

Pick between two models and see which one actually wins on coding, reasoning, speed, and cost - using dated numbers that link back to their source. No vendor marketing, no sold rankings.

85Models ranked
SourcedEvery number linked
DailyPrice & speed sync
Verified 2026 AI Benchmark Dataset

AI Model Leaderboard

Scored on reasoning (GPQA/MATH), SWE-bench coding, agentic autonomy, inference throughput (tps), and API pricing - sourced from OpenRouter, Artificial Analysis, and models.dev.

Showing 25 models
#1
OpenAI•Proprietary
Overall68.5
Reason96.0
Code74.1
Speed59 tps
Price/M$10.00
Scorecard→
#2
Anthropic•Proprietary
Overall66.8
Reason88.5
Code86.2
Speed42 tps
Price/M$5.00
Scorecard→
#3
Anthropic•Proprietary
Overall64.2
Reason93.7
Code67.4
Speed65 tps
Price/M$10.00
Scorecard→
#4
Anthropic•Proprietary
Overall62.0
Reason67.5
Code84.5
Speed110 tps
Price/M$10.00
Scorecard→
#5
OpenAI•Proprietary
Overall60.5
Reason78.5
Code84.8
Speed126 tps
Price/M$2.00
Scorecard→
#6
Anthropic•Proprietary
Overall58.2
Reason65.8
Code82.2
Speed50 tps
Price/M$0.00
Scorecard→
#7
Anthropic•Proprietary
Overall57.5
Reason64.5
Code81.8
Speed68 tps
Price/M$5.00
Scorecard→
#8
Meta•Proprietary
Overall57.4
Reason63.5
Code75.4
Speed221 tps
Price/M$1.25
Scorecard→
#9
OpenAI•Proprietary
Overall57.2
Reason65.5
Code81.2
Speed89 tps
Price/M$4.00
Scorecard→
#10
Google•Proprietary
Overall56.8
Reason62.5
Code73.7
Speed297 tps
Price/M$0.75
Scorecard→
#11
Xiaomi•Proprietary
Overall56.5
Reason65.2
Code78.4
Speed55 tps
Price/M$0.44
Scorecard→
#12
Moonshot AI•Open Source
Overall56.0
Reason64.2
Code80.5
Speed52 tps
Price/MFree/Open Source
Scorecard→
#13
Qwen / Alibaba•Proprietary
Overall55.8
Reason62.8
Code79.4
Speed39 tps
Price/M$2.00
Scorecard→
#14
Zhipu AI•Open Source
Overall55.8
Reason60.5
Code78.2
Speed61 tps
Price/MFree/Open Source
Scorecard→
#15
xAI•Proprietary
Overall55.6
Reason63.2
Code78.5
Speed59 tps
Price/M$2.00
Scorecard→
#16
Zhipu AI•Open Source
Overall55.4
Reason62.5
Code77.8
Speed57 tps
Price/M$1.40
Scorecard→
#17
Tencent•Open Source
Overall55.2
Reason58.5
Code77.0
Speed60 tps
Price/MFree/Open Source
Scorecard→
#18
OpenAI•Proprietary
Overall55.0
Reason59.2
Code77.4
Speed117 tps
Price/M$2.00
Scorecard→
#19
Anthropic•Proprietary
Overall54.8
Reason59.5
Code76.5
Speed57 tps
Price/M$5.00
Scorecard→
#20
Anthropic•Proprietary
Overall54.8
Reason58.6
Code75.8
Speed77 tps
Price/M$2.00
Scorecard→
#21
DeepSeek•Open Source
Overall54.4
Reason62.5
Code81.0
Speed176 tps
Price/MFree/Open Source
Scorecard→
#22
DeepSeek•Open Source
Overall54.2
Reason58.4
Code80.6
Speed50 tps
Price/M$1.74
Scorecard→
#23
Google•Proprietary
Overall54.2
Reason61.5
Code78.6
Speed169 tps
Price/M$0.75
Scorecard→
#24
Anthropic•Proprietary
Overall54.0
Reason57.2
Code72.4
Speed35 tps
Price/M$5.00
Scorecard→
#25
OpenAI•Proprietary
Overall53.8
Reason58.2
Code72.5
Speed14 tps
Price/M$2.50
Scorecard→
Showing top 25 of 85 scored AI models • Data synced every 6 hoursView All 85 Models on AI Leaderboard

Head-to-head comparison

See exactly which model wins

Pick any two models for a side-by-side score comparison, accuracy delta, and per-million-token cost - updated from the same dataset as the leaderboard.

Full Comparison Hub
Popular Head-to-Head Presets:
#1

GPT-5.6 Sol

OpenAI • Commercial API
Overall61.8
vs
#2

Claude Opus 5

Anthropic • Commercial API
Overall61.0
RankLLMs ScoreGPT-5.6 Sol Leads (+0.8)
100
50
61.8
61.0
GPT-5.6 Sol
Claude Opus 5
Composite IndexScale: 0-100
SWE-bench VerifiedClaude Opus 5 Leads (+0.7)
100
50
58.5
59.2
GPT-5.6 Sol
Claude Opus 5
Software DevSolved Task %
GPQA DiamondClaude Opus 5 Leads (+2.8)
100
50
55.2
58.0
GPT-5.6 Sol
Claude Opus 5
Hard Science & MathDiamond %
Streaming SpeedGPT-5.6 Sol +43 tps faster
85 tps
Tokens / Sec
42 tps
Inference PricingClaude Opus 5 is 7% cheaper
$7.78 / 1M
Cost / 1M Tokens
$7.22 / 1M
Comparing OpenAI's frontier GPT-5.6 Sol against Anthropic's Claude Opus 5.Open Full In-Depth Matchup Analysis

Visual tier showcase

Model tiers at a glance

Grouped from the verified overall index, the same composite behind the leaderboard. Every chip links to its scorecard. See methodology for the formula.

S · Frontier
5 models
A · Elite
13 models
B · Strong
18 models
C · Value
11 models
D · Niche
38 models
D<4538 models

Niche

Legacy, budget, or specialist picks. Judge on the job, never on overall alone.

Show all 26 niche models

Authoritative Research

Latest Benchmark Guides & Reviews

In-depth technical evaluations, CLI agent testing, and foundation model launch analysis.

View All (38)

Knowledge Base

Frequently Asked Questions About LLM Benchmarks

Clear answers to common technical questions about Large Language Model evaluation and API selection.

What is RankLLMs?

RankLLMs is an independent AI benchmark leaderboard and Large Language Model comparison platform. We aggregate data from OpenRouter, Artificial Analysis, and models.dev, normalize it into one dataset, and score it with a published formula - covering coding accuracy, mathematical reasoning, tokens-per-second speed, and real-world API inference costs. The engine is open source.

What is the highest-ranked AI model in 2026?

OpenAI's GPT-6 Astra currently holds the #1 overall position on RankLLMs with a composite score of 68.5, followed by Anthropic's Claude Fable 5.1 (64.2) and Claude Fable 5 (62.0). For open-source and open-weights models, Moonshot AI's Kimi K3 (56.0), Alibaba's Qwen3.8 Max (55.8), and Zhipu AI's GLM-5.3-Flash (55.8) lead the global rankings.

How are LLM benchmark scores measured on RankLLMs?

RankLLMs aggregates standardized evaluation frameworks including multi-file repository coding benchmarks, mathematical reasoning, Code Arena Elo rankings, and agentic tool-use capability, combined with verified inference speed (tokens/sec and Time-To-First-Token) and API token pricing per 1M tokens.

Which LLM is best for autonomous coding and software engineering?

OpenAI's GPT-6 Astra (88.5 Coding) and Anthropic's Claude Fable 5.1 (86.0 Coding) and Claude Fable 5 (84.5 Coding) rank highest among proprietary systems. For open-weights software development, Moonshot AI's Kimi K3 (80.5 Coding), Alibaba's Qwen3.8 Max (79.4 Coding), and Zhipu AI's GLM-5.3-Flash (78.2 Coding) provide near-commercial performance at a fraction of API token costs.

What are the fastest and most affordable open-weights models?

DeepSeek-V4-Flash-0731 ($0.14/M tokens at 176 tokens/sec), GLM-5.3-Flash ($0.19/M tokens), and Poolside Laguna S 2.1 ($0.11/M tokens) offer state-of-the-art inference efficiency for high-throughput enterprise pipelines.

How often is the RankLLMs leaderboard updated?

The RankLLMs leaderboard is continuously updated whenever foundation model providers (OpenAI, Anthropic, Google, DeepSeek, Zhipu AI, Moonshot AI, Alibaba Cloud / Qwen Team, Meta, xAI, ByteDance, MiniMax) release new model checkpoints, benchmark evaluations, or update their public API token pricing.

Transparent AI Benchmarks and Pricing

RankLLMs compares proprietary and open-weights models on coding (SWE-bench), reasoning (GPQA Diamond and MATH), agentic tool use, inference speed, and API price. Data is aggregated from OpenRouter, Artificial Analysis, and models.dev, normalized into one dataset, and scored with a published formula. We do not accept sponsored rankings or paid placements.

Pricing is in USD per million tokens and refreshed on every sync. The engine behind the dataset is open source, so you can reproduce and audit it yourself. Read the full scoring formula in our Evaluation Methodology or review our Editorial Standards.

Published and maintained by Lucky Yaduvanshi • Independent AI Benchmark Research.