Live Benchmark Telemetry50+ Frontier ModelsSeptember 2026

RankLLMs: AI Model Leaderboard & LLM Benchmarks

RankLLMs is the independent AI model leaderboard. Compare LLMs on real-world coding benchmarks, reasoning accuracy, tokens per second, and API inference pricing.

50+Verified Models
100%Empirical Ranks
DailySpeed & Pricing
Verified 2026 AI Benchmark Dataset

AI Model Leaderboard

Empirically evaluated on reasoning (GPQA/MATH), SWE-bench coding, agentic autonomy, inference throughput (tps), and API pricing.

#1
OpenAI•Proprietary
Overall68.5
Reason96.0
Code74.1
Speed59 tps
Price/M$10.00
Scorecard→
#2
Anthropic•Proprietary
Overall66.8
Reason88.5
Code86.2
Speed42 tps
Price/M$5.00
Scorecard→
#3
Anthropic•Proprietary
Overall64.2
Reason93.7
Code67.4
Speed65 tps
Price/M$10.00
Scorecard→
#4
Anthropic•Proprietary
Overall62.0
Reason67.5
Code84.5
Speed110 tps
Price/M$10.00
Scorecard→
#5
OpenAI•Proprietary
Overall60.5
Reason78.5
Code84.8
Speed126 tps
Price/M$2.00
Scorecard→
#6
Anthropic•Proprietary
Overall58.2
Reason65.8
Code82.2
Speed50 tps
Price/M$0.00
Scorecard→
#7
Anthropic•Proprietary
Overall57.5
Reason64.5
Code81.8
Speed68 tps
Price/M$5.00
Scorecard→
#8
Meta•Proprietary
Overall57.4
Reason63.5
Code75.4
Speed221 tps
Price/M$1.25
Scorecard→
#9
OpenAI•Proprietary
Overall57.2
Reason65.5
Code81.2
Speed89 tps
Price/M$4.00
Scorecard→
#10
Google•Proprietary
Overall56.8
Reason62.5
Code73.7
Speed297 tps
Price/M$0.75
Scorecard→
#11
Xiaomi•Proprietary
Overall56.5
Reason65.2
Code78.4
Speed55 tps
Price/M$0.44
Scorecard→
#12
Moonshot AI•Open Source
Overall56.0
Reason64.2
Code80.5
Speed52 tps
Price/MFree/Open Source
Scorecard→
#13
Qwen / Alibaba•Proprietary
Overall55.8
Reason62.8
Code79.4
Speed39 tps
Price/M$2.00
Scorecard→
#14
Zhipu AI•Open Source
Overall55.8
Reason60.5
Code78.2
Speed61 tps
Price/MFree/Open Source
Scorecard→
#15
xAI•Proprietary
Overall55.6
Reason63.2
Code78.5
Speed59 tps
Price/M$2.00
Scorecard→
#16
Zhipu AI•Open Source
Overall55.4
Reason62.5
Code77.8
Speed57 tps
Price/M$1.40
Scorecard→
#17
Tencent•Open Source
Overall55.2
Reason58.5
Code77.0
Speed60 tps
Price/MFree/Open Source
Scorecard→
#18
OpenAI•Proprietary
Overall55.0
Reason59.2
Code77.4
Speed117 tps
Price/M$2.00
Scorecard→
#19
Anthropic•Proprietary
Overall54.8
Reason59.5
Code76.5
Speed57 tps
Price/M$5.00
Scorecard→
#20
Anthropic•Proprietary
Overall54.8
Reason58.6
Code75.8
Speed77 tps
Price/M$2.00
Scorecard→
#21
DeepSeek•Open Source
Overall54.4
Reason62.5
Code81.0
Speed176 tps
Price/MFree/Open Source
Scorecard→
#22
DeepSeek•Open Source
Overall54.2
Reason58.4
Code80.6
Speed50 tps
Price/M$1.74
Scorecard→
#23
Google•Proprietary
Overall54.2
Reason61.5
Code78.6
Speed169 tps
Price/M$0.75
Scorecard→
#24
Anthropic•Proprietary
Overall54.0
Reason57.2
Code72.4
Speed35 tps
Price/M$5.00
Scorecard→
#25
OpenAI•Proprietary
Overall53.8
Reason58.2
Code72.5
Speed14 tps
Price/M$2.50
Scorecard→
Showing top 25 of 85 empirically verified AI models • Updated live dailyView All 50+ Models on AI Leaderboard
Side-by-Side Evaluation Engine

Compare Any Two AI Models

Select any two models from our verified dataset for an instant head-to-head spec duel, accuracy delta, and API cost calculation.

Full Comparison Hub
Popular Head-to-Head Presets:
#1

GPT-5.6 Sol

OpenAI • Commercial API
Overall61.8
VS
#2

Claude Opus 5

Anthropic • Commercial API
Overall61.0
Overall Benchmark

RankLLMs Score

GPT-5.6 Sol Leads (+0.8)
100
50
61.8
61.0
GPT-5.6 Sol
Claude Opus 5
Composite IndexScale: 0-100
Coding Benchmark

SWE-bench Verified

Claude Opus 5 Leads (+0.7)
100
50
58.5
59.2
GPT-5.6 Sol
Claude Opus 5
Software DevSolved Task %
Reasoning Benchmark

GPQA Diamond

Claude Opus 5 Leads (+2.8)
100
50
55.2
58.0
GPT-5.6 Sol
Claude Opus 5
Hard Science & MathDiamond %
Streaming SpeedGPT-5.6 Sol +43 tps faster
85 tps
Tokens / Sec
42 tps
Inference PricingClaude Opus 5 is 7% cheaper
$7.78 / 1M
Cost / 1M Tokens
$7.22 / 1M
Comparing OpenAI's frontier GPT-5.6 Sol against Anthropic's Claude Opus 5.Open Full In-Depth Matchup Analysis
Authoritative Research

Latest Benchmark Guides & Reviews

In-depth technical evaluations, CLI agent testing, and foundation model launch analysis.

View All (38)
Knowledge Base

Frequently Asked Questions About LLM Benchmarks

Clear answers to common technical questions about Large Language Model evaluation and API selection.

What is RankLLMs?

RankLLMs is an independent AI benchmark leaderboard and Large Language Model comparison platform. We provide objective, reproducible evaluations of proprietary and open-weights models based on coding accuracy, mathematical reasoning, tokens-per-second speed, and real-world API inference costs.

What is the highest-ranked AI model in 2026?

OpenAI's GPT-6 Astra currently holds the #1 overall position on RankLLMs with a composite score of 68.5, followed by Anthropic's Claude Fable 5.1 (64.2) and Claude Fable 5 (62.0). For open-source and open-weights models, Moonshot AI's Kimi K3 (56.0), Alibaba's Qwen3.8 Max (55.8), and Zhipu AI's GLM-5.3-Flash (55.8) lead the global rankings.

How are LLM benchmark scores measured on RankLLMs?

RankLLMs aggregates standardized evaluation frameworks including multi-file repository coding benchmarks, mathematical reasoning, Code Arena Elo rankings, and agentic tool-use capability, combined with verified inference speed (tokens/sec and Time-To-First-Token) and API token pricing per 1M tokens.

Which LLM is best for autonomous coding and software engineering?

OpenAI's GPT-6 Astra (88.5 Coding) and Anthropic's Claude Fable 5.1 (86.0 Coding) and Claude Fable 5 (84.5 Coding) rank highest among proprietary systems. For open-weights software development, Moonshot AI's Kimi K3 (80.5 Coding), Alibaba's Qwen3.8 Max (79.4 Coding), and Zhipu AI's GLM-5.3-Flash (78.2 Coding) provide near-commercial performance at a fraction of API token costs.

What are the fastest and most affordable open-weights models?

DeepSeek-V4-Flash-0731 ($0.14/M tokens at 176 tokens/sec), GLM-5.3-Flash ($0.19/M tokens), and Poolside Laguna S 2.1 ($0.11/M tokens) offer state-of-the-art inference efficiency for high-throughput enterprise pipelines.

How often is the RankLLMs leaderboard updated?

The RankLLMs leaderboard is continuously updated whenever foundation model providers (OpenAI, Anthropic, Google, DeepSeek, Zhipu AI, Moonshot AI, Alibaba Cloud / Qwen Team, Meta, xAI, ByteDance, MiniMax) release new model checkpoints, benchmark evaluations, or update their public API token pricing.

RankLLMs - The Open AI Model Leaderboard & Comparison Platform

Welcome to RankLLMs (also searched as rankllm), the definitive independent platform for real-time AI model comparison, benchmark analysis, and Large Language Model performance tracking. Whether you are an AI engineer selecting the optimal LLM API for production, a researcher evaluating frontier reasoning accuracy, or a software developer searching for autonomous CLI coding agents, RankLLMs provides transparent, data-driven evaluations across proprietary and open-weights artificial intelligence models.

Our tracking covers the current frontier - systems like GPT-5.6 Sol, Claude Fable 5, Claude Opus 5, Gemini 3.7 Flash, and Grok 4.6 - alongside leading open-weights families such as DeepSeek-V4, Qwen3.8, Kimi K3, and GLM-5.3. The full catalog of 80+ model scorecards is updated continuously, so the leaderboard reflects the market as it is now, not a launch-day snapshot.

Software Engineering & Coding Accuracy

Evaluating models on multi-file codebase ingestion, terminal command execution, real GitHub issue resolution, and precise syntax-valid patch generation without hallucination. Frontier models such as GPT-6 Astra, Claude Fable 5.1, and Kimi K3 lead public coding benchmarks with exceptional resolve rates.

Reasoning & Formal Logic (Chain-of-Thought)

Measuring step-by-step chain-of-thought verification, mathematical proof formulation, and scientific reasoning depth across frontier reasoning systems including GPT-6 Astra (69.8 Reasoning), Claude Fable 5.1 (68.5 Reasoning), and Claude Fable 5 (67.5 Reasoning).

Throughput, Latency & Generation Speed (c/s)

Measuring real-world Time-To-First-Token (TTFT) and streaming character throughput (characters per second) across cloud endpoints to ensure interactive applications and automated agent pipelines maintain minimal latency - highlighted by Gemini 3.5 Flash (348 c/s) and DeepSeek-V4-Flash-0731 (176 c/s).

API Economics & Blended Token Pricing ($/M)

Tracking standardized blended token pricing per 1 million tokens across models from high-efficiency open weights like DeepSeek-V4-Flash ($0.10/M) and Poolside Laguna S 2.1 ($0.11/M) up to frontier systems like Claude Opus 5 ($7.22/M) and GPT-5.6 Sol ($7.78/M).

Published and maintained by Lucky Yaduvanshi • Independent AI Benchmark Research.