How We Rank AI Models
The RankLLMs methodology: where our benchmark data comes from, how scores are calculated, how pricing and speed are verified, and when the leaderboard is updated.
Our Methodology, In One Paragraph
Every ranking on RankLLMs is built from measurable signals: published benchmark results, vendor-documented API pricing, and measured inference performance. We never accept payment for placement, we date every data point we publish, and when a number cannot be verified from a primary source we either mark it as unverified or leave it out. This page explains exactly how that works so you can judge our data the same way we judge the models we rank.
1. Where Benchmark Data Comes From
We use three tiers of sources, in this order of preference:
- Primary vendor sources. Official model cards, system cards, launch announcements, and API documentation from the labs that built the model (for example, Anthropic's published SWE-bench Verified results or Qwen's official evaluation tables). These always take priority over third-party replications.
- Independent benchmark organizations. Public leaderboards and evaluation suites such as SWE-bench, Terminal Bench, GPQA, MATH-500, OSWorld, and BrowseComp, cited with the exact version and date of the evaluation.
- Our own testing. For speed (tokens per second, time to first token) and hands-on CLI agent reviews, we run the workloads ourselves and describe the hardware, software versions, and prompts used.
When we republish a vendor's number, we link the source. When two credible sources disagree, we show the discrepancy rather than silently picking one.
2. How the Composite Score Works
Each model on the leaderboard carries a composite score from 0 to 100. It is a weighted blend of normalized results across four capability pillars:
- Coding - SWE-bench Verified accuracy and Terminal Bench performance for agentic software engineering.
- Reasoning - GPQA Diamond, MATH-500, and formal logic evaluations.
- Agentic / OS use - OSWorld, BrowseComp, and tool-use benchmarks.
- Arena performance - community Elo ratings where available.
Weights are reviewed when the benchmark landscape changes, and any change to the formula is announced on the changelog before it affects rankings. A model's individual pillar scores are always visible on its scorecard, so you never have to trust the composite alone.
3. Pricing and Speed Verification
API pricing is taken from each provider's official pricing page at the time of publication and re-checked during every data refresh; the price shown is per million tokens, input and output listed separately. Inference speed (tokens per second and time to first token) comes from our own measurement runs on documented hardware, or from provider disclosures where our own run is not yet available - in which case the source is noted on the model's scorecard. Prices are in USD.
4. Update Cadence and Versioning
The leaderboard and model scorecards are refreshed on a rolling basis as new benchmark results and pricing changes are published, with a target of reviewing the top 50 models at least weekly. Every scorecard and article shows when its data was last verified. When a model is updated or re-benchmarked, we keep the historical context visible rather than quietly overwriting it - our historical benchmark pages preserve version-by-version numbers with their exact release dates.
5. Known Limitations
- Benchmark scores measure benchmark performance, not your workload. Use pillar scores and our comparison guides to approximate your use case.
- Speed measurements vary by region, load, and provider tier; we publish our measurement conditions and encourage replication.
- Promotional or preview pricing changes frequently; always confirm current pricing on the provider's page before committing budget.
Corrections and Data Disputes
If you spot an error in our data, or you represent a lab with newer official results, we want to hear it. Send the source and we will verify and correct it publicly. See the contact page for channels, and the editorial policy for how corrections are handled.