Verified AI Benchmarks • Updated September 2026

LLM Benchmarks Explained

Plain-English guides to the benchmarks that rank AI models: SWE-bench Pro, Humanity's Last Exam, Terminal Bench, GPQA, and how to read leaderboard scores.

Why Benchmark Literacy Matters

A leaderboard is only as trustworthy as the benchmarks behind it. Each evaluation below measures a different slice of model capability, and none of them measures all of it. These guides explain what each benchmark tests, how it is scored, where it falls short, and how the current frontier models perform on it - so the numbers on our leaderboard mean something when you read them.

Benchmark Guides

The Benchmarks Behind Our Composite Score

Our methodology blends four pillars. These are the evaluations that feed them:

  • SWE-bench Verified & Pro

    Resolving real GitHub issues end to end. The gold standard for agentic coding.

  • Terminal Bench

    Multi-step tasks in a live terminal: builds, debugging, environment repair.

  • GPQA Diamond & MATH-500

    Graduate-level science questions and competition mathematics for reasoning.

  • OSWorld & BrowseComp

    Computer-use and web-research tasks for agentic operation.

Reading Scores Correctly: Three Rules

  • Check the version. A model's score changes across snapshots - always compare the same release, and watch for contamination-resistant benchmarks when comparing across time.
  • Check the harness. Agentic benchmarks depend on scaffolding, compute budget, and step limits; a score without those details is incomplete.
  • Check the date. The frontier moves monthly. A "state of the art" claim from two quarters ago is history, not guidance - our news hub tracks what changed and when.