Verified AI Benchmarks • Updated September 2026

Humanity's Last Exam Explained

What Humanity's Last Exam (HLE) measures, why 2,500 expert-written questions were built to resist memorization, how tools change the scores, and what a good HLE score is in 2026.

Humanity's Last Exam (HLE) is an expert-level benchmark built to sit at the very edge of human knowledge: thousands of questions written and reviewed by subject-matter experts across academia, designed so that answering them requires genuine, deep reasoning rather than retrieval. Its name is half joke, half warning - it was created precisely because older academic benchmarks (MMLU, GSM8K, GPQA) were being saturated by frontier models within months of release.

Why HLE Was Created

By late 2024, the standard evaluation stack had a problem. MMLU was above 88% for frontier models. GPQA Diamond, once brutal, had become a routine checkbox. When a benchmark saturates, it stops discriminating: a 90.4 and a 90.7 tell you nothing about which model is actually smarter. HLE responds by moving the goalposts to expert-frontier territory:

  • Expert-written, expert-reviewed questions spanning advanced mathematics, physics, biology, computer science, humanities, and linguistics - many at the level of active research.
  • Closed-book conditions for the headline score: no internet, no search, testing internalized knowledge and reasoning.
  • Contamination screening so questions that leak into training data can be identified and replaced.
  • Multi-modal and multi-step items, including questions where the answer format itself is demanding (exact values, proofs, structured outputs).

It was created through a public call for questions with experts compensated per accepted problem, and the result is a test where the first frontier models scored in the single digits - a shock after years of watching every benchmark fall within a year.

Closed-Book vs With-Tools: Two Very Different Numbers

HLE is reported two ways, and confusing them is the most common mistake in secondhand coverage:

  • Closed-book: the model answers with its parameters only. This measures internalized knowledge and reasoning at the expert frontier.
  • With tools: the model may browse, run code, and chain reasoning steps. This measures research agent capability, and scores run dramatically higher. In the RankLLMs dataset, GPT-5.6 Sol's 64.5 with tools is tracked as the leading result - far above any closed-book figure, because an agent with a browser turns HLE from a memory test into a research task.

When you read an HLE claim, check which condition it refers to. A "with tools" score is a benchmark of agentic pipelines; a closed-book score is a benchmark of the model itself.

What Counts as a Good HLE Score in 2026?

Context matters more than any single threshold. When the benchmark launched in early 2025, frontier models scored roughly 3-13% closed-book - near-random on many subjects, and a genuine ceiling event. Two years on, the frontier has moved: high-teens to mid-twenties closed-book is realistic for top models, and with-tools agentic results have crossed the 50-65% band, as tracked on our leaderboard.

The interpretation guide:

  • Single digits closed-book: 2025-era frontier. Excellent at college-level knowledge, still failing at research-frontier questions.
  • Teens to low twenties closed-book: 2026 frontier. Meaningful expert-level reasoning, far from saturation.
  • 50%+ with tools: a strong research agent - the model can find, integrate, and verify expert information reliably.

HLE's Known Limitations

  • Narrow audience. HLE measures expert-frontier knowledge. For everyday tasks - writing, business analysis, routine coding - it correlates weakly with useful performance. A model 5 points ahead on HLE may not be better for your work; check coding and agentic pillars separately.
  • Verifier dependence. Grading exact answers and proofs reliably is hard; scoring methodology differences between evaluators can shift results by points.
  • Clock is ticking. Like every benchmark before it, HLE will saturate. Its real legacy is proving that expert-frontier evaluation at scale is possible - and pushing the industry toward contamination-resistant design, the same principle behind SWE-bench Pro.

The Bottom Line

Humanity's Last Exam answers a question most benchmarks can no longer ask: how close is AI to the edge of what expert humans know? Use it to gauge raw intellectual capability and research-agent quality - and pair it with coding benchmarks, speed, and price when you are actually choosing a model. The current standings, across all pillars, are on the RankLLMs leaderboard.