Claude 3.5 Sonnet Benchmarks: SWE-bench, HumanEval and Full Scores

Lucky YaduvanshiLucky YaduvanshiSeptember 04, 20266 min readLoading views...
Claude 3.5 Sonnet Benchmarks: SWE-bench, HumanEval and Full Scores

Claude 3.5 Sonnet was the model that made agentic coding a mainstream benchmark category. Released by Anthropic in June 2024 and upgraded in October 2024, it held the SWE-bench Verified crown for over a year and became the default coding model for an entire generation of developer tools. This page documents its complete benchmark history with verified numbers, explains what each result meant, and puts the model in context against the 2026 frontier.

The reason this page gets searched so heavily is that Claude 3.5 Sonnet’s scores changed between versions, and most secondhand sources mix them up. Below, every number is tied to its exact release and snapshot date.

Claude 3.5 Sonnet at a Glance

Original (June 2024) Upgraded (October 2024)
Snapshot claude-3-5-sonnet-20240620 claude-3-5-sonnet-20241022
Released June 21, 2024 October 22, 2024
Context window 200K tokens 200K tokens
Input price $3 / million tokens $3 / million tokens
Output price $15 / million tokens $15 / million tokens
SWE-bench Verified 33.4% 49.0%
HumanEval Industry-leading at launch 93.7%
Headline feature Artifacts Computer use (public beta)

Pricing and context never changed between versions. Anthropic shipped the October upgrade “at the same price and speed as its predecessor,” which is a large part of why adoption was so fast: developers got a 15.6 point SWE-bench jump for free.

SWE-bench Verified: The Number That Defined the Model

SWE-bench Verified measures whether a model can resolve real GitHub issues in real codebases, and it became the benchmark that frontier models compete on largely because of this model’s trajectory.

  • June 2024 launch: 33.4%. At release, this was the best publicly available score, ahead of GPT-4o and every other shipped model. Anthropic also reported an internal agentic coding evaluation where Claude 3.5 Sonnet solved 64% of problems against 38% for Claude 3 Opus.
  • October 2024 upgrade: 49.0%. The refreshed version pushed nearly 16 points higher and, per Anthropic, outperformed all publicly available models on the benchmark at release. For the first time, a model resolved roughly half of real-world software engineering tasks end to end.

Why did this matter so much? Before Claude 3.5 Sonnet, “AI coding” meant autocomplete and chat about code. A 49% SWE-bench score meant a model could take an issue description, navigate an unfamiliar repository, edit multiple files, and produce a passing patch. GitHub Copilot’s later adoption of the upgraded model for multi-file editing flows is a direct consequence of this capability jump, and every SWE-bench leaderboard you see today - including the RankLLMs leaderboard - descends from the competitive era this model started.

HumanEval: 93.7% on the Upgraded Version

The October 2024 upgrade scored 93.7% on HumanEval, the classic function-level Python code generation benchmark. At that point HumanEval was nearly saturated - the remaining failures were mostly edge cases and underspecified prompts rather than realistic capability gaps - which is precisely why the industry shifted attention to SWE-bench and other agentic evaluations.

At the June 2024 launch, Anthropic reported new industry-leading results on HumanEval and GPQA (graduate-level reasoning) without initially publishing exact figures in the announcement text; the upgraded October release’s 93.7% is the widely cited number. For precise per-version GPQA and MMLU figures, the Claude 3 model card addendum remains the authoritative source.

Agentic Benchmarks: TAU-bench and OSWorld

The October upgrade also shipped major agentic improvements, measured on two benchmarks that were new at the time:

Benchmark June 2024 October 2024 What it measures
TAU-bench (retail) 62.6% 69.2% Multi-turn tool-use customer service tasks
TAU-bench (airline) 36.0% 46.0% Harder tool-use policies, more failure modes
OSWorld (screenshot-only) - 14.9% Real computer tasks from screenshots alone

The OSWorld result deserves special mention: 14.9% looks low until you compare it to the 7.8% of the next-best system, and it reached 22.0% when allowed more steps per task. This was the quantitative proof behind “computer use” - the ability to look at a screen, move a cursor, click buttons, and type - which shipped in public beta alongside the October upgrade and kicked off the entire computer-using-agent product category.

Pricing: $3 / $15 per Million Tokens

Claude 3.5 Sonnet cost $3 per million input tokens and $15 per million output tokens with a 200K token context window, across both versions. In 2024 this price-performance ratio was disruptive: it delivered frontier-tier coding at roughly half the price of Claude 3 Opus, which is why it became the workhorse model for startups and coding tools rather than a premium-only option.

For cost context against today’s models, check the pricing columns on the current leaderboard - the per-token cost of Sonnet-class quality has fallen substantially since, which is the clearest way to measure how the market moved.

How Claude 3.5 Sonnet Compares to the 2026 Frontier

Two years of progress have moved every goalpost the model set:

  • Coding: Current frontier models score far beyond the 49.0% SWE-bench Verified ceiling the upgraded 3.5 Sonnet established. Where 49% once led the field, today’s leaders on the RankLLMs leaderboard treat that level as a baseline for non-reasoning modes.
  • Computer use: The 14.9% OSWorld screenshot-only result that impressed in 2024 has been multiplied many times over by dedicated agentic OS models.
  • Pricing: Mid-tier 2026 models deliver comparable coding reliability at a fraction of the $3/$15 price point, and open-weight models now occupy much of the quality band that was proprietary-only in 2024.

The model’s practical legacy is structural rather than competitive: it established the evaluation methodology (SWE-bench Verified as the industry yardstick), the product pattern (artifacts, computer use), and the price-performance expectations that every successor - including the current Sonnet lineup - is measured against.

Is Claude 3.5 Sonnet Still Worth Using in 2026?

For new projects, no: the current Sonnet generation is faster, cheaper, and more capable, and Anthropic’s current model lineup has moved on. For existing systems still pinned to the 20240620 or 20241022 snapshots, the model remains a known quantity with stable, documented behavior - but migration to a current Sonnet model is a straightforward upgrade path with better scores at lower cost.

If you are here to reproduce 2024-era results or to verify what a historical benchmark claim actually said, the tables above tie every number to its official source and release date.

Share Article

Was this benchmark analysis helpful?

Lucky Yaduvanshi

Lucky Yaduvanshi(luckyyaduvanshi.in →)

Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.

Subscribe to AI Benchmark Intel

Get weekly AI model benchmark evaluations, LLM speed/cost breakdowns, and exclusive free API credit alerts delivered to your inbox.