Claude 3.5 Sonnet Benchmarks: SWE-bench Score, HumanEval & Coding (2026)

Every Claude 3.5 Sonnet benchmark in one place: SWE-bench Verified from 33.4% to 49.0%, HumanEval 93.7%, GPQA, and how it stacks up against the 2026 frontier models.

Lucky Yaduvanshi
Lucky Yaduvanshi
Founder & AI Lead
Sep 04, 2026•Updated Sep 24, 2026•3 min read
Independent technical benchmark • Primary data & verified methodology cited below
Claude 3.5 Sonnet Benchmarks: SWE-bench Score, HumanEval & Coding (2026)

Claude 3.5 Sonnet was the foundation model that established repository-level software engineering as the definitive benchmark for frontier artificial intelligence. Released by Anthropic in June 2024 and upgraded in October 2024, it held the SWE-bench Verified benchmark record for over a year and served as the default coding engine across tools like GitHub Copilot, Cursor, and Cline.

Because Claude 3.5 Sonnet was released in two distinct snapshots with identical names, many online sources conflate their benchmark scores. Below is the verified chronological data, calculated performance deltas between releases, and a comparative analysis against the 2026 frontier.


Claude 3.5 Sonnet Chronological Specifications

Specification Metric Original Snapshot (20240620) Upgraded Snapshot (20241022) Benchmark Trajectory
Release Date June 21, 2024 October 22, 2024 4-month upgrade cycle
Context Window 200,000 Tokens 200,000 Tokens Unchanged
Input Token Pricing $3.00 / 1M tokens $3.00 / 1M tokens Zero price increase
Output Token Pricing $15.00 / 1M tokens $15.00 / 1M tokens Zero price increase
SWE-bench Verified 33.4% 49.0% +15.6 pts (+46.7% relative gain)
HumanEval Score Industry-leading at launch 93.7% Function-level saturation
TAU-bench Airline 36.0% 46.0% +10.0 pts (+27.8% relative gain)
OSWorld Computer Use N/A 14.9% Inaugural computer-use beta
Headline Capability Claude Artifacts Computer Use & Desktop Control Shift to agentic interaction

What Actually Changed Between June and October?

When Anthropic deployed the October 2024 update, pricing and token limits remained unchanged. However, the architectural and RLHF refinements yielded massive relative gains on multi-turn software development:

  • SWE-bench Verified (33.4% $\rightarrow$ 49.0%): A +46.7% relative improvement. At 49.0%, Claude 3.5 Sonnet became the first model capable of resolving nearly half of real-world GitHub issues end-to-end, including locating the faulty files, writing the patch, and executing the project test suite.
  • TAU-bench Retail (62.6% $\rightarrow$ 69.2%): A +10.5% relative gain in conversational database operations and user tool interaction.
  • TAU-bench Airline (36.0% $\rightarrow$ 46.0%): A +27.8% relative improvement in complex multi-step state management.

Comparative Context: 2024 Benchmark vs 2026 Frontier

To understand how software engineering evaluation has progressed, the table below compares the peak Claude 3.5 Sonnet against current 2026 models on SWE-bench and Terminal-Bench:

Model Generation SWE-bench Verified Terminal-Bench 2.1 Input Price / 1M Output Price / 1M
Claude 3.5 Sonnet (Oct 2024) 49.0% 18.2% $3.00 $15.00
Claude Sonnet 4.6 (2025) 58.4% 24.8% $3.00 $15.00
Claude Sonnet 5 (2026) 66.8% 31.2% $2.00 $10.00
Claude Fable 5 Max (2026) 70.0% 34.1% $6.00 $18.00
DeepSeek V4 Pro 0813 (2026) 62.7% 87.9 $0.435 $0.87

As demonstrated by the data:

  1. SWE-bench Progression: The frontier has moved from resolving 49% of issues in 2024 to 70%+ in 2026 with models like Claude Fable 5.
  2. Economic Deflation: Modern successors like Claude Sonnet 5 deliver 17.8 points higher SWE-bench accuracy while reducing costs by 33% ($2/$10 vs $3/$15).
  3. Open-Weights Parity: Open models like DeepSeek V4 Pro now surpass Claude 3.5 Sonnet across all agentic metrics at less than 1/15th the price.

Sources, Disclosures & Primary Benchmark Data

Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.

Share Article
Lucky Yaduvanshi

Lucky Yaduvanshi(luckyyaduvanshi.in →)

Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.

Curated Research

Recommended Reading

Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Explore All 38 Guides→