Claude 3.5 Sonnet Benchmarks: SWE-bench Score, HumanEval & Coding (2026)
Every Claude 3.5 Sonnet benchmark in one place: SWE-bench Verified from 33.4% to 49.0%, HumanEval 93.7%, GPQA, and how it stacks up against the 2026 frontier models.


Synthesizing article benchmarks & model metrics...
Claude 3.5 Sonnet was the foundation model that established repository-level software engineering as the definitive benchmark for frontier artificial intelligence. Released by Anthropic in June 2024 and upgraded in October 2024, it held the SWE-bench Verified benchmark record for over a year and served as the default coding engine across tools like GitHub Copilot, Cursor, and Cline.
Because Claude 3.5 Sonnet was released in two distinct snapshots with identical names, many online sources conflate their benchmark scores. Below is the verified chronological data, calculated performance deltas between releases, and a comparative analysis against the 2026 frontier.
Claude 3.5 Sonnet Chronological Specifications
| Specification Metric | Original Snapshot (20240620) | Upgraded Snapshot (20241022) | Benchmark Trajectory |
|---|---|---|---|
| Release Date | June 21, 2024 | October 22, 2024 | 4-month upgrade cycle |
| Context Window | 200,000 Tokens | 200,000 Tokens | Unchanged |
| Input Token Pricing | $3.00 / 1M tokens | $3.00 / 1M tokens | Zero price increase |
| Output Token Pricing | $15.00 / 1M tokens | $15.00 / 1M tokens | Zero price increase |
| SWE-bench Verified | 33.4% | 49.0% | +15.6 pts (+46.7% relative gain) |
| HumanEval Score | Industry-leading at launch | 93.7% | Function-level saturation |
| TAU-bench Airline | 36.0% | 46.0% | +10.0 pts (+27.8% relative gain) |
| OSWorld Computer Use | N/A | 14.9% | Inaugural computer-use beta |
| Headline Capability | Claude Artifacts | Computer Use & Desktop Control | Shift to agentic interaction |
What Actually Changed Between June and October?
When Anthropic deployed the October 2024 update, pricing and token limits remained unchanged. However, the architectural and RLHF refinements yielded massive relative gains on multi-turn software development:
- SWE-bench Verified (33.4% $\rightarrow$ 49.0%): A +46.7% relative improvement. At 49.0%, Claude 3.5 Sonnet became the first model capable of resolving nearly half of real-world GitHub issues end-to-end, including locating the faulty files, writing the patch, and executing the project test suite.
- TAU-bench Retail (62.6% $\rightarrow$ 69.2%): A +10.5% relative gain in conversational database operations and user tool interaction.
- TAU-bench Airline (36.0% $\rightarrow$ 46.0%): A +27.8% relative improvement in complex multi-step state management.
Comparative Context: 2024 Benchmark vs 2026 Frontier
To understand how software engineering evaluation has progressed, the table below compares the peak Claude 3.5 Sonnet against current 2026 models on SWE-bench and Terminal-Bench:
| Model Generation | SWE-bench Verified | Terminal-Bench 2.1 | Input Price / 1M | Output Price / 1M |
|---|---|---|---|---|
| Claude 3.5 Sonnet (Oct 2024) | 49.0% | 18.2% | $3.00 | $15.00 |
| Claude Sonnet 4.6 (2025) | 58.4% | 24.8% | $3.00 | $15.00 |
| Claude Sonnet 5 (2026) | 66.8% | 31.2% | $2.00 | $10.00 |
| Claude Fable 5 Max (2026) | 70.0% | 34.1% | $6.00 | $18.00 |
| DeepSeek V4 Pro 0813 (2026) | 62.7% | 87.9 | $0.435 | $0.87 |
As demonstrated by the data:
- SWE-bench Progression: The frontier has moved from resolving 49% of issues in 2024 to 70%+ in 2026 with models like Claude Fable 5.
- Economic Deflation: Modern successors like Claude Sonnet 5 deliver 17.8 points higher SWE-bench accuracy while reducing costs by 33% ($2/$10 vs $3/$15).
- Open-Weights Parity: Open models like DeepSeek V4 Pro now surpass Claude 3.5 Sonnet across all agentic metrics at less than 1/15th the price.
Related Models & Discovery Resources
- Successor Scorecard: Specifications on Claude Sonnet 5
- Frontier Flagship: Technical specs on Claude Fable 5
- Compare Progress: Simulate generational matchups in the LLM Comparison Engine
- SWE-bench Explainer: Learn how contamination-free testing works in our SWE-bench Pro Guide
Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.
- •Anthropic Claude 3.5 Sonnet Model Card Addendum(Primary Source →)
- •SWE-bench Official Benchmark Repository(Primary Source →)

Lucky Yaduvanshi(luckyyaduvanshi.in →)
Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.
Recommended Reading
Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Best LLMs for Coding in 2026: SWE-bench Verified Data & Cost Math
The best coding LLMs in 2026, ranked by SWE-bench Verified, Terminal Bench, and cost per solved task. Claude Fable 5 leads raw accuracy; Gemini 3.7 Flash and GLM-5.3-Flash lead value.
Lucky Yaduvanshi
Claude Sonnet 5: Permanent $2/$10 Pricing and Developer Impact
Anthropic has formalized Claude Sonnet 5's $2/$10 API pricing permanently, canceling scheduled increases. Here is the token economics breakdown.
Lucky Yaduvanshi
GLM-5.3-FlashX: 200 Tokens/sec Throughput vs 2.5x Price Analysis
In-depth review of GLM-5.3-FlashX: 200 tokens/second inference throughput, latency benchmarks, and whether the 2.5x pricing premium is justified.
Lucky Yaduvanshi
TypeSafe AI Jev Access Guide: Free Trial Endpoints, API Pricing, and System One Architecture
How to access TypeSafe AI's Jev model for free: Vercel AI Gateway promotion, OpenRouter pricing at $0.042/1M tokens, latency benchmarks, and System One design.
Lucky Yaduvanshi