Grok 4.6: Benchmark Results, Pricing, and What Changed
Grok 4.6 delivers major improvements on coding and agent benchmarks over Grok 4.5 while maintaining $2/$6 per million token pricing. Here is the independent breakdown.


Synthesizing article benchmarks & model metrics...
SpaceXAI and xAI officially deployed Grok 4.6 on August 12, 2026. Rather than introducing a price increase, the release maintains the exact $2.00 per million input and $6.00 per million output API rate of Grok 4.5 while registering a 61 composite rating on the Artificial Analysis Intelligence Index.
In standardized evaluations, Grok 4.6 leads several knowledge-synthesis and reasoning evaluations—such as GDPVal-AA v2 at 1,753 and AA-Briefcase at 1,577—while showing substantial relative double-digit improvements on agentic software development benchmarks.
Below is an independent analysis of the published evaluation data, calculated deltas against Grok 4.5, comparative standings against GPT-5.6 Sol and Claude Fable 5, and practical token economics for production engineering teams.
Introducing Grok 4.6.
— SpaceXAI (@SpaceXAI) August 12, 2026
It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price. pic.twitter.com/RtTbpXcb3a
Quick Specifications & Deployment Summary
| Specification | Grok 4.6 High | Notes & Comparison |
|---|---|---|
| Developer / Foundation | SpaceXAI / xAI | Successor to Grok 4.5 |
| Input Token Pricing | $2.00 / 1M tokens | Parity with Grok 4.5 |
| Output Token Pricing | $6.00 / 1M tokens | Parity with Grok 4.5 |
| AA Intelligence Index | 61 | Tied with GPT-5.6 Sol Max (61); Fable 5 (62) |
| GDPVal-AA v2 | 1,753 | #1 Globally among tested frontier systems |
| AA-Briefcase | 1,577 | #1 Globally in enterprise document synthesis |
| CursorBench v3.2 | 69.9% | Up from 66.7% in Grok 4.5 |
| DeepSWE v1.1 | 65.9% | Up from 54.0% in Grok 4.5 |
| Terminal-Bench v3.0 | 26.0% | Up from 15.7% in Grok 4.5 |
| Full Model Scorecard | Grok 4.6 Specs & Rating | View global ranking in AI Leaderboard |
What Actually Changed? Calculated Deltas vs Grok 4.5
The most critical analytical signal in any model release is the relative performance gain across identical test harnesses. Analyzing the published metrics reveals where engineering effort was concentrated:
| Benchmark Evaluation | Grok 4.5 | Grok 4.6 | Absolute Delta | Relative Improvement |
|---|---|---|---|---|
| Terminal-Bench v3.0 | 15.7% | 26.0% | +10.3 pts | +65.6% relative gain |
| DeepSWE v1.1 | 54.0% | 65.9% | +11.9 pts | +22.0% relative gain |
| APEX-Agents | 47.1% | 57.5% | +10.4 pts | +22.1% relative gain |
| AA-Briefcase | 1,313 | 1,577 | +264 pts | +20.1% relative gain |
| GDPVal-AA v2 | 1,526 | 1,753 | +227 pts | +14.9% relative gain |
| AA Intelligence Index | 56 | 61 | +5 pts | +8.9% relative gain |
| CursorBench v3.2 | 66.7% | 69.9% | +3.2 pts | +4.8% relative gain |
| Harvey LAB (Legal) | 12.9% | 15.8% | +2.9 pts | +22.5% relative gain |
Key Takeaways from the Deltas:
- Agentic Tool & CLI Execution (+65.6%): The single largest jump is in Terminal-Bench v3.0, where Grok climbed from a mediocre 15.7% to 26.0%. This indicates targeted reinforcement training around bash shell execution, environment variable handling, and directory navigation.
- Software Patching & Issue Resolution (+22.0%): A 11.9-point absolute gain on DeepSWE v1.1 places Grok 4.6 in the top tier of automated bug-fixing models, closing the historical gap with Anthropic’s coding models.
- Price-to-Capability Multiplier: Because input and output pricing remained static at $2/$6, the effective capability per dollar increased by an average of 22% across reasoning and agentic tasks.
Frontier Model Comparative Results

To understand where Grok 4.6 stands across the wider frontier, the table below compares it against Grok 4.5, GPT-5.6 Sol Max, and Claude Fable 5 Max:
| Evaluation Suite | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Claude Fable 5 Max | Frontier Category Leader |
|---|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 | Claude Fable 5 Max (62) |
| GDPVal-AA v2 | 1,753 | 1,526 | 1,728 | 1,741 | Grok 4.6 High (1,753) |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% | Claude Fable 5 Max (70.5%) |
| DeepSWE v1.1 | 65.9% | 54.0% | 73.0% | 70.0% | GPT-5.6 Sol Max (73.0%) |
| FrontierCode v1.1 Extended | 61.3% | 56.6% | 60.6% | 64.9% | Claude Fable 5 Max (64.9%) |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% | Claude Fable 5 Max (59.2%) |
| Terminal-Bench v3.0 | 26.0% | 15.7% | 34.6% | 34.1% | GPT-5.6 Sol Max (34.6%) |
| AA-Briefcase | 1,577 | 1,313 | 1,502 | 1,574 | Grok 4.6 High (1,577) |
| Harvey LAB (Legal) | 15.8% | 12.9% | 2.5% | 11.3% | Grok 4.6 High (15.8%) |
Where Grok 4.6 Performs Well
Rather than declaring a single “best overall” model, technical evaluation requires isolating distinct capability envelopes:
- Complex Information Retrieval & Briefcase Synthesis:
- On AA-Briefcase (1,577), Grok 4.6 edges out Claude Fable 5 (1,574) and GPT-5.6 Sol (1,502). This test stresses long-context document understanding, synthesis across conflicting sources, and executive summary precision.
- Domain-Specific Legal Reasoning (Harvey LAB):
- Scoring 15.8% on Harvey LAB, Grok 4.6 substantially outscores GPT-5.6 Sol Max (2.5%) and Claude Fable 5 Max (11.3%). For legal tech teams analyzing contract clauses, statutory compliance, and case citations, Grok demonstrates distinct specialized tuning.
- General Knowledge Valuation (GDPVal-AA v2):
- At 1,753, Grok 4.6 ranks first among all evaluated frontier models, proving robust factual recall and low hallucination rates in open-domain knowledge queries.
Where Grok 4.6 Falls Behind & Practical Limitations
Honest benchmark analysis requires identifying where competing models maintain clear superiority:
- Repository-Scale Coding (DeepSWE v1.1): While Grok improved by 11.9 points, its 65.9% score still trails GPT-5.6 Sol Max (73.0%) by 7.1 percentage points and Claude Fable 5 Max (70.0%) by 4.1 points. For complex multi-file pull requests requiring deep architectural dependency tracing, OpenAI and Anthropic remain ahead. For a complete look at coding rankings, see our Best LLMs for Coding Guide.
- Autonomous CLI Operations (Terminal-Bench): Despite the +65.6% relative jump to 26.0%, Grok 4.6 lags substantially behind GPT-5.6 Sol Max (34.6%) and Fable 5 (34.1%). Autonomous DevOps scripts that require recovering from shell execution errors still encounter higher failure rates on Grok.
- IDE Inline Completions (FrontierCode & CursorBench): Claude Fable 5 Max remains the strongest inline coding model, leading CursorBench at 70.5% and FrontierCode Extended at 64.9%.
Token Economics & Production Cost Breakdown
Grok 4.6’s biggest structural advantage is not raw benchmark dominance, but cost-adjusted intelligence:
| Model Tier | Input / 1M | Output / 1M | Blended Cost (80% In / 20% Out) | Relative Cost Index |
|---|---|---|---|---|
| Grok 4.6 High | $2.00 | $6.00 | $2.80 / 1M | Baseline (1.0x) |
| GPT-5.6 Sol | $5.00 | $15.00 | $7.00 / 1M | 2.5x more expensive |
| Claude Fable 5 | $6.00 | $18.00 | $8.40 / 1M | 3.0x more expensive |
| DeepSeek V4 Pro | $0.14 | $0.435 | $0.20 / 1M | 0.07x (Open-weights budget leader) |
Production Scenario: 50,000 Multi-Turn Agent Invocations
Assuming an average task lifecycle of 200,000 prompt tokens (with context caching) and 10,000 generated tokens:
- Total Workload: 10 Billion Input Tokens, 500 Million Output Tokens.
- Grok 4.6 Total Bill: $(10{,}000 imes $2.00) + (500 imes $6.00) = $20{,}000 + $3{,}000 =$ $23,000.
- GPT-5.6 Sol Equivalent Bill: $(10{,}000 imes $5.00) + (500 imes $15.00) = $50{,}000 + $7{,}500 =$ $57,500.
- Net Infrastructure Savings: $34,500 per 50,000 complex agent sessions.
For engineering teams whose workloads are bounded by budget rather than absolute peak coding scores, Grok 4.6 provides 90–95% of frontier capability at less than half the operating expense.
What We Still Do Not Know
As with any major foundation release, several questions require independent long-term validation:
- Third-Party Contamination Audits: The GDPVal and AA-Briefcase numbers originate from initial partner benchmarks. Independent reproductions on contamination-resistant datasets like SWE-bench Pro and Humanity’s Last Exam will establish whether gains persist on fresh tasks.
- Context Window Degradation Under High Concurrency: Early API reports show solid low-latency streaming at typical prompt lengths, but throughput degradation on 128k+ token needle-in-a-haystack prompts remains to be verified under production enterprise load.
Related Models & Discovery Resources
- Model Scorecard: Detailed latency, context limits, and specs on the Grok 4.6 Model Profile
- Compare Head-to-Head: Simulate duels in the LLM Comparison Engine
- Global Rankings: View all frontier models on the AI Model Leaderboard
- Browse Catalog: Explore all 85 tracked models in the All Models Directory
RankLLMs independent evaluations verify official benchmarks against reproducible testing suites, community logs, and provider documentation.
- •SpaceXAI / xAI Official Release Announcement(Primary Source →)
- •Artificial Analysis Intelligence Index & Benchmark Data(Primary Source →)
Was this benchmark analysis helpful?
Thank you for your feedback! We update our benchmarks weekly based on developer input.

Lucky Yaduvanshi(luckyyaduvanshi.in →)
Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.
Recommended Reading
Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

DeepSeek V4 Pro 0813: Architectural Analysis, Benchmarks, and $0.435/1M Economics
DeepSeek V4 Pro 0813 achieves 87.9 on Terminal-Bench 2.1 using a 1.6T MoE architecture at $0.435/1M tokens. Here is the technical report and benchmark analysis.
Lucky Yaduvanshi
Qwen3.8 Flash vs GLM-5.3 Flash: Lightweight Frontier Coding Comparison
Head-to-head comparison of Qwen3.8 Flash and GLM-5.3 Flash: coding benchmarks, tool use, throughput latency, and per-token pricing.
Lucky Yaduvanshi
NVIDIA Nemotron 3.5 Lightning: Open Weights Architecture and Agentic Benchmarks
NVIDIA Nemotron 3.5 Lightning is a 30B MoE model (3B active) delivering 4x faster execution speed for agentic tool use. Here is the technical breakdown.
Lucky Yaduvanshi
DeepSeek V4 Flash 0731: Benchmarks, Features, and Inference Speed
DeepSeek V4 Flash 0731: post-training benchmarks across Terminal-Bench (82.7%) and DeepSWE (54.4%), 1M context efficiency, and $0.14/$0.28 token economics.
Lucky Yaduvanshi