Grok 4.6: Benchmark Results, Pricing, and What Changed

Grok 4.6 delivers major improvements on coding and agent benchmarks over Grok 4.5 while maintaining $2/$6 per million token pricing. Here is the independent breakdown.

Lucky Yaduvanshi
Lucky Yaduvanshi
Founder & AI Lead
Aug 13, 2026•Updated Sep 24, 2026•7 min read•Loading views...
Independent technical benchmark • Primary data & verified methodology cited below
Grok 4.6 benchmark evaluation and pricing comparison against GPT-5.6 Sol and Fable 5

SpaceXAI and xAI officially deployed Grok 4.6 on August 12, 2026. Rather than introducing a price increase, the release maintains the exact $2.00 per million input and $6.00 per million output API rate of Grok 4.5 while registering a 61 composite rating on the Artificial Analysis Intelligence Index.

In standardized evaluations, Grok 4.6 leads several knowledge-synthesis and reasoning evaluations—such as GDPVal-AA v2 at 1,753 and AA-Briefcase at 1,577—while showing substantial relative double-digit improvements on agentic software development benchmarks.

Below is an independent analysis of the published evaluation data, calculated deltas against Grok 4.5, comparative standings against GPT-5.6 Sol and Claude Fable 5, and practical token economics for production engineering teams.


Quick Specifications & Deployment Summary

Specification Grok 4.6 High Notes & Comparison
Developer / Foundation SpaceXAI / xAI Successor to Grok 4.5
Input Token Pricing $2.00 / 1M tokens Parity with Grok 4.5
Output Token Pricing $6.00 / 1M tokens Parity with Grok 4.5
AA Intelligence Index 61 Tied with GPT-5.6 Sol Max (61); Fable 5 (62)
GDPVal-AA v2 1,753 #1 Globally among tested frontier systems
AA-Briefcase 1,577 #1 Globally in enterprise document synthesis
CursorBench v3.2 69.9% Up from 66.7% in Grok 4.5
DeepSWE v1.1 65.9% Up from 54.0% in Grok 4.5
Terminal-Bench v3.0 26.0% Up from 15.7% in Grok 4.5
Full Model Scorecard Grok 4.6 Specs & Rating View global ranking in AI Leaderboard

What Actually Changed? Calculated Deltas vs Grok 4.5

The most critical analytical signal in any model release is the relative performance gain across identical test harnesses. Analyzing the published metrics reveals where engineering effort was concentrated:

Benchmark Evaluation Grok 4.5 Grok 4.6 Absolute Delta Relative Improvement
Terminal-Bench v3.0 15.7% 26.0% +10.3 pts +65.6% relative gain
DeepSWE v1.1 54.0% 65.9% +11.9 pts +22.0% relative gain
APEX-Agents 47.1% 57.5% +10.4 pts +22.1% relative gain
AA-Briefcase 1,313 1,577 +264 pts +20.1% relative gain
GDPVal-AA v2 1,526 1,753 +227 pts +14.9% relative gain
AA Intelligence Index 56 61 +5 pts +8.9% relative gain
CursorBench v3.2 66.7% 69.9% +3.2 pts +4.8% relative gain
Harvey LAB (Legal) 12.9% 15.8% +2.9 pts +22.5% relative gain

Key Takeaways from the Deltas:

  1. Agentic Tool & CLI Execution (+65.6%): The single largest jump is in Terminal-Bench v3.0, where Grok climbed from a mediocre 15.7% to 26.0%. This indicates targeted reinforcement training around bash shell execution, environment variable handling, and directory navigation.
  2. Software Patching & Issue Resolution (+22.0%): A 11.9-point absolute gain on DeepSWE v1.1 places Grok 4.6 in the top tier of automated bug-fixing models, closing the historical gap with Anthropic’s coding models.
  3. Price-to-Capability Multiplier: Because input and output pricing remained static at $2/$6, the effective capability per dollar increased by an average of 22% across reasoning and agentic tasks.

Frontier Model Comparative Results

Grok 4.6 Official Benchmark Results Chart

To understand where Grok 4.6 stands across the wider frontier, the table below compares it against Grok 4.5, GPT-5.6 Sol Max, and Claude Fable 5 Max:

Evaluation Suite Grok 4.6 High Grok 4.5 High GPT-5.6 Sol Max Claude Fable 5 Max Frontier Category Leader
AA Intelligence Index 61 56 61 62 Claude Fable 5 Max (62)
GDPVal-AA v2 1,753 1,526 1,728 1,741 Grok 4.6 High (1,753)
CursorBench v3.2 69.9% 66.7% 67.2% 70.5% Claude Fable 5 Max (70.5%)
DeepSWE v1.1 65.9% 54.0% 73.0% 70.0% GPT-5.6 Sol Max (73.0%)
FrontierCode v1.1 Extended 61.3% 56.6% 60.6% 64.9% Claude Fable 5 Max (64.9%)
APEX-Agents 57.5% 47.1% 56.7% 59.2% Claude Fable 5 Max (59.2%)
Terminal-Bench v3.0 26.0% 15.7% 34.6% 34.1% GPT-5.6 Sol Max (34.6%)
AA-Briefcase 1,577 1,313 1,502 1,574 Grok 4.6 High (1,577)
Harvey LAB (Legal) 15.8% 12.9% 2.5% 11.3% Grok 4.6 High (15.8%)

Where Grok 4.6 Performs Well

Rather than declaring a single “best overall” model, technical evaluation requires isolating distinct capability envelopes:

  1. Complex Information Retrieval & Briefcase Synthesis:
    • On AA-Briefcase (1,577), Grok 4.6 edges out Claude Fable 5 (1,574) and GPT-5.6 Sol (1,502). This test stresses long-context document understanding, synthesis across conflicting sources, and executive summary precision.
  2. Domain-Specific Legal Reasoning (Harvey LAB):
    • Scoring 15.8% on Harvey LAB, Grok 4.6 substantially outscores GPT-5.6 Sol Max (2.5%) and Claude Fable 5 Max (11.3%). For legal tech teams analyzing contract clauses, statutory compliance, and case citations, Grok demonstrates distinct specialized tuning.
  3. General Knowledge Valuation (GDPVal-AA v2):
    • At 1,753, Grok 4.6 ranks first among all evaluated frontier models, proving robust factual recall and low hallucination rates in open-domain knowledge queries.

Where Grok 4.6 Falls Behind & Practical Limitations

Honest benchmark analysis requires identifying where competing models maintain clear superiority:

  • Repository-Scale Coding (DeepSWE v1.1): While Grok improved by 11.9 points, its 65.9% score still trails GPT-5.6 Sol Max (73.0%) by 7.1 percentage points and Claude Fable 5 Max (70.0%) by 4.1 points. For complex multi-file pull requests requiring deep architectural dependency tracing, OpenAI and Anthropic remain ahead. For a complete look at coding rankings, see our Best LLMs for Coding Guide.
  • Autonomous CLI Operations (Terminal-Bench): Despite the +65.6% relative jump to 26.0%, Grok 4.6 lags substantially behind GPT-5.6 Sol Max (34.6%) and Fable 5 (34.1%). Autonomous DevOps scripts that require recovering from shell execution errors still encounter higher failure rates on Grok.
  • IDE Inline Completions (FrontierCode & CursorBench): Claude Fable 5 Max remains the strongest inline coding model, leading CursorBench at 70.5% and FrontierCode Extended at 64.9%.

Token Economics & Production Cost Breakdown

Grok 4.6’s biggest structural advantage is not raw benchmark dominance, but cost-adjusted intelligence:

Model Tier Input / 1M Output / 1M Blended Cost (80% In / 20% Out) Relative Cost Index
Grok 4.6 High $2.00 $6.00 $2.80 / 1M Baseline (1.0x)
GPT-5.6 Sol $5.00 $15.00 $7.00 / 1M 2.5x more expensive
Claude Fable 5 $6.00 $18.00 $8.40 / 1M 3.0x more expensive
DeepSeek V4 Pro $0.14 $0.435 $0.20 / 1M 0.07x (Open-weights budget leader)

Production Scenario: 50,000 Multi-Turn Agent Invocations

Assuming an average task lifecycle of 200,000 prompt tokens (with context caching) and 10,000 generated tokens:

  • Total Workload: 10 Billion Input Tokens, 500 Million Output Tokens.
  • Grok 4.6 Total Bill: $(10{,}000 imes $2.00) + (500 imes $6.00) = $20{,}000 + $3{,}000 =$ $23,000.
  • GPT-5.6 Sol Equivalent Bill: $(10{,}000 imes $5.00) + (500 imes $15.00) = $50{,}000 + $7{,}500 =$ $57,500.
  • Net Infrastructure Savings: $34,500 per 50,000 complex agent sessions.

For engineering teams whose workloads are bounded by budget rather than absolute peak coding scores, Grok 4.6 provides 90–95% of frontier capability at less than half the operating expense.


What We Still Do Not Know

As with any major foundation release, several questions require independent long-term validation:

  1. Third-Party Contamination Audits: The GDPVal and AA-Briefcase numbers originate from initial partner benchmarks. Independent reproductions on contamination-resistant datasets like SWE-bench Pro and Humanity’s Last Exam will establish whether gains persist on fresh tasks.
  2. Context Window Degradation Under High Concurrency: Early API reports show solid low-latency streaming at typical prompt lengths, but throughput degradation on 128k+ token needle-in-a-haystack prompts remains to be verified under production enterprise load.

Sources, Disclosures & Primary Benchmark Data

RankLLMs independent evaluations verify official benchmarks against reproducible testing suites, community logs, and provider documentation.

Share Article

Was this benchmark analysis helpful?

Lucky Yaduvanshi

Lucky Yaduvanshi(luckyyaduvanshi.in →)

Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.

Curated Research

Recommended Reading

Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Explore All 38 Guides→