Muse SparkMeta AIAI CodingAI AgentsNews

Muse Spark 1.2: Technical Specifications, Pricing, and Context Length

Technical review of Meta Muse Spark 1.2: 1M context window, context compaction, 82.9% Terminal-Bench score, token pricing, and co-training with Muse Code.

Lucky Yaduvanshi
Lucky Yaduvanshi
Founder & AI Lead
Aug 10, 2026•Updated Sep 24, 2026•4 min read•Loading views...
Independent technical benchmark • Primary data & verified methodology cited below
Meta Muse Spark 1.2 Model Specifications Architecture and Benchmark Analysis

Meta’s Muse Spark 1.2 is a specialized foundation model engineered specifically for autonomous software engineering, repository-scale refactoring, and multi-step agent trajectories. Launched on August 5, 2026 alongside the Muse Code terminal agent, Muse Spark 1.2 was co-trained directly with its agent runtime, incorporating context compaction, goal conditioning, and subagent worktree orchestration into the underlying weights.

On standardized benchmarks, Meta reports Muse Spark 1.2 achieving 82.9% on Terminal-Bench 2.1, surpassing Claude Sonnet 5 (80.4%) by +3.1% relative and scoring an industry-leading 90.3% on MCP-Atlas tool execution. Explore where Muse Spark ranks in our AI Model Leaderboard and Full LLM Model Directory.

Muse Spark 1.2 vs Claude Code Benchmark Comparison

Executive Summary & Core Specification Matrix

Specification Muse Spark 1.2 Specification
Developer Meta Superintelligence Labs
Release Date August 5, 2026
Primary Domain Focus Autonomous software engineering, CLI agents, multimodal code
Context Window 1,048,576 tokens (1M Standard)
Standard Input Price $1.25 per 1M tokens
Cached Input Price $0.15 per 1M tokens
Standard Output Price $4.25 per 1M tokens
Contributor Tier Rates $0.10 input / $0.20 output (Activity data shared)
Multimodal Perception Native Text, High-Resolution Image, Video
Tool Calling & MCP Native function calling, 90.3% MCP-Atlas score
Primary Runtime Meta Muse Code, Meta Developer API

Architectural Advancements: Model-Harness Co-Training

The critical technical differentiator in Muse Spark 1.2 is co-training with the Muse Code runtime:

  1. Context Compaction: Rather than carrying an ever-expanding conversational history that degrades attention fidelity, Muse Spark 1.2 was trained to compress past tool executions, bash outputs, and compiler traces into compact state representations.
  2. Subagent Worktree Coordination: The model understands Git worktree isolation, allowing it to delegate sub-tasks to parallel agent workers without generating conflicting file modifications.
  3. Continuous Execution Persistence: In an automated optimization experiment, Muse Spark 1.2 executed over 1,000 continuous tool calls across a 24-hour window, writing, profiling, and optimizing custom Triton GPU kernels on NVIDIA Hopper hardware without state collapse.
graph TD
    A["Developer Goal"] --> B["Muse Code Runtime"]
    B --> C["Muse Spark 1.2 Engine"]
    C --> D["Context Compaction & Planning"]
    D --> E["Tool Calls & Subagent Execution"]
    E --> F["Repository Diff & Test Suite"]
    F --> G["Final Code Verification"]

Benchmark Comparisons: Terminal-Bench 2.1

Meta evaluated Muse Spark 1.2 against top frontier models on Terminal-Bench 2.1:

Model Terminal-Bench 2.1 Score Leaderboard Rank Margin Analysis vs Muse Spark 1.2
Claude Opus 5 Max 86.7% #1 Opus leads by +3.8 pts (+4.6%)
Muse Spark 1.2 82.9% #2 Frontier Baseline
GPT-5.6 Terra 81.8% #3 Muse Spark leads by +1.1 pts (+1.3%)
Grok 4.5 High 81.6% #4 Muse Spark leads by +1.3 pts (+1.6%)
Claude Sonnet 5 80.4% #5 Muse Spark leads by +2.5 pts (+3.1%)
Gemini 3.6 Flash High 78.9% #6 Muse Spark leads by +4.0 pts (+5.1%)
Muse Spark 1.1 76.2% #7 Muse Spark 1.2 leads by +6.7 pts (+8.8%)

Methodological Note: Scores represent vendor-reported evaluations. Terminal-Bench evaluates live command-line tool execution, flag usage, and error recovery in real bash environments. For standard benchmarks, see our SWE-bench Verified Guide.

Token Economics: Standard vs Contributor Tier

Meta provides two commercial access tiers through the Meta Developer API:

  • Standard Developer Tier: $1.25/M input and $4.25/M output, offering standard zero-retention enterprise privacy terms.
  • Discounted Contributor Tier: $0.10/M input and $0.20/M output (a 92% input discount and 95% output discount). In exchange, Meta receives telemetry and interaction data to train future models.

For independent developers, open-source projects, and exploratory scripting, the Contributor tier provides exceptional value. For proprietary corporate codebases, standard rates or self-hosted models like Muse Glimmer remain required.

Final Verdict

Meta Muse Spark 1.2 represents the vanguard of harness-co-trained coding models.

By pairing an 82.9% Terminal-Bench score with 1M context length, context compaction, and $1.25 / $4.25 pricing, Meta delivers an engine purpose-built for the long-horizon autonomous workflows that define modern software engineering.

Read our complete How to Use Meta Muse Code Guide and track real-time benchmark scores on our AI Model Leaderboard.

Sources, Disclosures & Primary Benchmark Data

RankLLMs independent evaluations verify official benchmarks against reproducible testing suites, community logs, and provider documentation.

Share Article

Was this benchmark analysis helpful?

Lucky Yaduvanshi

Lucky Yaduvanshi(luckyyaduvanshi.in →)

Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.

Curated Research

Recommended Reading

Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Explore All 38 Guides→