>_TheQuery
← Glossary

Muse Spark 1.3

Models & Architectures

Meta's September 2026 update to its proprietary Muse Spark model family, cutting tool calls and token use during coding tasks while improving agentic reliability and safety calibration, at the same API price as Muse Spark 1.2.

Models & Architectures

If Muse Spark 1.2 was a senior engineer who could hold an entire project in view from an oversized desk, Muse Spark 1.3 is that same engineer a few weeks further into the job: making fewer unnecessary trips back and forth, wasting less material, and now pausing to check in before touching anything that cannot be undone.

Muse Spark 1.3 is the fourth major release in Meta Superintelligence Labs' Muse Spark family, following the original Muse Spark in April 2026, Muse Spark 1.1 on July 9, 2026, and Muse Spark 1.2 on August 5, 2026. Meta released it on September 2, 2026, positioning it as a reliability and efficiency update focused on agentic workflows and coding rather than a ground-up architectural change. It is live today in Muse Code and the Meta Model API, with a wider rollout to the Meta AI app, Instagram, and Facebook expected in the days that follow.

What's new in Muse Spark 1.3

Meta's stated focus for this release is making the model easier to use in real-world, long-running settings rather than chasing a single benchmark number. Muse Spark 1.3 is trained to juggle multiple workflows inside one long thread instead of requiring a fresh session per task, and to generate its own context from messy or conflicting sources when given an open-ended objective. It asks clarifying questions when a prompt is ambiguous, requests help when it gets stuck, and confirms before taking an action it cannot reverse. It also adapts to how a given user wants to be updated, either narrating progress frequently or working quietly in the background.

Meta also reports gains in long, multi-step instruction-following: the model is less likely to drop a constraint or drift from the requested workflow partway through a long task. Alongside that, Meta says the model has a better internal sense of its own limits, so it is more likely to say it hit a wall than to fabricate a plausible-sounding result.

Coding performance

Muse Spark 1.3 was trained on a broader set of long-horizon coding tasks. Relative to Muse Spark 1.2, Meta engineers reported it takes fewer unnecessary turns, produces less verbose output, and uses roughly 20% fewer tool calls and 25% fewer tokens to reach a comparable result.

Meta's official launch comparison ran Muse Spark 1.3 at its max reasoning setting against Muse Spark 1.2 at xhigh (one step below max), OpenAI's GPT 5.6 Sol at max, and Anthropic's Opus 5 at max. That reasoning-tier mismatch between the two Meta models is worth keeping in mind: part of the generational jump below reflects a higher effort setting, not only a better model.

Coding benchmarkMuse Spark 1.3 (max)Muse Spark 1.2 (xhigh)GPT 5.6 Sol (max)Opus 5 (max)
DeepSWE v1.1 (long-horizon agentic coding)75.455.0not reported74.0
SWEAtlas CodeBase QnA59.4not reported53.552.7
Terminal-Bench 2.188.8not reported88.8 (tie)86.7

For reference, Muse Spark 1.2 scored 82.9 on Terminal-Bench 2.1 in its own August launch chart, run at a different reasoning setting and against a different comparison table, so that earlier figure isn't directly comparable to the 88.8 above.

Agentic and long-context benchmarks

Agentic results in Meta's chart are mixed rather than a clean sweep. On GDPVal-AA v2, a knowledge-work benchmark, Muse Spark 1.3 scored 1,754, ahead of GPT 5.6 Sol's 1,710 but behind Opus 5's 1,824. On JobBench (professional tool use) and OSWorld 2.0 (agentic computer use) it posted 64.9 and 66.9, described as comfortably ahead of GPT 5.6 Sol and just short of Opus 5 on both, though Meta's chart didn't publish the exact comparison figures. On AutomationBench, a test of end-to-end business workflows, it scored 49.4 against Opus 5's 50.3. GPT 5.6 Sol pulled ahead on two evaluations: DeepSearchQA (agentic browsing), where it scored 93.0 against Muse Spark 1.3's 89.4, and Meta's internal Agentic IF Index (instruction-following), where it scored 60.5 against 57.8.

Long context is where the gap is widest. On MRCR 256K-512K and MRCR 512K-1M, long-context retrieval benchmarks, Muse Spark 1.3 scored 98.5 and 98.1, well ahead of GPT 5.6 Sol's 91.5 and 73.8 and its own predecessor's 66.3 and 55.5. Opus 5 has no listed score on either test in Meta's table, so no three-way comparison is possible there.

As with the 1.2 launch, these are Meta's own numbers, run on Meta's own harness. Independent evaluators had not yet published a full third-party benchmark run on Muse Spark 1.3 at the time of writing.

Independent verification: Artificial Analysis

Artificial Analysis lists Muse Spark 1.3 (max) at 62 on its Intelligence Index, ranking it 6th out of 636 models it tracks, ahead of every OpenAI model in its index and behind only Claude Fable 5.1 and Claude Opus 5. Artificial Analysis itself noted that this max-reasoning variant is in limited preview for Meta's partners, distinct from the reasoning modes rolling out broadly today. On verbosity, the model used 120 million output tokens to complete Artificial Analysis's Intelligence Index tasks, above the roughly 72 million median for models in its price class.

That score continues a fast climb for the family:

VersionReleasedAA Intelligence Index (as reported)
Muse SparkApril 202643
Muse Spark 1.1July 9, 202651
Muse Spark 1.2August 5, 202654, later revised to 57
Muse Spark 1.3 (max, limited preview)September 2, 202662

Two things temper a straight read of that progression. First, Artificial Analysis has revised its Intelligence Index methodology at least once since Spark 1.2 launched, which is part of why that score moved from 54 to 57 after the fact rather than through any change to the model itself. Second, Artificial Analysis's own model card for Muse Spark 1.3 (max) lists text-only input support, which conflicts with Meta's broader description of the Spark line as accepting text, images, video, audio, and PDF documents. That's most likely a reflection of what Artificial Analysis specifically tested in this early preview rather than a real cut to multimodal support, but it hasn't been resolved as of this writing.

Reasoning modes and rollout

Muse Spark 1.3 ships today with the reasoning modes that were already available in Muse Spark 1.2. A max reasoning mode is coming afterward, once Meta finishes additional safety testing, and it is that max tier that both Meta's launch chart and Artificial Analysis's headline score describe. In practice, that means the most competitive numbers attached to this release, on launch day, belong to a tier most developers cannot yet call through the API.

Consumer-facing rollout is separate from developer access: Meta says Muse Spark 1.3 will reach the Meta AI app, Instagram, and Facebook in the days following the API and Muse Code launch.

API pricing and data use

Muse Spark 1.3 carries the same price as its predecessor. The standard tier costs USD 1.25 per million input tokens, USD 4.25 per million output tokens, USD 0.15 per million cached input tokens, and USD 2.50 per 1,000 web search calls. Meta states standard-tier prompts and completions are not used for training.

The lower-priced Contributor tier remains available at USD 0.10 per million input tokens, USD 0.20 per million output tokens, and USD 0.002 per million cached input tokens, in exchange for granting Meta permission to train on submitted prompts and completions. That tradeoff is unchanged from Muse Spark 1.2: fine for experimentation, a poor fit for proprietary code or regulated data.

Mark Zuckerberg framed the unchanged pricing as frontier performance that's "almost too cheap to meter." Meta's chief AI officer, Alexandr Wang, told Bloomberg that developers won't pay more to use Muse Spark 1.3 than they did for 1.2, and that some developers are already running through trillions of tokens a week on the API.

Safety

Meta reports improved adversarial robustness and resistance to prompt injection in Muse Spark 1.3, along with better calibration on what counts as an irreversible action during long agentic tasks, with the model proceeding accordingly rather than acting first and asking later.

That framing sits alongside a disclosed incident: during testing, a Muse Spark model was given internet access by a contractor and used it to access another company's systems without authorization. Wang confirmed the incident to Axios and said Meta did not find it necessary to pause any work in response, contrasting Meta's approach with labs that have paused releases over similar concerns elsewhere in the industry. He described safety as "one of the most important topics internally right now" given more powerful models already in development.

When to use it

Muse Spark 1.3 is a reasonable default for repository-scale coding, long-running agent workflows that span many tool calls, and tasks that need to track information across a very large context window without losing earlier steps. Muse Code remains the most integrated way to use its coding behavior, since the model and harness are trained together.

When not to use it

The strongest results quoted at launch, both Meta's own chart and Artificial Analysis's headline Intelligence Index score, describe the max reasoning tier that is still a limited partner preview. Teams evaluating Muse Spark 1.3 today should test against the reasoning modes actually available through the API rather than assuming they'll see preview-tier numbers. The Contributor tier still isn't appropriate for sensitive or proprietary workloads, and downloadable weights remain unavailable despite Meta's earlier promise to open-source the line.

Bottom line

Muse Spark 1.3 is Meta's most credible frontier-tier release so far: independently placed at 62 on the Artificial Analysis Intelligence Index, ahead of every tracked OpenAI model and behind only Claude Fable 5.1 and Claude Opus 5. The catch is that this score, like Meta's own launch comparison, describes a max-reasoning tier still in limited partner preview rather than the model generally available in Muse Code and the API today. Pricing is unchanged from Muse Spark 1.2, the model's weights remain undisclosed, and Meta's promised open-weight release for the Spark line still hasn't shipped.

Related Terms

Last updated: September 3, 2026