Luxsynthesis Blog
All articles
Models

Muse Spark 1.1 Launched: Meta Releases a Multimodal Reasoning Model and a New Meta Model API

Meta Superintelligence Labs releases Muse Spark 1.1, a multimodal reasoning model built for agentic tasks with a 1M-token context window and major gains in tool use, coding, and multimodal understanding. A new OpenAI-SDK-compatible Meta Model API launches alongside at $1.25 / 1M input tokens.

Author
Luxsynthesis
Published
Reading time
12 min read
Muse Spark 1.1 Launched: Meta Releases a Multimodal Reasoning Model and a New Meta Model API

Meta Superintelligence Labs (MSL) officially released Muse Spark 1.1, the latest version of the Muse Spark model family and a major upgrade over Muse Spark (1.0).

Meta positions it as a multimodal reasoning model built for agentic tasks, with the biggest gains in three areas: tool and computer use, coding, and multimodal understanding.

Alongside the model, Meta launched the new Meta Model API (public preview), giving developers API access to Muse Spark 1.1 for the first time. The model is already live in the Meta AI app ("Thinking" mode) and on meta.ai.


Core Positioning

Key points from Meta's description of Muse Spark 1.1:

  • A multimodal reasoning model designed for agentic tasks.
  • 1 million-token (1M) context window, which the model can actively manage: remembering actions, retrieving information from earlier work, and preserving the key steps needed downstream during compression.
  • Zero-shot generalization to new native tools, MCP servers, and custom skills.
  • Meta says it advances the performance-efficiency frontier.
  • The product vision is positioned as personal superintelligence.

Muse Spark 1.1 core positioning: a unified intelligent core orchestrating coding, computer use, multimodal understanding, and other capabilities through an organic branching network, while actively managing a 1M-token context


Capability Breakdown

Multi-agent system topology: a main agent orchestrating multiple parallel subagents

Agents

  • Excels at personal-agent tasks that require planning and orchestration across multiple external apps and services.
  • Zero-shot generalization to new native tools, MCP servers, and custom skills.
  • Significantly faster than Muse Spark (1.0) on complex projects. The model is trained to orchestrate multi-agent systems to optimize end-to-end latency.
    • As the main agent: gathers context, forms a plan, and delegates execution to parallel subagents.
    • As a subagent: stays within its remit, understands the available tools, and knows when to escalate back to the main agent.
  • Actively manages a 1M-token context window: remembers actions, retrieves from earlier work, and preserves key downstream steps during compression.

Computer Use

  • Strong on computer-use workflows that span multiple applications and involve real-time information changes; maintains context across long sessions, adapts to shifting requirements, and minimizes human intervention when navigating unfamiliar interfaces.
  • Rather than clicking and reasoning step by step, it decides when to automate and when to use the UI directly: trained to write a script when scripting is faster, click through when direct interaction is simpler, and generate batched actions at each step.
  • Demo case (official video): "Agentic dinner party organization" — real-world scenarios surface new context that changes the task; Muse Spark 1.1 notices these changes while placing the order and updates proactively, without user intervention.

Coding

  • Substantial coding gains on real tasks involving large, complex codebases.
  • Can diagnose and fix complex bugs, implement new features in enterprise-grade systems, and carry out large-scale code migrations.
  • Major improvements over 1.0 on use cases such as building web apps and end-to-end Q&A.
  • Trained to adapt smoothly to a range of harnesses, reliably handling complex multi-turn dynamics and supporting popular agentic coding configurations; common features include planning mode, goal conditioning, subagent delegation, and context compaction.
  • Demo case (Debugging in OpenCode): Muse Spark 1.1 builds a chat web app, automatically screenshots it to identify user-visible failures, traces back to the relevant code to implement and verify a fix — simultaneously invoking coding, multimodal understanding, and tool-use capabilities.
  • Internal use at Meta: Internal Meta developers and researchers use Muse Spark 1.1 daily. On Meta Internal Coding Bench, Meta's primary internal coding benchmark, Muse Spark 1.1 significantly outperforms Muse Spark and is on par with leading competitors. Researchers have also started using it to automate model development and evaluation tasks.
  • Demo case (DeepSWE evaluation in OpenCode): Muse Spark 1.1 self-evaluates on a set of DeepSWE tasks (at varying reasoning effort) and generates an analysis dashboard based on the results.

Multimodal

  • Strong performance on perception, multimodal reasoning, and tool use, with the ability to interact with real environments and produce grounded outputs.
  • Strengths: visual-to-code artifact generation, ultra-descriptive image/video captioning, and executing agent workflows for multimodal use cases.
  • Especially valuable in scenarios where perception and action must happen simultaneously: the model can inspect visual and audio signals, retain details across long workflows, and leverage those details when operating a computer on the user's behalf.
  • Demo case (Facebook Marketplace agent): Given smartphone-recorded video, Muse Spark 1.1 extracts useful photos, reasons about the product, and drives the user's browser to publish a Facebook Marketplace listing on their behalf.

Perception and action occurring simultaneously: a perception core that observes, understands, and autonomously navigates and operates across multiple app interfaces in real time


Safety Evaluation (under the Advanced AI Scaling Framework)

Meta ran an extensive pre-deployment safety review of Muse Spark 1.1 under its Advanced AI Scaling Framework, which defines evaluations, threat models, and deployment thresholds for frontier models.

Overall Conclusion

  • After mitigations, Muse Spark 1.1's residual risk in three frontier-risk areas — Chemical & Biological, Cybersecurity, and Loss of Control — is moderate or lower.
  • Unmitigated: the model reaches the high-risk threshold in Chemical & Biological; in Cybersecurity it cannot be ruled out from hitting the high-risk threshold; Loss of Control remains moderate or lower.
  • Mitigated: residual risk in all three areas drops to moderate or lower, clearing the bar for release.

Frontier risks: unmitigated vs. mitigated

Key Points by Area

  • Chemical & Biological: Capability gains over 1.0, but no new category of risk beyond those already covered in the earlier Muse Spark Safety & Preparedness Report; targeted mitigations (including new API-deployment safeguards) bring residual risk to moderate or lower.
  • Cybersecurity: Stronger than 1.0 on cyber tasks, hitting the high-risk threshold unmitigated; layered mitigations (consistently refusing offensive cyber-abuse information and scalable blocking of sustained malicious use) reduce residual risk to moderate or lower.
  • Loss of Control: Rated moderate or lower; the model shows neither the capability nor the willingness to follow an out-of-control path.

Adversarial Robustness

  • Across content-safety, jailbreak, agent-abuse, and prompt-injection evaluations, robustness improves substantially over 1.0; Meta says it approaches SOTA among comparable models.
  • Third-party red-team validation: compared with other mainstream models, Muse Spark 1.1 carries lower risk and shows more balanced refusal behavior.
  • Meta recommends deploying it with system-level controls (application-layer policy guards, strict tool allow-lists, workspace isolation).

Benchmark Performance

Source: Muse Spark 1.1 Evaluation Report (Table 1 / Table 2 / Table 6). Comparison models were all evaluated via their respective APIs at high reasoning effort; some evaluations omit Claude due to high refusal rates.

Capability (catastrophic-risk-relevant)

AreaBenchmarkMuse Spark 1.1Muse Spark 1.0GPT-5.5Claude Opus 4.8Gemini 3.1 Pro
Chemical & BiologicalMBCT53.254.456.046.2
VCT52.049.752.144.6
HPCT61.955.767.762.9
WMDP-Bio89.088.490.489.5
WMDP-Chem87.085.685.586.3
ProtocolQA88.087.378.288.9
SeqQA (agentic)98.297.398.295.4
ABC Bench (Fragment Design)97.096.892.295.8
ABC Bench (Liquid Handling)93.792.393.299.2
ABC Bench (Screening Evasion)63.254.1
BioDesign Tools (avg)55.239.267.462.4
CybersecurityCybench (pass@1)92.965.4100.095.0
Curated CTFs (pass@1)89.972.0
CyberGym (pass@1)59.043.581.878.8
ExploitGym (pass@1)0.814.81.4
CyScenarioBench (pass@1)0.50.026.016.6
Social Engineering5.11.27.113.7
AIRS-Bench77.086.084.083.0
Loss of ControlSHADE-Arena6.80.57.61.6
GDM-Stealth (of 4)2/41/41/40/43/4
GDM Situational Awareness55.129.360.053.154.5

Table 1 source notes: snapshot of a capability subset; some Chemical & Biological items omit Claude due to high refusal rates. Bold indicates the best score.

Cybersecurity capability: the jump from Muse Spark 1.0 to 1.1

AIRS-Bench (end-to-end research capability, Table 6)

ModelValid submission rate (%) ↑Mean normalized score (≥0) ↑
Muse Spark 1.199.00 ± 2.000.77 ± 0.08
GPT-5.5100.00 ± 0.000.87 ± 0.08
Claude Opus 4.8100.00 ± 0.000.84 ± 0.06
Gemini 3.1 Pro99.00 ± 0.000.83 ± 0.01 *

Score meaning: 1.0 = matches human SOTA, >1.0 = exceeds human SOTA. * Taken from the Muse Spark Safety & Preparedness Report, computed over 10 seeds (others use 5).

AIRS-Bench: end-to-end research capability comparison

Safety / Robustness / Behavioral Tendencies (Table 2)

AreaBenchmarkMuse Spark 1.1Muse Spark 1.0GPT-5.5Claude Opus 4.8Gemini 3.1 Pro
Chemical & Biological (refusal rate)BioTIER97.798.069.294.957.7
Chemical Agents99.899.497.096.496.8
Cybersecurity (abuse ASR)Cyber Misuse Chat (ASR)2.839.142.531.6
Catastrophic Cyber Misuse (ASR)3.918.120.13.1
Poly-Guard Bench (ASR)6.948.428.146.2
Loss of ControlMASK95.789.190.194.440.5
Agentic Misalignment1.147.70.00.052.6
JailbreakStrongREJECT v2 (ASR)0.525.20.54.551.0
FORTRESS (ARS)12.320.216.318.229.8
Agent RobustnessAgentHarm (ASR)3.41.50.020.4
AgentDojo (pass@1 ASR)0.711.91.20.53.3
GraySwan ART (pass@1 ASR)0.36.10.80.11.0
False RefusalsOR-Bench (FRR)4.88.06.73.32.5
Cyber Misuse Chat (FRR)5.711.01.32.20.1
AgentHarm Verified (benign) (FRR)16.521.637.55.7
Model BehaviorSAVE-Bench (coding-agent risk escalation)90.732.285.917.8
Internal Sycophancy49.257.945.532.465.6
DeceptionBench (Honesty)0.10.70.125.5
HLE Calibration (Miscalibration)23.450.344.226.550.4
Alignment Faking (Scheming)-0.21.80.60.214.0

Table 2 source notes: ASR = Attack Success Rate (lower is better); FRR = False Refusal Rate (lower is better); bold indicates best safety performance.

Adversarial robustness: Attack Success Rate (ASR) comparison

Software Engineering / Coding Qualitative Findings (from the report body)

  • On SWE-Bench Verified Hard, Muse Spark 1.1 solved at least 24 of 42 independent tasks at least once, just clearing the capability-checkpoint threshold ("solve more than half of the independent tasks").
  • On coding benchmarks such as Terminal-Bench 2.1 and SWE-Bench Pro, it trails Claude Opus 4.8 and/or GPT-5.5.
  • On long-horizon agentic tasks such as DeepSWE and DeepSearchQA, it improves notably over 1.0 but still lags or matches the best-performing competitors.
  • AIRS-Bench (end-to-end research lifecycle): normalized score of 0.77, below GPT-5.5 (0.87), Claude Opus 4.8 (0.84), and Gemini 3.1 Pro (0.83); the model cannot yet reliably automate critical AI R&D tasks.

General Capabilities Evaluation Scope (Section 5, per-method)

Section 5 of the report covers the following benchmarks, though most scores are presented as figures (see the original figures) rather than enumerated in text:

  • Agents: OSWorld-Verified, OSWorld 2.0, WebArena-Verified, GDPval-AA, JobBench, MCP Atlas, Toolathlon-Verified, DeepSearchQA, WideSearch, Finance Agent v2
  • Coding: Terminal-Bench 2.1, SWE-Bench Pro, DeepSWE v1.1, VibeCodeBench v1.1
  • Health: HealthBench Pro
  • Multimodal: CharXiv Reasoning, BabyVision
  • Reasoning: Humanity's Last Exam (HLE), MRCR v2 (1M-context 8-needle retrieval)

Evaluation setup: Muse Spark 1.1 was evaluated entirely through the Meta Model API at xhigh reasoning effort; comparison models used their respective highest reasoning effort (Gemini=high, Claude=max adaptive, GPT=xhigh).


Pricing and Access

Meta Model API: an OpenAI-SDK-compatible unified developer entry point, backed by web search, visual input, and metered billing

Meta Model API (public preview)

  • Open to developers for the first time, self-serve, OpenAI SDK-compatible.
  • Pricing (pay-as-you-go): $1.25 / 1M tokens input, $4.25 / 1M tokens output.
  • A $20 one-time free credit on sign-up.
  • Built-in web search grounding (web search + citations): add the {"type": "web_search"} tool to any Responses API call to fetch real-time information and produce cited answers.
  • Supports visual input (e.g., base64-encoded PNG).
  • Currently available to US-based developers.

Official Resources


Early Partner Feedback

The blog quotes three early partners:

"What's most impressive about Muse Spark is how much it packs into a single model: million-token context, full multimodal support (images, video, PDF), built-in cited search, strong reasoning, top-tier coding (especially frontend and design), structured output, parallel tool calls — all in a clean OpenAI-compatible package. A complete agent foundation." — Amjad Masad, Replit CEO

"Meta is clearly building for serious agentic coding — strong tool use combined with a price point that makes large-scale real-world coding workloads economically viable. That combination is rare, and exactly what we want Cline developers to get their hands on early." — Saoud Rizwan, Cline CEO

"When testing on our enterprise work evaluation set, Muse Spark demonstrates enterprise-grade capabilities on par with today's leading frontier models. That level of intelligence, combined with its strengths in structured, process-driven workflows across industries (professional services, public sector, industrial operations), makes it very attractive to organizations." — Yashodha Bhavnani, VP of Product, Box AI


References

All articles