Artificial Intelligence
●13 min read●October 5, 2026

Meta Muse Spark 1.3: Pricing, Benchmarks & API Guide

The cheap tier has a data-sharing tradeoff. Max reasoning needs Standard. Here is what to check, with current sources and coding examples you can open.

Paras Tiwari
Paras TiwariFounder, Spectrum AI Labs
Official blue Muse symbol on a white tile against a pale blue background

Get weekly AI tool reviews

We test tools so you don't have to. No spam.

TL;DR

Muse Spark 1.3 is Meta's current hosted reasoning model for coding and agent work, with a million-token context window. Contributor is cheap at $0.10 input and $0.20 output per million tokens, but it allows training on your content and rules out confidential or personal inputs. Standard costs more and unlocks max reasoning. Public builds are worth opening; independent benchmarks show enough variation that I would test it on a few real tasks before switching an entire workflow.

Muse Spark's pricing makes a small experiment easy to justify. Choosing the wrong tier can make the same experiment a bad idea. A private repository, customer emails and a disposable game prompt should not all go through the cheapest model ID.

This guide covers Muse Spark 1.3 as of October 5, 2026. We checked Meta's documentation, independent evaluations and original developer reports. We also opened and briefly played one published game. We did not run a paid Spark comparison or reproduce the builders' generation times.

What is Meta Muse Spark 1.3?

Muse Spark is the model; Muse Code is the coding agent; Meta Model API is the service that hosts it. That distinction matters when somebody says they tried “Muse.” A consumer chat, a terminal agent and a custom API integration can use different settings and data policies.

Meta released Spark 1.3 on September 2. The current model catalog recommends it for new work. It accepts text, images, video and PDFs, and returns text. Its documented context budget is 1,048,576 tokens. Meta's long-context guide says input and requested output share that limit.

There is an easy-to-miss footnote: audio understanding is not fully supported in 1.3. Meta recommends Spark 1.2 or Muse Voice Transcribe for audio work. Muse Image is another separate model; Spark itself is not the image generator.

Official Muse Spark model table with input modalities, text output, context limit and the audio-support footnote
Meta’s current model catalog, captured October 5, 2026. The footnote warns that Spark 1.3 audio understanding is not fully supported. Meta model catalog

In the 1.3 announcement, Meta says its engineers saw roughly 20% fewer tool calls and 25% fewer tokens than 1.2 in internal comparisons. Fewer unnecessary steps would be useful in a coding agent. Those figures belong to Meta's workload, though. They do not tell us what your repository will cost or how many bugs it will leave behind.

The API is still labeled public preview. Access has expanded, but supported locations and product rules still apply. Check the geographic policy before building for users in another country.

Muse Spark 1.3 pricing: Standard vs Contributor

Meta's direct API rates, checked October 5, are in US dollars per million tokens:

TierUncached inputCached inputOutput
Standard$1.25$0.15$4.25
Contributor$0.10$0.002$0.20
Meta API documentation showing Standard and Contributor token prices and their training-data difference
Meta’s Standard and Contributor rate cards, captured October 5, 2026. Contributor’s discount permits training on prompts and completions. Meta API pricing

The difference extends beyond price:

  • Standard: prompts and completions are excluded from model training. Spark 1.3 supports max reasoning. The published team limit is 3,000 requests per minute.
  • Contributor: Meta can use content for training. Personal, confidential and sensitive inputs are prohibited. Max is unavailable, and the team limit is 100 requests per minute.

Those are team limits, shared across keys. The same pricing page lists no long-context surcharge. Built-in web search adds $2.50 per 1,000 queries. Higher reasoning effort also uses billable output tokens, even when the final answer is short.

A small bill example

Suppose a run uses 100,000 uncached input tokens and 10,000 output tokens, including reasoning. With no search calls, the arithmetic is $0.1675 on Standard or $0.012 on Contributor.

The Standard calculation is 0.1 × $1.25 + 0.01 × $4.25. If 90,000 input tokens qualify for caching, its total falls to $0.0685. Ten search queries add another $0.025.

These are example token bills. We have not measured a coding task that uses exactly those amounts, and the tiers cannot be compared at max because Contributor does not offer it. Retries, subagents and repeated context can easily matter more than the neat little example.

What about the $5 Muse Code plan?

Meta's Muse Code subscription page lists Everyday at $5/month, High at $15 and Power at $50. Everyday advertises 10–50 prompts per five hours, with usage depending on the task; availability and benefits vary by region.

The subscription credential works in the Muse Code CLI. Separate API-key traffic is still pay-as-you-go. Do not budget an app backend as if the $5 plan covers its API requests.

Read the data terms before connecting private code

For a public toy project, Contributor's discount may be useful. For client work, an internal repository or a document containing personal information, the Contributor restrictions make that tier unsuitable. This is an explicit input restriction, not just a checkbox about improving the model.

Standard's no-training commitment is useful, but Standard does not automatically mean zero data retention. Meta's October 2 terms still allow processing and retention for stated service, legal, safety and policy purposes.

Zero data retention requires separate approval for a qualifying organization. It applies to direct calls with that organization's keys, disables features including stored conversation continuation, background requests, uploads and web search, and has exceptions for policy-flagged content. A third-party provider has its own arrangements too.

Before committing a product to it, read the service terms and acceptable-use rules. They include restrictions on competing-model training and public benchmarking used to promote a competing service. This is a summary of the published terms, not legal advice or a complete compliance review.

Is Muse Spark 1.3 good at coding?

There is credible coding evidence. There is also a large gap between passing a short repository task, finishing a long terminal job and producing a nice-looking demo.

What Meta reports

These selected results come from Meta's current comparison page, checked October 5. The Spark versions use different reasoning efforts:

TestSpark 1.3 maxSpark 1.2 xhighWhat it measures
DeepSWE v1.175.455.0Repository software tasks
SWEAtlas Codebase QnA59.446.2Questions about a codebase
Terminal-Bench 2.188.882.9Terminal tasks in the reported setup
OSWorld 2.0, partial credit66.947.6Progress through computer workflows
OSWorld 2.0, binary success32.017.9Fully completed computer workflows

Not sure which AI model to use?

21 models · Personalized picks · 60 seconds

Take the Quiz
Meta’s Muse Spark 1.3 scorecard with agent, long-context and coding results, including separate partial and binary OSWorld scores
Meta’s official release scorecard, checked October 5, 2026. It compares 1.3 max with 1.2 xhigh and release-era competitors, using differing sources and harnesses. Open the image for the full-size chart. Meta’s 1.3 announcement

The coding rows report task-pass or mean pass@1 results on a percentage scale. OSWorld uses both mean partial credit and strict completion. The OSWorld rows are worth reading together. A 66.9 partial-credit score does not mean that 66.9% of workflows finished successfully; the reported binary score is 32.0.

Meta's methodology PDF describes a mixture of its own runs, provider results and public leaderboards. Harnesses vary, and the OSWorld version for Spark 1.2 differs from the others. These numbers support progress in the tested configurations, not one perfectly controlled ranking of every model.

Independent results are more useful when you keep the test version

On October 5, Artificial Analysis shows Spark 1.3 max at 48 on its Intelligence Index and $1.60 weighted cost per index task. Its xhigh page shows 45 and $1.37. The score is a composite, not the percentage of your tasks it will finish.

If you remember a launch score of 62, you are remembering a different snapshot. AA's September 2 article reported that number; its current methodology is v4.3.2, after changes to the tasks, grading and baselines. Comparing 62 then with 48 now would tell the wrong story about the model.

A more demanding coding check comes from Vals' Terminal-Bench 4.0 results, updated October 1. Spark 1.3 max completed 24.75% of tasks at a reported $6.65 per test. Vals uses 66 new tasks, mini-swe-agent and three runs averaged. Meta's 88.8 on Terminal-Bench 2.1 uses a different task set and setup, so subtracting the two scores would be meaningless.

My takeaway is fairly simple: there is enough here to justify an evaluation for coding. There is no reason to skip your tests because a launch chart looks good. If you are choosing across model families, use our task-by-task model guide as a starting point, then test the exact model, effort and agent you plan to use.

What developers have actually built with Muse

The most useful examples have something you can open, source code you can inspect, or a failure with enough detail to learn from.

A one-file Asteroids game you can play

Dustin Davis published an Asteroids prototype on September 3, made with Spark 1.3 through OpenCode Zen's free route. He reports one prompt, no follow-up edits and 26.3 seconds. The post gives the prompt and links the playable result.

We opened that result on October 5. Starting the game, moving, firing, increasing the score and pausing all worked in a brief desktop-browser check. We did not generate it ourselves or verify the author's timing. Audio, touch controls and every collision edge case were outside that check.

Dustin Davis's published Asteroids prototype paused after the score reached 20 during a brief browser check
The published Asteroids game during our October 5 check. This is a screenshot of the existing demo, not a model run performed for this article. Play Dustin's Asteroids

Aakib Ansari's September 9 Bot Arena write-up is useful for a different reason. His Spark 1.3 xhigh output had bots that moved but did not shoot. A follow-up fixed firing and added difficulty levels; he still reported occasional overlapping bots. He published prompts, code excerpts and an embedded demo. We read the write-up rather than play-testing that game.

MineBench counted the rejected attempts too

MineBench's merged Spark 1.3 integration is a better piece of evidence than an unlinked claim that a model “built Minecraft.” It connects Spark to an existing voxel-building benchmark. The recorded run reports 15 finalized builds, 53 provider calls and 28 completed attempts, including 13 rejected responses.

The PR reports $6.57 across the covered attempts. That is the project's reported run cost, not our measurement or a recalculation at today's rates. MineBench lists AI-lab and API-credit support; this is a project record, not an unaffiliated comparison we conducted. Fifteen finished outputs took more than fifteen tries, which is exactly the detail I want before treating a low token price as a low project cost. The release and live project are public; we did not rerun the paid benchmark.

A real code-review integration, and a real repair trail

The open-source pair-review app added Muse Code as an analysis provider in August and moved its built-in models to 1.3 in September. Its provider implementation uses muse exec --json, with unit tests alongside it. This gives developers a released integration to inspect. It does not tell us how often its reviews find real bugs, or mean Muse wrote the application.

The Hermes issue opened September 5 shows the other side. A user running Spark 1.2/1.3 Contributor through OpenCode Go reported unrelated fragments at the end of tool-heavy sessions. Maintainers traced one trigger to conversation replay, landed repairs and closed the issue on September 19, while documenting cases their detector deliberately leaves alone.

That is a dated integration failure and repair, not evidence that every current Spark session has the same bug. It is a good reason to inspect the provider route and conversation adapter when an agent starts behaving strangely.

How to use the Muse Spark API

Meta's developer overview offers three routes: direct API calls, Muse Code, or a compatible third-party coding agent. The base URL is https://api.meta.ai/v1. Standard uses muse-spark-1.3; Contributor uses muse-spark-1.3-contributor.

The API has Responses, OpenAI-compatible Chat Completions and Anthropic-compatible Messages interfaces. Start with Meta's quickstart, create a key in the official console and keep it server-side in an environment variable. Never paste it into a browser app or commit it to your repository.

This small Responses example follows the documented request shape. We have not executed it for this article. A successful request can incur token charges; the output budget includes internal reasoning.

curl https://api.meta.ai/v1/responses \
  -H "Authorization: Bearer $MODEL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "muse-spark-1.3",
    "reasoning": {"effort": "low"},
    "max_output_tokens": 1024,
    "input": "Reply with one short sentence confirming you received this message."
  }'
Meta reasoning documentation showing effort levels, billed reasoning tokens and the Standard-only max restriction
Meta’s reasoning reference, captured October 5, 2026. Max is limited to Standard Spark 1.3; reasoning tokens count toward the billed output budget. Meta reasoning documentation

Before moving from a greeting to an agent, check four things:

  1. Choose a supported reasoning level. Minimal through xhigh are available; max requires Standard 1.3. Setting none returns a 400 error.
  2. Preserve conversation state correctly. Responses supports previous_response_id or encrypted reasoning replay. External Chat Completions calls cannot carry private reasoning across turns. Copying only visible text between endpoints changes the behavior of a long-running agent.
  3. Budget for reasoning and tools. A 1,024-token connection check is not a sensible limit for a repository-wide change. Give a real job enough output room, with a spending limit and a stopping condition.
  4. Keep tool permissions narrow. A model's request to run a command is only part of the system. Your agent executes it. Use a disposable branch or sandbox and require review before deployment or destructive actions. Our MCP guide explains the separate tool layer.

If you want a ready-made terminal agent, follow the official Muse Code guide. Check the selected model explicitly: parts of the CLI guide still name 1.2 as the default. Its tools, approval behavior and handling of long sessions matter as much as the model name. Substituting Spark into another agent will not automatically reproduce a Muse Code demonstration.

Is Muse Spark open source, and can you run it locally?

Spark 1.3 is currently a hosted model in the sources we checked. Meta's September announcement still puts a Spark open-weight release on the roadmap. We found no released Spark 1.3 checkpoint or weight license to point readers to, and no verified release date for one.

Meta has separately released Muse Glimmer, a 30B open-weight model distilled from Spark with an Apache 2.0 license. Glimmer's size and architecture should not be quoted as Spark's specifications. It is a different model to evaluate for local use.

If downloadable weights are a requirement, our MiMo V2.6 guide covers Xiaomi's available checkpoints and the difference between a model's total and active parameter counts. Download availability alone does not mean the full model will fit on a laptop.

What I would test before switching a coding workflow

Take three tasks you already know how to judge: a small bug fix, a change spanning several files, and a feature with a visible UI behavior. Write the acceptance checks first. Keep the starting repository, tools and permissions the same across models.

For each attempt, record the tier, provider, reasoning effort, agent version, passed checks, total bill, elapsed time and manual corrections. Count abandoned runs as well. Start on non-sensitive material, and compare more than one attempt before making a decision.

I would try Contributor for disposable, non-sensitive prototypes when its training terms are acceptable. I would evaluate Standard for private work only after checking the applicable data requirements, and compare high or xhigh with max rather than assuming the largest setting is always worth it. For audio-heavy work, follow Meta's current recommendation to use a different model.

The result I want is a patch I can accept with little correction. A cheap failed run is still work somebody has to redo.

Sources and what we checked

Prices, model details and terms were checked on October 5, 2026. The three documentation screenshots are actual captures of Meta pages. The benchmark chart is Meta’s original release image. Each has a source link and can be opened at full size. The text-free banner uses the official Muse product symbol; it is not a separately announced Spark-only logo.

  • Product and access: Meta's 1.3 announcement, model catalog, pricing and reasoning reference
  • Independent evaluations: Artificial Analysis' dated model pages and methodology, plus Vals' Terminal-Bench 4.0 results. Scores and costs are attributed to their specific test configurations
  • Builds and experiences: original author posts, public code, release records and issue discussions. Only the published Asteroids demo received a brief browser interaction check; generation claims and benchmark runs were not independently reproduced

FAQ

How much does Muse Spark 1.3 cost?

As checked October 5, 2026, Standard costs $1.25 per million uncached input tokens and $4.25 per million output tokens. Contributor costs $0.10 and $0.20 respectively, permits training on content, and excludes max reasoning. Cached input costs less; reasoning and search can add to the bill.

Is Muse Spark 1.3 open source?

Meta still describes Spark open weights as a future release in the sources checked October 5, 2026. Spark 1.3 is currently a hosted model. Muse Glimmer is a separate downloadable model distilled from Spark, not a local copy of Spark 1.3.

Does Meta train on Muse Spark API data?

Standard prompts and completions are excluded from Meta model training. Contributor permits training and prohibits personal, sensitive and confidential inputs. Standard is not automatically zero retention; separate organization-level approval and feature restrictions apply to ZDR.

What is the difference between Muse Spark and Muse Code?

Muse Spark is the model. Muse Code is Meta's terminal coding agent, which gives the model tools to work in a project. Meta Model API serves the model to Muse Code and other applications. The agent's tools, permissions and conversation handling affect its results.

Paras Tiwari
Written by
Paras Tiwari
Founder, Spectrum AI Labs

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.

More about Paras →

Stay ahead of the AI curve

We test new AI tools every week and share honest results. Join our newsletter.