Muse Code and Muse Spark 1.2

(research.meta.ai)

159 points | by paulkrush 5 hours ago

35 comments

  • andai 1 minute ago
    The most interesting thing here is the kernel optimization graph.

    It look like all models were still improving, when they cut off the experiment.

    It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.

    The only difference between the models seems to be how quickly they arrive.

  • tristanj 4 hours ago
    Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.

    https://developer.meta.com/ai/models/muse-spark/

    • moonu 26 minutes ago
      I've been surprised by the reception to this, as OpenAI, for a while now, has had free API usage when data sharing is enabled (https://help.openai.com/en/articles/10306912-sharing-feedbac...)
      • OsrsNeedsf2P 6 minutes ago
        I tried following this page, and it's certainly a lot more complex than what Meta is offering. Different price tiers, opt-in configurations, usage based availability.. I'll take the 10x discount for flipping a param switch over this all day long.
    • jjcm 2 hours ago
      I actually really like that pricing strategy. It's very transparent
      • bdangubic 1 hour ago
        I love the idea of this pricing strategy but there is no way meta is not training on your data regardless of your monthly invoice
        • simonw 7 minutes ago
          So you think the only difference between the $1.25/million token plan and the $0.10/million token plan is that you pay them more to both lie to you and breach their contractual obligation to you?
    • ray_kay777 3 hours ago
      This makes it a very interesting alternative to Deepseek for personal work where I don't care about the training - judging by the AA benchmarks it seems like overall cost per task is similar to the new Deepseek Flash but with better benchmarks (and inbuilt vision capabilities).
      • embedding-shape 37 minutes ago
        Except with one you know they'll release the weights and architecture back to the community, with the other, it leans towards they won't do that.
    • GodelNumbering 4 hours ago
      I think that's a fair offering tbh
    • dan15 2 hours ago
      Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)
      • dghlsakjg 1 hour ago
        FYI: there are providers of deepseek that offer the same or lower pricing and zero retention policies.
        • greyb 1 hour ago
          Unfortunately, none with the same caching performance as DeepSeek proper.
      • ForHackernews 2 hours ago
        DeepSeek is really crazy cheap, though, and they don't have a giant pool of other invasive personal data to correlate it with.
        • jofzar 2 hours ago
          I hate to say this and this is because I fucking despise meta. But between DeepSeek and Meta, and trust they handle the training data correctly, I trust meta.
          • aand16 1 hour ago
            What do you mean by "correctly"?
            • handfuloflight 11 minutes ago
              For example, OpenCode says they have a ZDR with DeepSeek. Some of us are skeptical that's going to be properly honored. There's no way to know.
    • HDBaseT 2 hours ago
      Meta, please offer this on OpenRouter too (ZDR + Non-ZDR, official Meta Provider).
    • gigatexal 2 hours ago
      Yikes that’s compelling pricing.
    • deno 4 hours ago
      I think it's limited to US or at least EU is excluded.
  • WhitneyLand 4 hours ago
    They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

    They left Opus in and got beat in all but one benchmark.

    Nothing wrong with trying to improve, but why the marketing games?

    Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.

    Then when your ready, come back and talk frontier without playing hide the model.

    • spmurrayzzz 1 hour ago
      Given the current throughput figures on OpenRouter (~180 tk/s), its likely a much smaller param count on the order of something like Luna. I think the better, more timely comparison (re: your point on Chinese labs) would be to DeepSeek-V4-Flash-0731.

      It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.

      (They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)

    • ac29 2 hours ago
      > They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

      If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2

    • jjice 3 hours ago
      While I won't take their limited benchmarks with much salt, if it actually is this close to opus, but at a third the cost, that's pretty solid. Now, Terra is pretty damn affordable too and you're right that it's suspicious that they don't put Sol in there at all.
    • krm01 4 hours ago
      We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
      • deepsquirrelnet 1 hour ago
        If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.

        Cherry picking the benchmarks you present is where the falsehoods lie.

      • lacker 3 hours ago
        My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
        • villish 24 minutes ago
          Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.

          Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..

          You have to test each model on your actual use case to see how well it really performs.

        • nrub 3 hours ago
          Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
          • brokencode 3 hours ago
            Or they did try to game the benchmarks and just didn’t do it well enough.

            Benchmarks are one data point, not the only one, but the easiest one to compare.

            • nrub 1 hour ago
              Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.
  • simonw 8 minutes ago
    It can cyberattack other companies, too: https://www.cnn.com/2026/08/05/tech/meta-ai-hacking
  • bradfa 4 hours ago
    If you got the $20 in free credits from Meta for signing up when muse-spark-1.1 was release, please note that there's now small print stating "While using free credits your content may be used for product improvement" which was not present at muse-spark-1.1 launch when the credits were given out.

    If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.

    • giancarlostoro 4 hours ago
      The API costs for the version of their model that feeds things back to meta is also drastically lower.
  • wiradikusuma 58 minutes ago
    Hey guys I'm just wondering. Usually when someone announces a new model, they'll show you some fancy viz/video/images: "These are what my model can produce." I'm wondering if anyone is keeping track of these? Like in a gallery form, "Use this prompt to produce this output".

    By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.

  • drivebyhooting 36 minutes ago
    There’s no way I’m giving Zuck any of my data.
  • mchusma 5 hours ago
    This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.
    • handzhiev 4 hours ago
      If you are happy to share data for training, the contributor mode offers amazing price $0.10 / $0.20
      • mchusma 3 hours ago
        Yes, that is the really compelling thing here IMO. Its a viable deepseek competitor for many people, and I missed that on the first pass.
  • conradkay 5 hours ago
    https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg

    Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?

    Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data

  • wxw 4 hours ago
    Last I heard, everyone at Meta was using Claude Code.

    Any insiders know how Muse Code is doing internally?

    • paxys 14 minutes ago
      Meta lets engineers use the best tools for the job. I doubt anyone internally is going to be rushing to switch from Claude Code or Codex.
    • dxxmxnd 3 hours ago
      Everyone is still using claude or codex if they aren’t forced off of it. Nobody is going to use a worse tool in this culture.
      • baby 2 hours ago
        it's actually interesting that they're not being forbidden to use claude/codex, is Meta paying for it or is it personal accounts?
    • GodelNumbering 4 hours ago
      If there were, do you believe it would be in their interest to answer this publicly?
      • georgemcbay 4 hours ago
        > > Any insiders know how Muse Code is doing internally?

        > If there were, do you believe it would be in their interest to answer this publicly?

        If it were being adopted like gangbusters in their organization, sure!

        So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...

    • youre-wrong3 4 hours ago
      [dead]
  • ipsum2 5 hours ago
    I wonder why they didn't compare with GPT-5.6-sol, only Terra?
    • wmf 5 hours ago
      Clearly they're positioning it as a mid model.
      • minimaxir 5 hours ago
        Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive).

        Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.

        • ukblewis 2 hours ago
          I don’t know where you get your statistics, but I love Terra and use it all of the time. It is the default fastest model in ChatGPT/Codex today. I saw today a notice saying that the model had hit capacity briefly
      • redox99 4 hours ago
        But why include Opus then?
    • woadwarrior01 5 hours ago
      Haven't you seen the kernel optimization case study at the bottom of the page? They compare against GPT-5.6 Sol and their model is worse.
    • logicchains 5 hours ago
      Presumably because it's worse than Sol, same reason they compared it to Opus 5 not Fable.
    • Handy-Man 5 hours ago
      Their bigger model is not ready - watermelon code name was still being prepared for release as of a month ago
  • simonw 3 hours ago
    Here's the Muse Spark 1.2 pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

    I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/

    • meric_ 3 hours ago
      Lol is that a helmet? Clearly they're taking the Anthropic path and are safety pilling their models
    • Marciplan 2 hours ago
      I wonder if at this point the labs ha.. please Simon, stop this bs everytime. Its time.
  • daemonologist 4 hours ago
    Pricing: https://dev.meta.ai/docs/pricing-rate-limits

    Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.

  • IceWreck 4 hours ago
    Ive been poking with the muse code binary - seems to be written in rust, looks similar to codex but either its a very hard fork (i also see dissimilar things like config format is different, no acp, etc) or is just heavily inspired by it (more likely).
  • liviux 5 hours ago
    Does this muse code have any muse spark 1.2 usage included? Can't understand from the docs.
  • Bolwin 5 hours ago
    > Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access

    Wasn't the previous one us only? This is probably the biggest part of the post

    Anyone know if muse code is open source?

  • sarjann 4 hours ago
    I do think some of features in their harness seem interesting (workers in separate worktrees at once), recovery from crashes seem interesting.
  • Cappybara12 4 hours ago
    Is this becoming a race where we have a usual flow of a company .. AI models, Coding agents, image generation tools, and more AI models ?
  • kcb 5 hours ago
    Open the weights.
  • king_crimson 4 hours ago
    Why does every AI lab feel the need to build their own coding agent…? Don’t we have more than enough already?
    • HDBaseT 2 hours ago
      Outputs are a little bit more deterministic if you control the harness.

      It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.

    • conception 2 hours ago
      Telemetry, marketing
  • batuhandumani 1 hour ago
    Why should I leave Claude or GPT and switch to Meta's aMUSEment model?
  • dilyevsky 2 hours ago
    the soak tub in the kitchen was nice
  • paulkrush 5 hours ago
    Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.
  • fcoury 5 hours ago
    Interesting, it seems like their muse code is built upon Codex CLI?
  • AtlanticThird 4 hours ago
    I wish they would add a ZDR endpoint on OpenRouter
  • alex1138 1 hour ago
    Fun, this is currently on the front page at the same time this https://news.ycombinator.com/item?id=49187977 is
  • giancarlostoro 4 hours ago
    Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary.

    I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.

    • greyb 4 hours ago
      I honestly think they're kinda banking on piggybacking off of Facebook account integrity systems to avoid the problems that other LLM providers are facing in trying to prevent mass free trial signups for token relays and so forth.

      It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).

    • sunaookami 3 hours ago
      I could sign in with my Meta account that is independent and not linked to IG or Facebook. Just click "Login with Email" on dev.meta.ai.
    • hahahaa 4 hours ago
      China says hi.
      • giancarlostoro 2 hours ago
        Not in my case, I don't see any of my employers (past or current) trusting a country like China with their data.
        • HDBaseT 2 hours ago
          The decades of US brainwashing children into thinking China is the big bad guy has worked unfortunately.
    • aanet 4 hours ago
      +10000 to that
  • Laurel1234 4 hours ago
    The only company less trustworthy than OpenAI and Anthropic is meta.
  • qphe95 5 hours ago
    Theres no actual evidence they didn't just distill Kimi K3
    • toephu2 5 hours ago
      At this point, it doesn't matter who is distilling from who.
    • Jabrov 4 hours ago
      Is there any actual evidence that they did?
  • vcryan 4 hours ago
    It seems like one day, Google or Meta might produce a coding model worth discussing. That day is not today.
  • esafak 5 hours ago
    If anyone from Meta is reading, please can you publish the cost and latency for each of your benchmarks, like OpenAI does? Show us how the reasoning effort level affects them in 2D charts. This needs to become standard practice.
  • Readerium 4 hours ago
    Lol worse than DeepSeek
  • rvz 5 hours ago
    First of all, you have login to use it. Why?

    After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.

    Think twice before falling for this announcement and ask yourself what they are not telling you.

  • minimaxir 5 hours ago
    Muse Spark 1.1 was released July 16th, less than a month ago. A new version release this soon (particularly after Kimi K3's release drastically overshadowed it) is a bit sus and it appears that Meta is trying a first launch do-over.
    • ac29 2 hours ago
      Doesnt seem suspect to me, training runs have checkpoints and there is no reason you cant release a checkpoint even if you are still training the model
    • gaogao 4 hours ago
      Frequent minor version bumps are pretty common these days. Opus 4.7 -> 4.8 was 42 days.
      • minimaxir 4 hours ago
        Which was in itself a do-over because Opus 4.7 received a lot of bad press on suspicion of being a regression from 4.6.
  • arjie 4 hours ago
    Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model.

    The use traces must be crucial to functionality which is why they’re keeping prices so low.

    • wmf 4 hours ago
      They rebooted less than one year ago so this is decent progress. Obviously users don't care about progress though.
      • arjie 3 hours ago
        Yeah, progress is useful as an internal metric, but I'm going to measure against the present frontier unfortunately. Eager to see what they come up with in the future.