9 ms·
Muse Spark 1.3
https://research.meta.ai/blog/introducing-muse-spark-1-3 https://research.meta.ai/blog/introducing-muse-spark-1-3
- frozenseven 1mo agoThis should probably be primary: https://news.ycombinator.com/item?id=49541149 https://news.ycombinator.com/item?id=49541149
- simonw 1mo agollm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff902fb6c340a3c5fc0bea317ef7bef79 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fce974a21202b0595e36ec2a5ddb51480#response https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F4f34f84caa12a306bded637ea495698d https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The most expensive was reasoning level xhigh - 7.5 cents, 1m34s. And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F950ba8b7ed5baa0be56f52425f2315ad https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
- tomrod 1mo agoWhat does the mean pelican look like at this point? Also 3X token use vs. 1.2
- _puk 1mo agoRed eyes and a tattoo?
- drusepth 1mo agoIs there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.
- simonw 1mo agoIt's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction. The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
- piker 1mo agoI was going to ask the exact same question earlier but deleted it after thinking “I’m sure Simon has done some sort of discussion on this.” Since it does seem novel to you, too, it would be really interesting to read more about this phenomenon.
- m12k 1mo agoIt's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you prompted it in Hebrew, Arabic or another right-to-left language?
- jermaustin1 1mo agoYears ago, I lived in NYC, and my roommate was a director of photography for National Geographic, and various other nature documentaries. I loved photography (still do, but much less time for it as a late 30s adult than a mid 20s adult), and she was kind enough to answer any question I had regarding film/photo. She told me that "left to right" denoted progression in the story, "right to left" told the viewer the subject was "exiting" the current scene. She didn't go into the details of WHY, and I probably didn't probe deeper, but it stuck with me, and I notice it all the time in film and television.
- jmkni 1mo agolol Definitely an upgrade over 1.2
- jonahx 1mo agoIf you have a grading rubric, huge points off for adding arms instead of using the wings as arms!
- Fergusonb 1mo agoI think it's hilarious that this detail is enough for me to dismiss looking into the model, but here we are, and it is.
- jonplackett 1mo agoHas any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that?
- NoOneCares44 1mo ago[flagged]
- jttnr 1mo agoI wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?
- EugeneOZ 1mo agoAbsolutely BRUTAL! :) Thank you for doing this, I love your benchmark the most!
- drob518 1mo agoSimon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.
- gpt5 1mo agoWe should just consider the pelican bench as saturated and mostly meaningless.
- simondotau 1mo agoBut the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable, or a pink tennis ball with strawberry seeds and a reset button) and you can see that improvements are not narrowly tailored.
- crimsoneer 1mo agoSomeone tested this, and it doesn't look to be saturated. https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world. https://simonwillison.net/2026/Jul/16/kimi-k3/ https://simonwillison.net/2026/Jul/16/kimi-k3/
- kaoD 1mo ago> Someone tested this, and it doesn't look to be saturated. They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>". Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
- simondotau 1mo agoThey're still not yet at the point where pelicanmaxxing is the best way to win this benchmark. Earlier models sucked because their SVG skills sucked. Newer models are likely better because more/better SVG models are being added to their training data.
- hollowturtle 1mo agoIs there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too
- tintor 1mo agoIt would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.
- nojs 1mo agoThe amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data. It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
- fc417fc802 1mo agoYou can't benchmaxx spatial awareness without solving the fully general problem (at least I figure).
- BeetleB 1mo agoYou win this thread's prize: https://news.ycombinator.com/item?id=49538333 https://news.ycombinator.com/item?id=49538333
- cheesecakegood 1mo agoIt also works as extremely effective engagement farming, for lack of a better phrase
- tintor 1mo agoDid any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans. Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
- TiredOfLife 1mo agohttps://en.wikipedia.org/wiki/Bird_feet_and_legs#/media/File:Bird_leg_and_pelvic_girdle_skeleton_EN.gif https://en.wikipedia.org/wiki/Bird_feet_and_legs#/media/File... Bird knees bend same way human ones do
- gnatolf 1mo agoIt's clear they mean the 'exposed' joint where humans assume the knees, and where one can see the leg bend. Technically you're correct, but it's just that. Please answer in better faith instead of well akshually.
- tintor 1mo agoHere is a photo of Pelican: https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2Fas1.ftcdn.net%2Fjpg%2F12%2F53%2F44%2F44%2F1000_F_1253444423_R7InsJ5maiWprpbOeLEnZkSLXyEqmxn1.jpg&f=1&nofb=1&ipt=7bbb69bb412c2651e84da0ef9a0526980f1dc64e67fecda60ba29a582b1fa2ff https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
- ImprobableTruth 1mo agoThat's the ankle. The actual knee is hidden in the feathers of the body.
- 0xbadcafebee 1mo agoFor all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html https://dylancastillo.co/posts/pelicanmaxxing.html "Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing. Or at least they’re not doing it in a plainly obvious manner."
- ipsum2 1mo agoAll of the links show "Error: Gist API returned 403".
- leumon 1mo agonext, try: "generate an svg of a human hand". this is a prompt where many models fail imo.
- wewewedxfgdf 1mo ago[flagged]
- panarky 1mo agoIf you could write the SVG on the whiteboard then I'd hire you.
- fuddle 1mo agoI also aced my interview by focussing on pelicancode problems, instead of leetcode problems.
- deleted 1mo ago[deleted]
- sroussey 1mo agoYou should post your source code you wrote here… ;)
- labrador 1mo agoI interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.
- hunterpayne 1mo ago"software developer"...you keep using that word. I do not think it means what you think it means.
- labrador 29d agoWhat does it mean for you, sex robot developer?
- rattray 1mo agoIs this for real
- pavs 1mo agoFYI, your renderer breaks with error "git api access error 403", rate limiting error from git, when using cloudflare vpn. I am guessing its not super common, but it happens just so you know.
- andytratt 1mo agoexcellent thread
- m00dy 1mo agoI see no point having these pelicans used for anything related model qualification.
- simonw 1mo agoAt this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
- menaerus 1mo agoI don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doing good on actual work you need to do is close to zero.
- dwaite 1mo agois there a reason there are so many common base decorative elements across pelicans on bicycles? For instance, there's a lot hats/helmets and scarfs/capes across models.
- dhon_ 1mo agoI'm waiting for the models to start responding with "Oh hi Simon!"
- rexthonyy 1mo agoI'm not sure why it had to have the pelican wearing a red scarf seeing as that was not in the prompt
- 6r17 1mo ago"The LLM is better because the pelican hat is better" Benchmarking like never before
- coverband 1mo agoThe pelican is for the last gen of LLMs -- have you tried a penguin instead?
- Melatonic 1mo agoIt's actually the other way around - the evil Penguin villain from Wallace and Gromit is secretly controlling SimonW !
- MagicMoonlight 1mo ago[dead]
- tcp_handshaker 1mo agoWould it not make more sense, assuming the purpose is to have a quick smoke test of model quality...to do a different animal, in a different setting each time, so as to defeat any tuning for your benchmark? Then go back and do the same for other models? Keep the pelican as a side baseline?
- simonw 1mo agoI do that any time I'm suspicious that a model has done too well. My dream is to catch a lab that does a perfect pelican on a bicycle but is bad at other animals on other forms of transport.
- nightmunnas 1mo agoHave you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"
- w4yai 1mo agoInteresting question
- benjamintelliot 1mo agoI asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.
- Zambyte 1mo agoI decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right. Kimi 2.6: treated the question as a riddle, did not know. Kimi 3: Simon Willison GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme. Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names. Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog. Qwen 3.7 Plus: Did not know. Qwen 3.8 Max: Simon Willison GPT OSS 120B: Did not know. GPT 5.6 Luna: Treated it as a riddle, guessed wrong. GPT 5.6 Terra: Treated it as a riddle, guessed wrong. GPT 5.6 Sol: Treated it as a riddle, guessed wrong. DeepSeek V4 Flash: Treated it as a riddle, guessed wrong. DeepSeek V4 Pro: Treated it as a riddle, guessed wrong. Gemma 4 31B: Treated it as a riddle, guessed wrong. Gemini 3.1 Flash Lite: Guessed wrong Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model). Gemini 3.7 Flash: Simon Willison Muse Spark 1.2: Treated it as a riddle, guessed wrong. Grok 4.3: "I have no idea" Grok 4.6: Simon Willison Mistral Medium 3.5: No way to know Mistral Small 4: I don't have enough information Hermes-4-405B: Guessed wrong MiniMax M3: Treated it as a riddle, guessed wrong. Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
- Gecko4072 1mo agoUsed Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.
- WASDx 1mo agoI'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.
- scotty79 1mo agoIs the fact that everybody almost catches up with the frontier a sign that we are entering a new region of sigmoid curve?
- dominotw 1mo agometa fails at everything yet is frontier on this one
- redox99 1mo agoNo because the frontier keeps advancing very fast.
- samuelknight 1mo agoMeta has an enormous amount of compute. They are either going use it making and inferencing models or they are going to sell their excess capacity to model providers. Zuck had to completely rebuild his AI team after the Llama 4 launch mess.
- schopra909 1mo agoProgress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up. Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether. In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree. That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.
- refulgentis 1mo agoI don’t know why people think DeepSeek did reasoning models / RLVR before OpenAI, there was a gap of months.
- wxw 1mo ago“contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol. Definitely shows how important a user data flywheel is for RL and model improvement.
- finnjohnsen2 1mo agoSo one model is "Not used to improve our products" and is 10-20 times more expensive to the "Used to improve our products"-model. Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
- whimsicalism 1mo agothe meaning is pretty obvious - they want to train on your chats & tasks and are willing to subsidize for the privilege of doing so.
- thefreeman 1mo agoaren't they explicitly saying this with both their pricing and their wording? I'm not sure what you are alluding to?
- bigyabai 1mo agoGiven OpenAI and Anthropic's behavior, do you really expect them to be singled out for this practice? Zero trust has been in "LGTM" territory for years now. Meta's bet against people taking a principled stance arguably paid off great.
- Jcampuzano2 1mo agoI'm confused what your surprise is here. It's plain and simple right to the point wording. I don't see the wiggle room at all.
- duplessitous 1mo agoWhat is the confusion? They directly state that you are the product if you use their discounted offering. It isn't an assumption that should lead you to this, it is Meta's very direct communication that should lead you to this
- IshKebab 1mo agoI think it's more that the "not used to improve our models" is expensive because companies need that. It's simple price differentiation. In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
- majerep 1mo agoThe previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free.
- mgaunard 1mo agoThey could have just called the article "struggling to remain relevant"
- geooff_ 1mo agoCould this be best intelligence / $ if you're willing to let zuck digest your data?
- oofbey 1mo agoYeah. Super icky. But this might be the first time in Zuck’s life he’s being honest about the business model.
- HDBaseT 1mo agoBy default, even without the training endpoint the pricing is pretty competitive, especially against Opus and Fable. [1] The 'muse-spark-1.3-contributor' endpoint is by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too. This price/intelligence beats even legacy DeepSeek V4 Flash pricing. [1] https://artificialanalysis.ai/#total-cost-tabs https://artificialanalysis.ai/#total-cost-tabs
- meerita 1mo agoI declined the use of cookies and everything went black. No content at all. Dissapointed.
- 7734128 1mo agoPractically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.
- 0xbadcafebee 1mo agoI'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models
- HDBaseT 1mo agoA small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily. It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways. [0] https://www.anthropic.com/research/small-samples-poison https://www.anthropic.com/research/small-samples-poison
- owaiswiz 1mo agodoesn't mean the raw text goes into training. they most likely have a pipeline to clean out any secrets before they train on it?
- deleted 1mo ago[deleted]
- superfrank 1mo agoI started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it. I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be. I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be. Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
- WASDx 1mo ago[dead]
- MangoCoffee 1mo ago>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
- KptMarchewa 1mo agoI would imagine your interactions with it are more important than the output.
- tinyhouse 1mo agoI had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits.
- meric_ 1mo agoYou should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks
- dcl 1mo agoAny more info on this?
- meric_ 1mo agoCline experiment: https://x.com/cline/status/2085237843379519737 https://x.com/cline/status/2085237843379519737 Muse code: https://developer.meta.com/ai/resources/blog/build-with-muse-code/ https://developer.meta.com/ai/resources/blog/build-with-muse... > Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you already use.
- dcl 1mo agothank you
- tinyhouse 1mo agoThanks. Just downloaded and pretty impressed so far. It's fast and nice to work with.
- bertili 1mo agoDeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
- cbg0 1mo agoBut is the score really reflective of the quality or are both models benchmaxxing?
- bermudi 1mo agoMuse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD. This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
- gpt5 1mo agoBoth versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.
- dominotw 1mo agohow much of it is from reallocation of staff to ai training and labeling
- WASDx 1mo agoWith the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
- dakolli 1mo agoand they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
- Lucasoato 1mo agoA model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim). Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
- dbbk 1mo agoHow is it not SOTA? It's beating 5.6 Sol.
- ctolsen 1mo agoYou gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now.
- pqdbr 1mo agoThats 3 months in AI years.
- cdelsolar 1mo agowhat is it in dog years
- tclancy 1mo agoSeptember.
- redbear2026 1mo agoeternal september.
- likestowatch 1mo agowe did it reddit!
- ChrisArchitect 1mo agoBlog post: https://research.meta.ai/blog/introducing-muse-spark-1-3 https://research.meta.ai/blog/introducing-muse-spark-1-3 (https://news.ycombinator.com/item?id=49541149 https://news.ycombinator.com/item?id=49541149)
- sunaookami 1mo ago>Previously available reasoning modes are available today with max reasoning coming shortly after we finish additional safety testing Lmao. And their benchmark table only shows max reasoning.
- mromanuk 1mo agoI didn't like 1.2, It make some mistakes in a web app, so I quickly went back to Claude, Kimi K3 or Deepseek V4. Hope this one can clear agentic development, because Muse Spark models are fast and cheap.
- tyre 1mo agoMeta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI. I feel the same about Grok w/ Elon. I will pay extra to use someone else. I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money. And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
- loeg 1mo ago"Avoid generic tangents" / "Please don't complain about tangential annoyances."
- reaperducer 1mo ago"Avoid generic tangents" / "Please don't complain about tangential annoyances." That's pretty much 90% of HN these days. Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens. Microsoft releases a new version of Windows? Here come the gripes about Azure. Google changes something in GMail? Play Store! It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.
- _diyar 1mo agoHow is this a tangential annoyance or a generic tangent? > Meta announces they have a new model, demonstrating its capabilities. > Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“
- loeg 1mo agoGrandparent comment has zero to do with the article. It's just GP generically bitching about Meta. (Your "quote" of the comment does not appear anywhere in the actual comment.)
- souvlakee 1mo agoWhy they didn't use LLM to create html table instead of https://lookaside.fbsbx.com/elementpath/media/?media_id=1048442011123823&version=1788377103&transcode_extension=webp https://lookaside.fbsbx.com/elementpath/media/?media_id=1048...?
- jumploops 1mo agoThe "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data. The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $). Stats: 1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
- 2001zhaozhao 1mo agoI have a feeling that Meta is not gonna like what people actually use the contributor model for lol. (It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)
- HDBaseT 1mo agoNot to mention, this is hyper competitive against even Chinese providers given its multi-modal support. Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.
- a012 1mo agoMuse Spark may be competitive in capabilities but it’s not for serious works since Meta trains on your prompts so no ZDR, in contrast Chinese provider like Z.AI promises ZDR which is more attractive to big corps.
- fibonacci112358 1mo agoIs everyone rushing to launch something before Astra tomorrow?
- PunchTornado 1mo agoWhat is astra?
- lovelymono 1mo agoProbably https://openai.com/index/path-to-astra/ https://openai.com/index/path-to-astra/? > We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step. > We plan to make Astra available soon[, but access to its most advanced cybersecurity capabilities will be more limited].
- IshKebab 1mo agoLol "not used to improve our models" is AI's enterprise SSO.
- r_lee 1mo agothat'd be ZDR, the one you need to beg from their Sales teams with $$$
- cnxhk 1mo agoartificial analysis results: https://x.com/ArtificialAnlys/status/2095247787277553929 https://x.com/ArtificialAnlys/status/2095247787277553929
- johnnyApplePRNG 1mo agothanks, google ... are you kidding me?! posting an x.com link to a cheating benchmarking website? get out
- lostmsu 1mo agoWhat a day. OpenAI is behind basically all major competitors - at least for a some amount of time.
- dangoljames 1mo agoIf it's from meta, pit h in the bin.
- LZ_Khan 1mo agoHa, even with monitoring engineers keystrokes and mouse movements not SotA on OSWorld.
- doublerabbit 1mo agoMeta is all okay now because they've released a new LLM model. /s
- maciejgryka 1mo agoDoes anyone know what the license for this model is? Specifically any word on restrictions about what it can be used for?
- lovelymono 1mo agoThis: https://dev.meta.ai/legal/terms-of-service https://dev.meta.ai/legal/terms-of-service and this: https://dev.meta.ai/legal/acceptable-use-policy https://dev.meta.ai/legal/acceptable-use-policy, looks like it.
- apodolny 1mo agoI like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent.
- mmastrac 1mo agoAny idea what size this is?
- adrian_b 1mo agoMark Zuckerberg said that they will release soon Muse Spark as open weights, in which case we will see the size. However, the statement did not include any details, so it is not clear if the open weights variant will be the same that they are hosting now, or some scaled down version.
- gehsty 1mo agoAs a product, would developers switch to a meta model/harness? I don’t think so. Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI? I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?
- phyrex 1mo agoMeta also has 50k engineers. Not to mention that tons of meta infrastructure - including ads! - use AI. Would you want that sort of business be this dependent on someone else?
- improgrammer007 1mo agoAll people here care about is hating Meta. Just look at the top voted comment. No one cares about the merits of the model, etc. HN has become nothing but an echo chamber.
- jmward01 1mo agomuse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?
- popularonion 1mo agoThis has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!" We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for. I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.
- bilalnpe 1mo agoBoth of them let you opt out of it on subscriptions.
- cma 1mo agoWe already had a good idea of how valuable it is from how much X.ai acquired Cursor for, and the near-immediate improvements to their coding scores.
- bluecalm 1mo agoThis is also really smart business wise imo. For hobby projects, toys, quick scripts you don't really mind if they train on it. It's a win-win. Once you get used to the tools and you want to do more serious business you are more likely to buy a more expensive sub from them.
- anjel 1mo agoNot mentioned in pricing: Surveillance costs of using Muse Spark
- ryanschaefer 1mo agoFor all of the comments about training: I thought that subscription plans for other models allow the same. Am I mistaken?
- Aurornis 1mo agoIt's a toggle. Some will automatically enable it and you have to turn it off. People who rapidly click through setup flows can miss it and leave it enabled.
- coolcoder613 1mo agoI have not tried Muse Spark for code, but I've been using it for a while to write Latin. I find it's one of the best at it, alongside Gemini. For example, I've recently been using it to translate the subtitles of the show I'm watching into Latin, to provide me with a bit more input. (I'm learning Latin, for context)
- ejboy 29d agoProbably because Zuck is a big fan of ancient Rome.
- deleted 1mo ago[deleted]
- dcl 1mo agoVery keen to try this after using Claude Code over the last few months. Should I just point Claude Code to Muse Spark endpoint (because I'm familiar with Code)? What do people think of Muse Code or other coding agent harnesses?
- alexboehm 1mo agoJust try opencode, it comes with 1.3 contributor free.
- dcl 1mo agoWell thats very interesting. Thank you. Will be interesting to see how hard/easy it is to translate my Claude skills, loop design, etc to the new harness. This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?
- jonahhorowitz 1mo agoComing from Claude Code, I initially went with opencode but switched to pi.dev after a while and I think I like it more. It's lighter weight. It's worth trying both.
- wkcheng 1mo agoHow do people actually use this? Do they use it through some sort of subscription plan, or via OpenRouter?
- dv35z 1mo agoYou can check out Muse Spark 1.3 by using OpenCode (https://opencode.ai/ https://opencode.ai/ - open-source AI / coding harness). There's a terminal version and a GUI / desktop version. Good luck!
- yanjunnf 1mo agoIt's true that there hasn't been any meta news about LLM for a while now
- keyle 1mo agoI am very impressed by this model so far. It's faaast and it seems to be just intelligent enough to do really well. It's UI work (simple python UI) is very clean and functional. The UX was 'there'.
- rho138 1mo ago[dead]
- m00dy 1mo ago$META has everything it needs, great team, great models coming out, great infrastructure (GPUs), great userbase and distribution channels. $META is underrated.
- israrkhan 1mo agoit seems like gemini 3.8 flash is more capable and cheaper. The only reason i would use this is if i was willing to share my data with meta, and allow them to train on my data. In that case it becomes dirt cheap.
- esafak 1mo agoFunny how quickly Meta caught up after Lecun left.
- ydna404 1mo agoFor folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?
- MitziMoto 1mo agoSome of us own and run companies? Cost per performance is a huge deal.
- Mashimo 1mo agoEven with subscriptions, it means you get more: https://opencode.ai/go https://opencode.ai/go On this 10 USD / month sub you can do over 250 times more request compared to Kimi 3 or Grok. Or 20 times as much as ChatGPT Luna.
- IIIIIllIIII 1mo agoIm a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.
- roytam87 1mo agoI have opposite result: MS1.2 wrote C code without following original source code writing style, and no descriptive info why writing such code, DS4F or even Mimo seems better to me.
- Athanase000 1mo agoI think the person you are answering to was saying the same thing. They wrote "worth" instead of "worse".
- geoffbp 1mo ago> /taste: an anti-slop filter: a flat checklist of visual defaults not to use, so generated UI stops looking machine-made. This is interesting
- xyzkoi 1mo ago[dead]
- water-drummer 1mo agoStill waiting on them to release weights for Muse Spark 1.2, like they promised to. Wonder if they plan on doing the same for 1.3 which would be crazy
- Iolaum 1mo agoZuck hinted at it n his twitter post but I doubt it.
- luciana1u 1mo ago[flagged]
- unsupp0rted 1mo agoI'm annoyed my (US-bought) Meta glasses still block me from using the AI features, months after moving back to a country where it's generally enabled.
- bdlowery 1mo agobenchmaxxed model
- goatydev 1mo agoI used 1.2 for free for a while, and it was a pretty good experience. 1.3 would also be worth using, provided the price is reasonable.
- sourcecodeplz 1mo agoprice is same, also for contrib version
- 1saadcodes 1mo agoGemini 3.8 Flash still looks like the better pick to me. Muse Spark 1.3 is nice, but Gemini gets you similar performance for a cheaper price. Not to mention with the pace at which Google is moving with their Flash models I expect a new one to release soon
- sourcecodeplz 1mo agoi've been using muse spark 1.2 contribs since launch exclusively. no other models. i like it very much. it is different than all other chinese models distilled from claude. just ask it to do some front-end work and you will see its not the same UI as all other claude/distils. also the price is unbeatable, $0.002 input caching. its the same as old dsv4-flash prices.
- vinhnx 1mo agoMuse Spark 1.3 Max is the first Meta model to surpass OpenAI’s best on Artificial Analysis’ Intelligence Index. https://artificialanalysis.ai/models#intelligence https://artificialanalysis.ai/models#intelligence
- wartywhoa23 1mo agoI wonder if any commenters here were among those who used to ridicule the rate at which new JS frameworks kept popping up in the 2010s, and the amount of heroic zeal required to never miss the bandwagon?
- podgorniy 1mo agoQuite similar to those times. The difference is that swapping existing LLM with new one is waaay easier than the frameworks. And competition reflects on the price for consumers. So more LLM options/providers/opensources appears better than rain of js frameworks.
- lylo 1mo agoI've tried it via OpenCode and I'm impressed. So fast compared to Opus, and the results so far are comparable I'd say.
- deleted 1mo ago[deleted]
- grkn 1mo agoIt's funny that it comes with *-contributing model on in the CLI as default. All code examples are like that as well. Any company without bad intentions would do the opposite, but no not with Meta. I'm super impressed with their level of evilness on every product.
- XCSme 1mo agoNice, Muse Spark is so good and keeps improving, but it's still not the best choice for any use-case. The Sol models are in their own league currently in terms of cost/speed/performance. Good improvements from 1.1 and 1.2[0], but when I tested 1.3 it was very slow (through openrouter). [0]: https://aibenchy.com/compare/meta-muse-spark-1-3-high/meta-muse-spark-1-2-high/meta-muse-spark-1-1-high/ https://aibenchy.com/compare/meta-muse-spark-1-3-high/meta-m...
- sscaryterry 1mo agoCannot agree more with the Sol models. Everything else I try just seems "dumb".
- deleted 1mo ago[deleted]
- kkkamur 1mo agoDamm this is so cheap literally, I have been running 100s of subagents and it is cheap - with the contributor model ofc :)
- hnjbx769kd 1mo ago[dead]
- kmike84 1mo agoThe big news is that it's going to be open weight - https://x.com/finkd/status/2095232032896946311 https://x.com/finkd/status/2095232032896946311.
- vickychijwani 1mo agoThat tweet says “Muse Spark open weights releases coming soon”, not that this specific model is going to be open weights.
- dinga 1mo agoI dislike meta so much.. I try to avoid that company as much as possible.
- ifahimreza 1mo agochatgpt, grok, and claude are down, but muse is working, that's a relief.
- mark_l_watson 1mo agoMeta is getting most of my casual vibe coding business as long as they keep giving about 90% discount in the ‘we train on your prompt interactions data.’ I have had sos-so results with their local muse 30b model, but the hosted API is very fast and I have been getting good results.