10 ms·
Muse Spark: Scaling towards personal superintelligence
https://meta.ai/ https://meta.ai/
- bustah 6mo ago[flagged]
- warthog 6mo agoHoping the benchmarks are correct this time...
- htrp 6mo agoAnyone done vibe testing at meta ai yet?
- moab 6mo ago"Muse Spark is available now, and Contemplating mode will be rolling out gradually in meta.ai." How does one get their hands on these models? They are not open-source, right? I go to meta.ai, but it's just a chat interface---no equivalent to codex or claud code? Can you use this through OpenCode? Is meta charging for model access, or is the gathering of chat data a sufficiently large tithe?
- monkeydust 6mo agoTBD it seems. So far the only explained usage pattern is through a Meta product (Whatsapp, Facebook, Instagram).
- moab 6mo agoSo to verify their claims and see how strong these models are, the answer is "believe us"? Note: I'm expressing some skepticism here largely due to how recent rollouts from Meta flopped. Sincerely hoping that they do better this time around!
- nemomarx 6mo agoI assume the answer is try it out in the chat mode? You could run your usual benches through that right
- pstuart 6mo agoI appreciate that they build this stuff for their own benefit, but I don't want to feed even more of my private info. Hopefully the models will become public or lead to equivalent models from other sources.
- meetpateltech 6mo ago"It will be available in private preview via API to select partners, and we hope to open-source future versions of the model." from Facebook Newsroom: https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs-first-model-built-to-prioritize-people/#:~:text=It%20will%20be%20available%20in%20private%20preview%20via%20API%20to%20select%20partners%2C%20and%20we%20hope%20to%20open%2Dsource%20future%20versions%20of%20the%20model. https://about.fb.com/news/2026/04/introducing-muse-spark-met...
- tempaccount420 6mo agoI can't think of any "select partners" that would want to use this non-SOTA model. Just put it on OpenRouter.
- giancarlostoro 6mo agoIf Microsoft is a select partner, maybe they could shove it into Copilot for VS or something, but yeah, I'm wondering the same, maybe Apple could be one of their partners too?
- mark_l_watson 6mo agoThat would be my question also. I like it when companies have easy to sign up for, pay as you go models. Being able to buy $5 worth of tokens and get an API key - in less than a few minutes - is ideal.
- ddp26 6mo agoThe second paragraph starts "Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts. To support further scaling, we are making strategic investments..." This article is about Meta, not about the user. Who signs off on these? Is the intended audience other people at Meta, not the user?
- tjkrusinski 6mo agoThe article is published primarily to signal to the market that Meta is serious in its efforts to compete in building frontier ai models. They want to 1) attract talent, 2) tell wall street they can play in this space as well, 3) help employees feel the company is moving in the right direction. A frontier LLM doesn't apply to their core consumer products.
- Lihh27 6mo agothe blog is the product. investor deck posted as a tech launch
- conradkay 6mo agoStock up 9% today, very pleasant for Zuck if you do the math on his net worth :)
- hungryhobbit 6mo agoI mean, kinda? It's not like Zuck is selling his stock tomorrow, so daily fluctuations in stock price don't really affect him.
- throwaway173738 6mo agoHe can borrow against that, so it actually does matter.
- zurfer 6mo ago> Muse Spark is available today at meta.ai and the Meta AI app. We’re opening a private API preview to select users.
- m4r1k 6mo agoSo no Open-weight .. why one would choose Muse Spark instead of Anthropic, OpenAI, or Google models all featuring from good to amazing harness?
- Artgor 6mo agoI'm cautiously waiting for the feedback from the first users. Meta has produced a lot of great models (LLama), maybe this is a comeback... but I'm cautious, as the jump in the quality is almost too high. Also, I think people aren't used that using such models requires meta.ai or meta ai app.
- solenoid0937 6mo agoMy Meta friends say it's benchmaxxed af
- loeg 6mo agoWe used to call this "overfitting," but I suppose everything has to be maxxed now. Fitmaxxed?
- conradkay 6mo agoIt doesn't seem benchmaxxed, ARC AGI 2 score is quite bad (42.5%, GPT 5.4 is 76.1%) and coding is okay. But maybe this is the best Meta can do even benchmaxxing The impressive part is multimodality, very plausible since there's less focus there by other labs (especially Anthropic)
- dbgrman 6mo agoGiven llama 4 mucked up benchmark numbers, I’d take spark announcement with a many grains of salt.
- khalic 6mo agoOh good, if they built a lab, I’m sure they took the time the precisely define what they mean by super intelligence? Right? …
- gallerdude 6mo agoThis would have been an amazing release 6 months ago. But the industry moves so fast, this is a trite release. Maybe it’s best for Meta to sell their superintelligence division. I don’t think Zuck’s vision is particularly compelling.
- gordonhart 6mo agoA new model comparable (ish) to the Claude/Gemini/GPT flagships is a big deal for the industry and for Meta even if it doesn't set the new frontier.
- zozbot234 6mo agoTheir new Contemplating mode gives this model a Deep Research ability (akin to existing models from GPT and Gemini) that might make it quite comparable to the just-announced Mythos.
- solenoid0937 6mo agoMythos is a much bigger pre train, Contemplating is not the same thing.
- zozbot234 6mo ago> Mythos is a much bigger pre train Do we have data to substantiate that claim?
- solenoid0937 6mo agoIt's pretty common knowledge. Spud is the only other PT comparable with Mythos. Both Spud and Mythos can also scale via inference time compute. Meta simply did not have enough compute online, long enough ago, to have a similar PT.
- temp_praneshp 6mo ago> might make it quite comparable to the just-announced Mythos Do we have data to substantiate that claim?
- alyxya 6mo ago[dead]
- chrsw 6mo agoSo Meta is not releasing open source models anymore?
- sidcool 6mo agoWill experiment with the model. But I am scared of sharing any information with the Zuck ecosystem.
- toddmorey 6mo agoQuestion: since they've rebooted their approach to AI... have they given up on open models? There's no mention of open source or open weights or access to the models beyond their hosted services.
- thegeomaster 6mo agoAlexandr Wang on Twitter [0] mentioned open source plans: "this is step one. bigger models are already in development with infrastructure scaling to match. private api preview open to select partners today, with plans to open-source future versions. incredibly proud of the MSL team. excited for what’s to come!" https://x.com/alexandr_wang/status/2041909388852748717 https://x.com/alexandr_wang/status/2041909388852748717
- prodigycorp 6mo agoSo the answer is: no. lol. Remember Llama 4 Behemoth, and how we were supposed to get more great models from it?
- wmf 6mo agoThis may be too large to run locally anyway. Maybe they will distill down some smaller open versions later.
- OsrsNeedsf2P 6mo agoThe only benchmark they show against SOTA models is in bioweapons refusal. Edit: nvm I can't read, regular benchmarks against SOTA are there
- santiagobasulto 6mo agoThis looks like a very interesting model and very promising, especially after llama lost so much ground recently. I hope they release the weights
- visioninmyblood 6mo agohttps://meta.ai/ https://meta.ai/ this is where you can try it seems like the API is not publicly accessable yet. I feel they are very late to the game and do not show value to customers over other models.
- p_stuart82 6mo agolate isn't the problem. private preview api and no reason to switch. that's just another hosted model
- throwaw12 6mo agoHow is that Meta spent so much money for talent and hardware, but the model barely matches Opus 4.6? Especially, looking at these numbers after Claude Mythos, feels like either Anthropic has some secret sauce, or everyone else is dumber compared to the talent Anthropic has
- wotsdat 6mo ago[dead]
- zozbot234 6mo ago> has some secret sauce Yup, it's called test-time compute. Mythos is described as plenty slower than Opus, enough to seriously annoy users trying to use it for quick-feedback-loop agentic work. It is most properly compared with GPT Pro, Gemini DeepThink or this latest model's "Contemplating" mode. Otherwise you're just not comparing like for like.
- throwaw12 6mo ago> it's called test-time compute. Why can't others easily replicate it?
- coder68 6mo agoI have not delved into the theory yet but it seems that the smaller open-source models do this already to an extent. They have less parameters, but spend much more time/tokens reasoning, as a way to close the performance gap. If you look at "tokens per problem" on https://swe-rebench.com/ https://swe-rebench.com/ it seems to be the case at least.
- strulovich 6mo agoMeta did a bunch of mistakes, and look like Zuckerberg spent a lot of money on talent and made big swings to change it (that happened about a year ago) I think it’s unrealistic to expect them to come back from that pit to the top in one year, but I wouldn’t rule them out getting there with more time. That’s a possible future. They have the money and Zuckerberg’s drive at the helm. It can go a long way.
- rvz 6mo agoUntil you actually try the model itself, assume any benchmark presented to you as being part of the marketing material of the model, as it is not independently verified and completely biased. The same is true with any other model, unless otherwise stated. In the next few days, we'll see who Meta has paid to promote this model on social media.
- daft_pink 6mo agoThis really reinforces the idea that the AI race and the Railroad Mania of the 19th century are very similar. So many different companies are going to have similarly powerful ai that there will be no moat around it and it will be cheap. They will never earn their investment back.
- dist-epoch 6mo agoThe moat is in the compute and the energy access. And further down the line in chips, which is why Elon is building a fab now. There are plenty of capable models on HuggingFace, yet I have no way of running them.
- cedws 6mo agoThat fab will never be delivered. In five years you might see the manufacturing equivalent of a person dancing in spandex.
- QQ00 6mo agoThe only company that musk own and actually achieve something is spacex. so I believe you. He likes to hype things beyond what is actually possible. spacex is engineering masterpiece with how they revolutionize the space industry.
- declan_roberts 6mo agoDo you know about his other company and what they do?
- khalic 6mo agoGive it a few years, or month. Tiny models are getting outrageously good
- mobattah 6mo agoExactly. We’ll see the cost of AI continue to drop. I was saying this for years about Tesla’s FSD - they finally had to give in and drop the price to stay competitive.
- _2d30 6mo agoRan some of my internal benchmarks against this and I'm very unimpressed. I don't think this moves them into the OAI v Anthropic v Gemini conversation at all. Major analytical errors in their response to multiple of my technical questions.
- _2d30 6mo agoPlaying with this some more and it's actively not good. Just basic mathematical errors riddling responses. Did some basic adversarial testing where its responses are analyzed by Gemini and Gemini is finding basic math errors across every relatively (relative to Opus, Gemini or GPT can handle) simple ask I make. Yikes.
- smlacy 6mo agoPost actual results, make a blog post. Don't just say "this sucks" without tangible evidence. Otherwise you're doomed to "sample size of one" level of relevance.
- thorum 6mo agoI have the opposite experience: random HN/Reddit comments saying “this sucks” or “whoa this is a huge improvement” are the only benchmark that means anything. Standard benchmarks are all gamed and don’t capture the complexity of the real world.
- titanomachy 6mo agoThen your internal benchmarks will be in the post-training set and you’ll have to make new ones.
- _2d30 6mo agoI may already have but I'm pseudonymous on this website.
- smlacy 6mo ago[flagged]
- oliver236 6mo agoso glad its beating all the others on bioweapons refusal. this is what i most wanted out of the latest SOTA model
- wmf 6mo agoZuck has a lot more experience being summoned before Congress than you.
- deleted 6mo ago[deleted]
- chankstein38 6mo agoPersonal Superintelligence made me think this was an open-source model being released and I was excited. Then I continued reading and I'll just wait until the model comes out.
- gardnr 6mo agoI was really excited until I realised that “personal” meant “owned by meta“. I’m trying to decide is I find the doublespeak a bit offensive or not.
- dbgrman 6mo agoI wonder if Zuck will ever internalize that the words ‘personal’ and ‘meta’ will not be taken seriously together for another decade (if they don’t make another gaff).
- ComputerGuru 6mo agoSo does this confirm the end of llama?
- jansport123 6mo agodid they just copy the chatgpt ui?
- ChrisArchitect 6mo agoAssociated Meta news post with consumer-friendly takes: https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/ https://about.fb.com/news/2026/04/introducing-muse-spark-met...
- ehutch79 6mo agoHow's the metaverse doing? It was the next big thing and how we're all going to be working inside it in... was it like 3 months ago? Maybe they need to mine more libra coin first? or is it diem now? is that even still part of meta? I'm sure this new AI is super intelligent and super awesome and will be writing all the code, making all the blog posts, and generating all our youtube shorts in 6 months.
- captn3m0 6mo agoLibra/Diem got sold to the bank they were partnering with (Silvergate) for $200M, which then filed for Bankruptcy. https://en.wikipedia.org/wiki/Diem_(digital_currency) https://en.wikipedia.org/wiki/Diem_(digital_currency)
- serf 6mo agowhat's with the negativity? yeah, the metaverse got abandoned. Also: Meta was the only one to try the concept for the past X-umpteen years even though everyone in the industry ga-gas over virtual reality worlds and workplaces at every opportunity. It's literally Meta and Linden Labs (which has been on life support for 10+ years.) The alternative is : no one does it and nothing gets abandoned, which the industry has shown itself to be exceedingly good at w.r.t VR for the past 40+ years. To be clear: I have no faith in meta as a company; my problem lies in kicking an entity because they attempted something different.. I don't think that's productive, and it produces stuff like the past AI winters because groups get afraid of touching experimental concepts ever again lest they incur the wrath of the shareholder.
- ehutch79 6mo agoIt's not the failure here or there, it's a pattern. It's not even the failing, it's the excessive hype cycle. We keep seeing things being overhyped, with not much thought behind it. Meta is particularly bad about it. They changed their name for the hype of their VR product, when VR was still niche and had a long way to go, and still does. They couldn't even figure out legs for launch. Now they have a 'superintellegence'? Yeah, that sounds like just the latest in a line of bullshit. Why would this be different.
- sidcool 6mo agoMeta.ai has muse spark
- tty456 6mo agoI don't get the comments trashing this. If it slightly beats or even matches Opus 4.6, it means Meta is capable of building a model competitive with the leading AI company. Sure, they spent a lot of money and will have on-going costs. But how much more work would it take to turn that into a coding agent people are willing to try (and pay for) along side their usage of a collection of agents (Claude, Codex, etc)? Also means Meta doesn't have to pay another company to use a SATA model across all their products (including IG and WhatsApp, vr) which will matter to their balance sheet long term (despite the constant r&d spend).
- prodigycorp 6mo agoComments trashing this are rightly correct skeptics who remember the benchmaxxing of llama 4. This model was out in the woods as early as like a couple months ago but they didn't release it because it was at gemini 2.5 pro levels.
- zozbot234 6mo agoThe llama4 series was one of the earliest large MoE's to be made publically available. People just ignored it because they were focused on running smaller and denser models at the time, we should know better these days.
- prodigycorp 6mo agothe models were objectively horrible
- NitpickLawyer 6mo agoThey really weren't horrible. They were ~gpt4o, with the added benefit that you could run them on premise. Just "regular" models, non "thinking". Inefficient architecture (number of active out of total) but otherwise "decent" models. They got trashed online by bots and chinese shills (I was online that weekend when it happened, it's something to behold). Just because they were non-thinking when thinking was clearly the future doesn't make them horrible. Not SotA by any means, but still.
- 1970-01-01 6mo agoI can remember when AOL was an unstoppable giant. Except it wasn't. People eventually realized they could get a better, cheaper, faster experience with ISPs and search engines. The same path is unfolding before Meta. People have much better options, and plethora of Meta users will slowly leave until the big moat is drained. Zuck, go retire to your NZ bunker before Meta is forced to merge with another media company.
- hackrmn 6mo agoThe hero image on the linked page, which consists of a muted teal background with the words "Introducing Muse Spark", weighs in at 3,5MB. I don't even...
- Invictus0 6mo agocomplaining about sand on the beach
- hungryhobbit 6mo agoSomeday our robot overlords will be intelligent enough to ... optimize images! (But today is not that day.)
- ruszki 6mo agoThe proper optimization in this case is to not use images at all.
- eranation 6mo agoSo this is why Anthropic rushed the weirdest "pre-responsible-disclosure-totally-not-for-marketing" announcement yesterday? To make sure Spark doesn't steal their thunder? (Spark beats Opus 4.6 on some benchmarks...). Or did I become a bitter cynical old man.
- hnav 6mo agoIt's giving "OpenAI says its new model GPT-2 is too dangerous to release (2019)"
- reducesuffering 6mo ago[because it would start an arms race]. The very arms race we're in... They were right
- signatoremo 6mo ago13 days ago. https://news.ycombinator.com/item?id=47538795 https://news.ycombinator.com/item?id=47538795
- spindump8930 6mo agoYes, it's far more certain that meta released this, which is less convincing on evals, as a result of the mythos previews.
- levocardia 6mo agoAnthropic had their mythos post (and model) basically ready a few weeks ago, as evidenced by the blog content leaks. Also I highly doubt they just threw together a 250-page PDF model card in a "rush."
- dbgrman 6mo agoLast i checked with friends at meta they are pretty deeply invested in using claude for coding etc. anthropic has nothing to be scared of at MSL. If spark beats opus 4.6, why is meta wasting money on opus internally?
- Kuyawa 6mo ago> Meta AI isn't available yet in your country Not my loss, will keep using DeepSeek then. Wake me up when my country is no longer in the wrong/right side of history.
- deleted 6mo ago[deleted]
- bguberfain 6mo agoWe all know it... but I think they were very bold in this warning about using your private messages to train public models. _Your messages with AIs will be used to improve AI at Meta. Don't share information, including sensitive topics, about others or yourself that you don't want the AI to retain and use_
- discopicante 6mo agometa doesn't exactly instill confidence on using personal data responsibly. hard pass
- vinni2 6mo agoI have to create meta account to access. No thanks.
- gritspants 6mo agoI would like someone to tell me how stupid I am. If I were Meta/Zuck I'd open source a great model the moment my company developed it. This just looks like a pitch to investors, otherwise.
- jamiequint 6mo ago"This just looks like a pitch to investors" The goal of public companies is generally to generate profit for their investors.
- samrus 6mo agoIm beginning to think thats the mantra we'll keep reciting as this whole country slowly falls apart
- SoftTalker 6mo agoThis is also the goal of private companies.
- gritspants 6mo agoThank you for telling me how stupid I am.
- kzrdude 6mo agopitch to investors sounds like working for the opposite goal though - to convince investors to give more money to the company.
- ge96 6mo agofunny how websites do that thing where it looks like you can use the product but soon as you hit enter, nope login first
- edwcross 6mo agoWhat is the "BioTIER-refuse" thing mentioned in the "Bioweapons Refusal" graph? I Googled it and found absolutely nothing. Well, to be honest, I got 100% of websites containing the French word "boîtier" (box) with a typo. Even on Google Scholar, the closest match is "BioTiER (Biological Training in Education and Research) Scholars Program", which is at least 10 years old and has nothing to do with that. Is that an AI-generated image with an AI-generated name that has no physical existence?
- EnderWT 6mo agohttps://securebio.org/biotier/ https://securebio.org/biotier/
- nharada 6mo agoSaying nothing about the actual performance of this model, it does strike me how .... minimal(?) this announcement is. Their safety section is like 2 paragraphs about bioweapons. Go look at the reports for OpenAI and Anthropic's model releases. It's like 50+ pages of tests, examples, reports, and benchmarks across a bunch of safety and wellfare metrics. If Meta wants to be seen as a cutting edge massive lab they need to come across as one instead of looking like a school project version of a frontier model.
- WarmWash 6mo agoRumor on the ground is that they expected a much stronger model than this one.
- nubg 6mo agoCan you elaborate?
- WarmWash 6mo agoThat's it. It's just a rumor. A model, which I don't even know of it's this one specifically, fell short of expectations. This rumor came up around mid March.
- htrp 6mo agollama4 behemoth problems?
- levocardia 6mo agoFunny contrast with Anthropic. Ant does a "hero run," gets a model much more powerful than they expect. Meta does a hero run, gets a model much more mediocre than they expect. Read into this what you will, I guess?
- binaryturtle 6mo agoLooks like it needs a meta account? As soon you hit enter it wants to log-in. I guess I won't try this any time soon. :)
- khurdula 6mo ago"we hope to open-source future versions of the model." Love to see it. Cheers!
- tekacs 6mo agohttps://meta.ai/share/pe4HxOfv2Bp https://meta.ai/share/pe4HxOfv2Bp Finding a little bit tricky to evaluate because the harness is unfortunately very, very bad (e.g. search is awful). Can't wait to try this in some real external services where we can see how it performs for real. Definitely getting ordinary high-quality results, overall. But hard to test agentic behavior and hard to test prose quality, even, when just working off of the default chat interface. One thing that stands out is that _for_ the quality it feels very, very fast. Perhaps it's just only very lightly loaded right now, but irrespective it's lovely to feel. I'm quite impressed with the tone overall. It definitely feels much more like Opus than it does, like, GPT or Grok in the sense that the style is conversational, natural and enjoyable.
- nl 6mo agoThis seems pretty good.
- glerk 6mo agoPersonal as in Meta gets your personal data so they can sell you more ads.
- 2pointsomone 6mo ago[flagged]
- CrzyLngPwd 6mo agoIf I'm a claw, then they can send me as many ads as they like.
- flufluflufluffy 6mo agoAnd slowly siphon your personal essence away from yourself and into the model.
- napolux 6mo agoI can't login. It sends me always the same code and it's not correct for them
- nubg 6mo agoNOTHING about this is personal! No weights were released!
- hvass 6mo agoGenuine question: Why release this the day after Mythos? It does not appear SOTA (just based on benchmarks). OpenAI will likely release Spud tomorrow.
- eranation 6mo agoThat's a really good question, my sarcastic mind thinks that Anthropic rushed the Mythos announcement of fears of Meta stealing their thunder... (I guess someone leaked that, a LOT of anthropic folks are ex meta... so, you know) Just a speculation, I have no real knowledge about it.
- MattRix 6mo agoI think Anthropic did the mythos announcement to undercut OpenAI’s upcoming next model announcement, not Meta’s.
- MattRix 6mo agoWhy not? Not everything has to be SOTA to be interesting.
- paxys 6mo agoMythos is a news article. This is an actual model you can use.
- syntaxing 6mo agoKinda crazy, it really felt like Meta had the lead in LLMs, especially during the early LLaMa days. What happened for them to fall so far behind? I don’t get how LLaMa 4 was such a big train wreck and they couldn’t correct the course like Google.
- aivillage_team 6mo ago[dead]
- plombe 6mo agoLooks like a lightweight article. But memory usage went from 316MB -> 502 MB when I hit refresh. Not sure why? Any one have any ideas? Why does it need half a gig of ram in the first place?
- GalaxyNova 6mo agoIt is unfortunate that they decided to stop doing open-weight releases. What could have been interesting has been reduced to simply another subpar LLM release.
- dhruvyads 6mo agoSad to see it's not going to be open source.
- eranation 6mo agoSarcasm aside, tried it (with instant mode), it's an impressive model. It nailed all the ChatGPT meme gotchas (walk to the carwash, Alice 50 brothers, upside down cup, R's in strawberry, which number is bigger, 9.11 or 9.9?) I guess all that money poaching OpenAI / Anthropic talent went somewhere... Now, would I use "Meta Muse Code" or "Muse CoWork" if I have to have a facebook account to all of my developers? Maybe not. Would I use it via an API key? I might, depends on the pricing!
- turtlesdown11 6mo agoso since they hard programmed all of the meme gotchas, they built a good model?
- nh23423fefe 6mo agolazy snark < playing around with it
- yalogin 6mo agoMeta is in a weird spot. They caught up late to the game and instead of releasing llama as a chat bot they open sourced it, precisely because they lost the mind share. They thought chatbot is not their product and I am sure they are regretting it now. Mark is obsessed with becoming the android of something and he poured billions into the metaverse thinking he is first and failed. He then open sourced llama and wanted to be the android of llms. He ended up enabling groq but it didn’t benefit meta directly at all. They have no revenue or mind share path from llms but continue to pour billions into it. The only 1-1 mapping is with the glasses but that is a tough fit for the company given they are extremely allergic to privqcy and security. Not sure what this is now.
- deleted 6mo ago[deleted]
- gardnr 6mo agoThe llama weights were leaked. It open sourced itself. You are right though. Meta could have been in lockstep releasing ChatGPT features into some chat bot on Facebook.com but instead it seemed like their FAIR arm was hell bent on commoditising this stuff by publishing their research models before the Chinese companies took the lead in that. It’s hard for me to be mad at FAIR even though I general disagree with the outcomes that Meta produce for their users.
- IceWreck 6mo ago> He then open sourced llama and wanted to be the android of llms. Well the original llama did kick off the era of open source LLMs. Most original open source LLMs were based on the llama architecture. And look where we are now OSS modles are very close to frontier. It may not have benefitted Meta but it commoditizatised LLMs.
- solarkraft 6mo agoHell, most of us are still using llama.cpp for inference in some form
- btown 6mo ago
- BugsJustFindMe 6mo agoI'm struck by all these independent announcements saying "look at our new model that we only spent $N Billion in acquisitions and hardware time to build and operate that's just like those other ones but this one is ours." Because if any of these companies would simply pool resources and work together, and if the government actively participated in providing funds, they'd be able to accelerate AI so much faster. It all feels incredibly wasteful. But I guess that's communism or something.
- victorbjorklund 6mo agoCompetition often foster innovation. Why are they innovating so fast and spending so much money? Because they don’t wanna get behind. If there was no competition at all then there would be much less reason to innovate and spend resources.
- BugsJustFindMe 6mo ago> Competition often foster innovation. So does cooperation in any framework that values public good over pure obedience to an inherently-abusive late stage capitalism. I know that's passé in a world where the US government no longer believes in funding science, and yet. Competition is also inherently wasteful. And if you're talking about wasting a few K or a few Mil here or there, fine, whatever. But here we're talking about waste on the order of trillions of dollars at the end of the day.
- nathan_compton 6mo agoTheir product could literally teleport gold into my hands and I wouldn't use it.
- cvhc 6mo agoCan't login. No error message in the UI. But the URL changes to "https://www.meta.ai/?error=Token%20exchange%20failed https://www.meta.ai/?error=Token%20exchange%20failed".
- cvhc 6mo agoI switched to Chrome (from Firefox) and tried again. Now it's "https://www.meta.ai/?error=Invalid%20CSRF%20token https://www.meta.ai/?error=Invalid%20CSRF%20token" :facepalm:
- nh23423fefe 6mo agosame. closed tab and will forget to ever use it now
- fede_dp 6mo ago[dead]
- spprashant 6mo agoSounds like a good effort. They are choosing to focus on multi-modality - perhaps they are taking a different route here to Anthropic. I don't like that I need to login to my FB/Instagram account to access this.
- spearman 6mo agoUploading images requires logging in. Logging in is broken. It redirects to https://meta.ai/?error=Token%20exchange%20failed https://meta.ai/?error=Token%20exchange%20failed and doesn't show any error message. Impressive.
- brianmcnulty 6mo agoIt has been up and down today, specifically with authentication breaking. I also saw an error message with backend SQL in it (in my 6 years of Meta bug bounty security research, I have never once seen backend SQL before). I suspect it is because they also refactored Meta AI entirely to use Next.js instead of their normal stack they use for literally everything else. Not sure why they would do this, but I guess it works (...or maybe not) for them.
- btown 6mo agoBenchmarks are meaningless until the pelican benchmark comes out: https://simonwillison.net/ https://simonwillison.net/
- dbgrman 6mo agoLitmus test: what % of meta engineers are using muse vs Claude code? Last i heard it was mostly claude code. Tell you everything you need to know about how serious these benchmarks are.
- upmind 6mo agoSure it's not as good as Claude right now but for their first model in years it's certainly not bad. I hope they continue to develop models, having another competitor in the space would be nice.
- granzymes 6mo agoComes impressively close to GPT 5.4 / Gemini 3.1 Pro / Opus 4.6! Mostly behind OpenAI on coding/agentic benchmarks, behind Google on text reasoning, behind Anthropic on Humanity's Last Exam with tools (surprisingly the only benchmark where Anthropic leads currently). Meta hasn’t fully caught up, but they came close and I think can solidly claim to be a frontier lab again. I’d call it a 3.5 horse race right now, and hopefully their next model improves. More model competition is good! Poor Grok 4.2 should probably be dropped from the table.
- fancy_pantser 6mo agoIt's looking rather low on reasoning and long-range problems with the approach described. For example, even with 16 agents and compaction, the HLE score is significantly below Anthropic's Mythos. Like you, I can see the release as a net Good Thing, but apples-to-apples for each org's latest models do have Meta holding steady in the middle pack.
- zozbot234 6mo agoHLE encompasses very hard problems where the larger pretraining of Mythos probably matters quite a bit. I'm not saying that Mythos is not showing some amount of genuine improvement compared to e.g. the latest Opus; just that if you're going to compare models, you should at least make sure that the overall test-time workload is in the same ballpark given how high it seems to be for Mythos.
- deanc 6mo agoGrok code was my daily driver for months while it was free and it was fantastic - it is certainly no worse than it was a few months ago. Unfortunately with LLMs everything is based off your use case, domain and the context you give it. I also use Grok daily for health questions as the other models are too afraid to give input on medical matters
- flufluflufluffy 6mo agoWhy do you need to ask any AI questions regarding your health every day?
- redlewel 6mo agoI am already somewhat concerned with companies like Anthropic and especially OpenAI having personal data via chats. Typing that sort of information into a Meta AI product feels completely irresponsible. You could make some very sophisticated ads/psyop attacks with data from daily ai chats. I doubt its better than Opus and even if it was its not worth the privacy concerns.
- LZ_Khan 6mo agoOne word: distillation
- pixel_popping 6mo agoMeta back in the commercial race is actually exciting, despite not being a fan of the company.
- leumon 6mo agopelican riding a bicycle (svg): https://files.catbox.moe/u5yc0x.png https://files.catbox.moe/u5yc0x.png
- TobTobXX 6mo ago> Muse Spark is a natively multimodal reasoning model with support for [...] visual chain of thought [...]. Do they mean "the chain of thought is visible to the user" (ie. not hidden like ChatGPT), or "the medium of the chain of thought is not text, but visuals" (ie. thinking in images). I'd guess the former, since it wouldn't be economical to generate transient images, just for thinking. But I'm not sure why they'd highight that in that case. If it were the second thing, that'd be extremely interesting. The first model not to think in text.
- rain-princess 6mo agoActually I believe that behavior shows up in Gemini chats (if you are doing a visual task) it will generate intermediate diagrams and research papers have created approaches to that effect (generating turtle diagrams) since 2024
- fc417fc802 6mo agoPerhaps more importantly, will their chain of thought be "real"? So far the ones I've seen seem to be elaborate fakery. They look good unless you dig in at which point you often find that it merely looks plausible on the surface but that something else is going on under the hood.
- seanhunter 6mo agoI don't know what you mean by that. We know what's going on under the hood always: linear algebra, the attention mechanism etc. To my first approximation all "Chain of thought" means is that instead of having to prompt the model to discuss everything in text and then decide at the end[1], now it sort of automatically does that so you don't need to prompt it. [1] Which used to bring about very substantial improvements in performance on some tasks
- fc417fc802 6mo agoI think it was clear from context that "under the hood" wasn't referring to the math but rather to the contents of the trace. What's written (often?) isn't what's actually being "thought" about. The trace is a trained output similar to the final output, which is to say that it's fake. There are research papers on the topic, particularly that models can be trained to print other arbitrary stuff during the "thinking" phase instead. You can easily see this for yourself by carefully walking through a given trace with a critical eye. Here's an example from myself a few days ago. https://news.ycombinator.com/item?id=47623324 https://news.ycombinator.com/item?id=47623324
- KoolKat23 6mo agoPerhaps I'm wrong, but definitely seems to be SOTA. Although looking at it's ARC-AGI-2 score it's reasoning isn't very good. I suspect it's got the benefits of scale but lacks that human added element, understandable considering they claim to be building it from the ground up. This should come in time if they have a good team. In real life, I'd imagine one would worry about overfitting when using it. (I'm not using it as I'm not agreeing to their ad terms).
- adt 6mo agoCongrats to the Meta team on being model #800 on the Models Table, I suppose. https://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- 2001zhaozhao 6mo agoThe "AIME Evolution" graph seems interesting. I wonder if other labs are doing this too to improve the reasoning performance of their models. > Think longer to solve harder problems > Compress > Think longer again
- try-working 6mo agoLooks alright for a "first" but there's no reason for anyone to really use until they open source it.
- anigbrowl 6mo agoKinda off topic but I wonder why they picked this name, knowing of Nvidia's Spark. They're different products, obviously, but the potential for confusion is real as both brands are competing for mindshare in the AI space. I opened this story expecting to read they'd deployed on a cluster made of Spark machines or somesuch.
- supermdguy 6mo agoAnd also OpenAI’s codex spark?
- zmmmmm 6mo agoThe real question for me, if we assume they once again have a competitive frontier model, is what this means for Meta's strategy now. In particular, have they abandoned all their philosophy of the open ecosystem / open model play they were pursuing before? While it's true, llama4 sucked, I still can't help feeling they have lost ground compared to where they would have been if they maintained that strategy. Due to llama, they were considered a peer with the other frontier model providers. Now they are not even in the conversation. It would take an incredible shift in performance to make me even consider using their new model. They may have a model, but the other providers have been busy building whole ecosystems around their tech which Meta has none of. Maybe they could dump $1b into OpenCode or something and reignite the open ecosystem play with an open harness. They need something to get back in the conversation, if that's where they want to be. Otherwise, it will just be another closed, hidden proprietary AI model driving user facing Meta apps, but which nobody else cares about.
- jatora 6mo agono need for an open harness when anthropic so kindly gifted the community theirs :)
- simonw 6mo agoPelicans: https://simonwillison.net/2026/Apr/8/muse-spark/ https://simonwillison.net/2026/Apr/8/muse-spark/ I also had a poke around with the tools exposed on https://meta.ai/ https://meta.ai/ - they're pretty cool, there's a Code Interpreter Python container thing now and they also have an image analysis tool called "container.visual_grounding" which is a lot of fun.
- wsgeorge 6mo agoAlexandr Wang suggesting this might be open-weights/source in the future gives me hope. Hopefully they stay on this path.
- lemonish97 6mo agoI have a feeling it won't be this exact model, but rather smaller distilled variants, similar to the gemma line
- sbinnee 6mo agoIt is fair to think so because that is what everyone is doing. But being Meta and considering Llama, if MSL is going to keep releasing models and wants to join back the AI war, they may actually open weights just to get more attention. Once they establish a sizable community, they can start guarding their frontier models.
- sbinnee 6mo ago> but you can try it out today on meta.ai (Facebook or Instagram login required). I guess I will have to wait. I hope at least soon it will be available on Openrouter. Overall, I am really excited to try it out.
- nickvec 6mo agoThe only benchmark I care about! Just curious Simon - which model do you think has created the best pelican riding a bicycle thus far?
- laser 6mo agoFirst thing I tried is a visual reasoning test on floor plan documents that applies directly to something I'm working on and needed that I posed to ChatGPT, Claude, Gemini, and Grok yesterday (lowest tier paid plans on each). In that test only Gemini succeeded while the other models hallucinated/incorrectly reported the relative location of building units. I just posed the identical prompt/document to Muse Spark and it knocked it out of the park, extracted and displayed the pertinent pages from a multi-page PDF inline in the chat and rendered a correct answer. This may be a one-off or lucky start but given the incredible result out of the gate I'm optimistic and will continue testing in parallel against other models before potentially making it my primary daily driver, excluding coding where the harnesses of claude code and codex are still needed (although hopefully they release something in this space too). That being said Meta has the most adversarial data-usage policies I've seen among LLM providers so that's unfortunate for handling anything sensitive, but it also stands to reason that they have a long term advantage with such a massive proprietary data set. I'd prefer to also have a paid plan like the other services that allows me to keep my data out of training, rather than a free service and my usage being monetized in other ways.
- RandyOrion 6mo agoNo open weights. Besides, I'm old enough to recall that META has trained a version of LLAMA 4 specifically for LM arena elo benchmaxxing and PR things, and proceeded to release a different version of LLAMA 4.
- senor26 6mo agowho trusts meta on anything!!
- thebiggestloser 6mo agowhy is it behind a login? Such bad UX.
- leentee 6mo agoLook at their benchmark charts to understand how desperate they're. A lame duck now.
- anxtyinmgmt 6mo agoI wanted to root for Mark and Meta as another frontier lab especially focused on open source but at this moment I have to say who cares. Gemini has a better OS track record thus far. Alex Wang is a reputational hazard. It is hard to get over the bias that this too might be benchmaxxed. I'd love to see demos of products actually using these models to overcome that but with the current pace of progress now my intuition says skip all this.
- maxaravind 6mo agoPersonal superintelligence sounds nice until you actually try to use it. We spent time yesterday arguing through an architecture decision. Today I ask the Agent to help implement it - it knows nothing about any of that. You’re effectively starting over. Feels like the real problem isn’t intelligence, it’s continuity. And most benchmarks don’t even touch that.
- manyminds 6mo agoYes this feels very new from a product and harness design perspective but it's brand new! Nine months old. The mobile and web sessions don't even real-time sync between each-other yet there's endless work to be done and time will tell if they can bring all the people to bring it together. The underlying model seems like a great foundation now but securing the supremacy of usage is multilateral requiring both machine learng advancements and product/harness/usage design.
- samarth0211 6mo ago[dead]
- gloosx 6mo ago>Text field. >"Ask Meta AI..." placeholder. >Colourful blue Send button. >Eager to try, entering question... hitting Send. >Log in or create an account to access. >15 seconds of loading time >Continue with Facebook or Instagram Typical meta move, throwing a dark pattern at you from the beginning instead of just letting you try it Won't even bother to continue, somehow OpenAI got this right.
- damian_pol 6mo agoI hate that they ask to log in with facebook/instagram account. I tried to create a new one with proton's hide-my-email and it got suspended 30 seconds later. When I tried to log in they require a selfie proving that I am not a robot. Ridiculous that in order to use dev tool you need to link it to social account or send a selfie
- deleted 6mo ago[deleted]
- supermatt 6mo agoDoes "personal" here mean "run the model on your personal hardware", or just "give your personal data to meta"?
- ssxx 6mo ago[dead]
- RestlessAPI 6mo agoToken cost really matters here. I want too know what API pricing is. As we see, this model is like 85% as good as the frontier models? What if its priced at $0.2 in / $0.5 out Mtok? All of a sudden, this model is A LOT more appealing to me.
- voidUpdate 6mo agoWhat makes this "superintelligence" instead of regular artificial intelligence?
- Alifatisk 6mo agoDo we have any numbers on input, output and conversation context window limit? I tried multiple riddles, graphs and questions I know some LLMs fails at, but this one seems to do well. But I still don't have much trust in Meta after the scandal of them fiddling with their previous models to look good.
- ForgeSynapse 6mo ago[dead]