13 ms·
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
- wegothimyay 2mo ago[flagged]
- tosh 2mo agogood to see new open weights releases from meta
- jauntywundrkind 2mo agogood looking showing too, which is excellent.
- InfiniteLoup 2mo agoThe least they could do, after ruthlessly bombarding my employer's servers with requests, ignoring the robots.txt, scraping everything, and incurring significant Google Maps costs for us in the process.
- ninjin 2mo ago2a03:2880: by any chance?: https://news.ycombinator.com/item?id=48137854 https://news.ycombinator.com/item?id=48137854 Have asked them to stop numerous times and they just keep hitting for about eight months now.
- richardfey 2mo agoLooking forward to giving this a try with llama.cpp. I’m watching the open-weights competition with high expectations.
- gunalx 2mo agoMeta did not abandon opensource. I would love to see a smaller distill, or a moe of this size but the benchmarks seems competetive as long as it isnt benchmaxed witch i would not be suprosed if it is.
- ignoramous 2mo ago> Meta did not abandon opensource Open weights* I don't think outside of the Big 3 (Ant, OAI, GDM), given the strong competition from China, any other Lab has a chance at capturing the coding market if they aren't open weights (save for xAI whose latest Grok looks every bit good & will probably rely on Cursor for distribution instead of going open weights). There's literally no other selling point, as the capabilities have mostly converged by now among the chasing pack.
- ComputerPerson 2mo agoThere was a good discussion yesterday on the DeepSeek Flash release thread about this. There's a large market, very large, who want the best regardless of what it costs. Probably a large enough market to keep that domain of research afloat (as opposed to shifting research manpower to cost cutting). The reasoning is just that the marginal cost of AI is very secondary to fixed costs of the businesses themselves; it's not an excuse to sacrifice performance.
- dannyw 2mo agoDon’t sleep on NVIDIA and Nemotron. It’s not completely open source, but they actually release their pretraining and post-training datasets with some redactions for (cough) pirated content. They also have very good code and playbooks for actually doing a fine-tune, CPT, etc. Even if you’re not tuning a Nemotron model, its mixes are very excellent for your replay data slice; or general experiments. Way better curation and quality than Dolma, etc; or other large huggingface data mixes I tested.
- nickludlam 2mo agoYes, I second Nemotron. I'm using Ultra remotely and Super locally, and I find them very useful for RAG-like problems. I wouldn't really use them for coding.
- HardCodedBias 2mo agoGDM -- Ok, I'll bite. Why are you including them?
- scrlk 2mo agoWill be interesting to see how Qwen3.8 27B compares against this once it releases this week. Seems like dense 30B is back in fashion? EDIT: An open weight version of Muse Spark 1.2 is going to be released as well: https://x.com/alexandr_wang/status/2086756152034066792 https://x.com/alexandr_wang/status/2086756152034066792 https://xcancel.com/alexandr_wang/status/2086756152034066792 https://xcancel.com/alexandr_wang/status/2086756152034066792
- wronglebowski 2mo agoIt’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.
- dannyw 2mo agoQwen thinking is really good in Mandarin; and probably natively trained the most there. Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language.
- kadoban 2mo agoIs the quality of the thinking better or it's just shorter since Mandarin is more compact?
- yiyu_earth 2mo agoThis is most likely because the vast majority of the information the model absorbed during training was in Chinese. As a native Mandarin speaker, I frequently need to convert the prompt into English and output it in English in order to avoid that the model falls back into Chinese reasoning logic. PS: Switching the thinking process from Chinese to English can also significantly circumvent certain self-censorship mechanisms built into the model.
- ComputerGuru 2mo agoJust to play devil’s advocate: you can’t compare Qwen to a (proprietary/closed source) hosted model and deduce that Qwen is overthinking, as Qwen gives you the full reasoning/thinking trace while all the proprietary models now give you only a summary “to prevent distillation”, making it hard to properly compare apples to apples here.
- Gecko4072 2mo agoWhat I think would be perfect is a model that could run on a single DGX spark and be competitive with DSV4 Flash 731. Flash is already a game changer. Hopefully meta plans on this, like the old 70b. V4 flash is smart enough for any use but slightly too big. 27b-30b isn’t intelligent enough.
- 127 2mo agoDSV4 Flash 0731 already runs on RTX 4090 24GB + 128GB system RAM at a usable tok/s and quantization.
- Gecko4072 2mo agoYou personally? Just curious. Context window is also a factor and ram isn’t really cheap. Sparks are assembled units which I like.
- dannyw 2mo agoFor the same price as a DGX Spark here (A$8499) I can buy roughly 544GB of DDR5-5200MHz from retail; which on a quad channel platform would deliver ~160gb/s real world; and ~320gb/s with octa channels (Xeon, Threadripper Pro). If you can afford it or somehow find a used unit, you can go Epyc for 12 channels. 8/12 channel DDR5 will beat DGX Spark in inference/decode even without a GPU of any kind, as it’s memory bandwidth bound, and the Spark tops out at ~240gb/s real world. With some optimisation and maths, it’s entirely plausible to ach You are paying an extraordinary amount of money for the convenience of a super small unit, with still mediocre software support, but at least a community. Expect to be crawling through forum posts regularly, as SM121/Spark has many quirks and ecosystem issues still. Please don’t pay another 70-80% gross margins on top of already inflated DRAM prices unless you need. The Spark IS really nice if you want to test out ConnectX or if you really need something small and compact and quiet. Also consider: used Adas or even Ampere NVIDIA workstation GPUs can come with a lot of VRAM and be “reasonable”, with CUDA.
- danielEM 2mo ago
- sajithdilshan 2mo agoStill needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
- Gecko4072 2mo agoThere have been discussions on language specific not really being a relevant change to reduce size.
- Manfrednotfunny 2mo agoI would love to see any good research projects about it but i have the feeling that Frontier with MoE is making too fast of a progress so that a customized model would always be worse and that the MoE part is actually going somehow in this direction. On the other hand, at the GTC was a talk about coding in different lanugage (like spanish) and explaining that the quality between spanish and english is relevant different. But i have not found a good article about the impact of learning data with practical experiments or even if the order of the learning data matters. At least I think i remember that Meta mentioned having better and less data can be better than more data with lower quality. As long as these models can explain to you facts about any other topics, its still overfitted for the task though.
- mapontosevenths 2mo agoCapability in LLM's is distributed throughout the manifold in subspaces. Even worse, the subspaces exist in superposition. That is to say, there is no single 'python' part of the model. The python bit is spread throughout the entire model and overlaps with other pieces that have similar, but unrelated, capabilities. For example the python subpspace might be partially in superposition with cupcake recipes, Esperanto, and calculus. We need calculus in a coding agent but not the other two. However, separating them cleanly is almost impossible, and even identifying them is tough. Internally the manifolds are highly inefficient and nothing like you would imagine something humans built would be designed. It's more like something that evolved in nature.
- solarkraft 2mo agoWow, Meta is back (at least for now)! I like this class of model. Multi-token prediction makes it viable to run dense models at not-too-far-off speeds as MoE models with much better intelligence. The submission’s title (open weights 30B local coding model) is luckily wrong: This is meant to be a general agentic model. It even comes pre-quantized and with a MTP/drafter model. Looking good! Let’s hope they aren’t dishonest with the benchmarks this time …
- akazantsev 2mo ago> The submission’s title (open weights 30B local coding model) is luckily wrong: This is meant to be a general agentic model. https://xcancel.com/alexandr_wang/status/2086756152034066792 https://xcancel.com/alexandr_wang/status/2086756152034066792 It's correct. See the OpenCode demo. Generic models are good enough for coding without necessarily being designed specifically for coding.
- solarkraft 2mo agoRight, so it's as correct as me claiming it to be an E-Mail sorting model. It may be good at that, but that's not its primary purpose.
- bwfan123 2mo ago> It even comes pre-quantized and with a MTP/drafter model Glad to see the extra engineering effort that went into creating this local model and making it run well on a consumer device. I use qwen3.5-coder, and am waiting to kick the tires on this one. I hate to say this, but kudos to Meta ! I hope apple and others follow suit and create similar local models for other use cases like audio, images and video that can run on a laptop.
- Havoc 2mo agoThe favourable comparisons to Gemma 4 and qwen3.6 look promising!
- cmrdporcupine 2mo agoThose two offer MoE variants, this doesn't seem to. Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant.
- petu 2mo ago3090/4090 probably would do 40 t/s, for 5090 75 t/s is shown in the blog.
- EddieRingle 2mo agoDense models run at a very usable speed (Qwen 3.6 was running at ~50t/s last I looked) on my dual 7900 XTX desktop. (And before anyone brings it up, I did not buy them for this purpose, so the up-front cost is irrelevant in my case.)
- Havoc 2mo agoThe benchmark comparison is against the dense variants not MoE
- nutjob2 2mo agoThe more open weight models get released the greater the market for personal and small business oriented hardware to run these models. This will drive lower cost hardware, which has stagnated in recent years due to most software not needing the performance and capacity.
- cmrdporcupine 2mo agoThe opposite happening because foundries are full to capacity making higher margin stuff.
- nutjob2 2mo agoYou have to look a little past the current hysteria.
- grim_io 2mo agoHigher demand for 5090's did not make them cheaper, because Nvidia got much higher margin products to focus on.
- jkwang 2mo ago[flagged]
- zmmmmm 2mo agoMeta knows how to win back developer's hearts .... let's see if they have the goods
- xandrius 2mo agoIf there is anything meta can do to regain hearts other than owning up their evil deeds, radically change their business model and paying up for taxes and damages, then the world is truly fucked and corporations will continue to win.
- maxignol 2mo agoOptimizing speed is really the way to go. Yet 24GB is not what everyone can afford. Maybe we could take some of those 56tk/s and transfer into some free RAM space using MoE loading ? I'd be glad with a less than 10GB and more than 6tk/s model.
- Manfrednotfunny 2mo agoI don't thinnk just MoE will solve it. If you hit constantly different expert layers, you can't outsource layers efficently and have to swap it in. MoE will be faster because it will read less memory for sure, you still have to have it though.
- lisplist 2mo agoUnfortunately this is just the entry price for LLMs. With the exception of the Qwen 27B models, I personally haven’t found a ton of use cases for models less than 200B. With the right setup, fine tuning, etc, you can make small models do cool things, but hard to please everyone given the insane hardware costs at the moment and the comparably cheap API costs.
- dist-epoch 2mo agoGemma4-E4B (4B params) works pretty well as a local wiki, or when you don't have connectivity.
- dannyw 2mo agoNitpick: Gemma4-E4B is actually a 8 billion param model, but only 4.5B params worth of memory bandwidth needed per decode.
- dannyw 2mo agoSmall models are still great for lots of “simple intelligence” use cases, like annotating or summarising files and media; or even just basic chat when given web search tools. My local NAS is private and I’m not going to send it off to APIs for captioning or metadata; but even Qwen3VL 8B does an excellent job at this, despite being quite old. They are also really excellent for fine tuning. Unsloth and Tinker (from Mira’s TML) are great places to start. If your use case is narrower than “coding agent for everything”, you can probably match frontier performances on that narrow domain with ~30b and exceed it with ~100b+.
- bronxbomber92 2mo agoI wish they would release the quantized versions in a safetensor format. Many frameworks can't load PTE and GGUF.
- ddzzz 2mo agoTry quantizing yourself, just a suggestion. https://github.com/pytorch/executorch/tree/main/examples/models/muse-glimmer#:~:text=For%20a%20custom%20local%20quantization,%20export_solo%20also%20accepts%20a%20consolidated%20BF16%20checkpoint%20with%20--checkpoint-dir%20and%20a%20recipe%20selected%20by%20--quant-recipe). https://github.com/pytorch/executorch/tree/main/examples/mod...
- _ache_ 2mo agoIt is interesting but it does look like a careful distillation of (Spark and) biggers open-weight models. The progress compared to Qwen3.6 27B is good, not that impressive, it's a 4 months old model. (kuto to them to compare to 27B dense and not 35B MoE, it's more fair to do so). It is very probable that Qwen3.8 27B will crush Glimmer-30B on most benchmarks.
- skohan 2mo agoStill great if they want to play in this space. Having competition for the 24-32GB VRAM target is only good for the end user.
- drob518 2mo agoAgreed, the trend in this consumer-accessible range is encouraging.
- pettijohn 2mo agoI'm so excited about these two new models. Qwen 3.6 27B has been my sweet spot so I cannot wait to try 3.8. Glimmer looks really strong, I'm encouraged that Meta compared it to 3.6 in the model card! Exciting times!
- petcat 2mo agoAs an industry, I wish we would stop calling these things "open weight" because it is too easy to confuse with actual "open source", which they are not. Photoshop source code+ OSI license = open source Photoshop binary you can run on your own computer = open weight Photoshop SaaS web app = closed, proprietary (Opus, GPT, etc.) "Open weight" models are still just binary blobs that are completely inscrutable. It's like bringing home a dog from the rescue and just hoping that it doesn't have a tendency to bite kids in the face. You just can't know. The only thing that you can do is try to add more training (fine tuning) telling it not to bite kids. I don't think the FOSS community has ever accepted this, but somehow we're feeling like it is okay now.
- piker 2mo agoIt is useful to indicate you can run the weights on your own hardware. That’s categorically different from most other commercial offerings. It’s as if your adobe example ignores the reality that would exist had photoshop been invented in 2019: cloud only.
- microtonal 2mo agoPhotoshop source code+ OSI license = open source Photoshop binary you can run on your own computer = open weight I don't think this is a correct analogy. You are not allowed to distribute modified versions of the Photoshop binary. Most open weight model licenses allow you to make and distribute your own finetunes, etc.
- monster_truck 2mo agoThis analogy is terrible and seems to be extremely misinformed about how rescues evaluate dogs before they are put up for adoption
- petcat 2mo agoI am extremely well aware of how rescues evaluate dogs. And I'm also fully aware that they do not know the full history of the dog. They go through a limited set of testing and interrogation to evaluate the safety of the dog. That's it.
- bentt 2mo agoMeta seems like the one American bigtech that would distill the the other American frontier models. My enemy’s enemy is my friend?
- grim_io 2mo agoThey do distill, their own bigger Muse model.
- dev_daftly 2mo agoYou think the company buying up all the books, cutting off the bindings, and feeding them through a scanner isn't also distilling other models?
- Maxious 2mo ago> Some have tried to frame distillation as harmful, but I think it is important to protect the principle that you can learn from anything you can observe. - Mark Zuckerberg https://www.meta.com/thefutureisforeveryone https://www.meta.com/thefutureisforeveryone
- vibe42 2mo agoMeta released their own 4-bit quant of this model for devices with 24GB VRAM. That's a modern gaming laptop; cheapest I see in the US with 24GB is $3.5k. Should be quite a bit faster than the new M5 MacBook Pro, and you can run Linux on it!
- sgt 2mo agoCan I run this on my RTX 5090?
- skohan 2mo agoYes they have quants for 32GB and 20GB use-cases (including mmproj and kv cache + context)
- polymorph1sm 2mo agoSome interesting findings from the chat template designs: 1. The template name is Onyx ATEM as found in the tool call exception message 2. It appears to be following a harmony-style chat template. But the tool use seems to be a xml like :<atem:function_calls> / <atem:invoke> / <atem:parameter> 3. atem: a internal joke of meta in reverse? https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/chat_template.jinja https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/mai...
- dannyw 2mo agoThe XML tags are similar to <antml:xxx>, which is obviously Anthropic ML (or ANTrophic xML). I think it’s likely 3; meta in reverse. While tokenisers and preprocessing can catch it, you want your special tokens to be unique and not present in the original corpus. <meta: is likely too common.
- jszymborski 2mo agoLikely inverted "meta" to avoid collision with HTMLs meta tags
- dudefeliciano 2mo agoatem also means breath in German
- kristjansson 2mo ago> atem also perhaps taking some small joy from the lexical similarity to aten[0] namespace that lies at the heart of pytorch [0]: https://github.com/pytorch/pytorch/blob/main/aten/src/README.md https://github.com/pytorch/pytorch/blob/main/aten/src/README...
- cmiles8 2mo agoWith the business model for API based LLMs looking iffy at best it seems like we’re heading back to the “server under your desk” era of IT again.
- cube00 2mo agoConsidering how all the big players are playing fast [1] and loose [2] with limits, billing [3] and adding undisclosed changes that burn your tokens on autopilot [4], it can't happen soon enough. [1]: Limits may change without notice, including due to capacity constraints. - https://support.google.com/gemini/answer/16275805?sjid=14713165163333653664-NC#zippy=%2Cusage-limit-changes:~:text=Limits%20may%20change%20without%20notice%2C%20including%20due%20to%20capacity%20constraints https://support.google.com/gemini/answer/16275805?sjid=14713.... [2]: "standard limits" are never defined - https://support.google.com/gemini/answer/16275805?sjid=14713165163333653664-NC#:~:text=an%20AI%20plan-,Standard%20limits,-AI%20Plus https://support.google.com/gemini/answer/16275805?sjid=14713... [3]: https://tobyonfitnesstech.com/blog/anthropic-refund-scam/ https://tobyonfitnesstech.com/blog/anthropic-refund-scam/ [4]: https://news.ycombinator.com/item?id=48947776 https://news.ycombinator.com/item?id=48947776
- xscott 2mo agoNot to mention all the other ways they can screw you: - Middle of the day, servers busy? Swap to Sonnet while pretending it's still Opus. Many people won't notice, and nobody can prove anything if they suspect. - Middle of the night, server load is light? Put it into extra thinky mode so it burns more tokens to ramp up the bills. Flip the switch where it gets really pedantic about writing lots of extra test cases and verifying against documentation. - Demand increases, but don't feel like running more hardware? Switch to low bit quants, but have a monitor model swap back to quality if it can tell you're running a benchmark. Assuming model capability plateaus (I think it will), token providers will be in a race to the bottom to maximize profits at the expense of quality that's very difficult to measure.
- dannyw 2mo agoThese kind of tricks will completely break API customers and be super visible, since most companies deploying API at scale have ample telemetry, evals, etc. Although, selectively applying it to consumer subs is probably beyond likely at this point.
- avaer 2mo agoI lament the comments saying this in any way redeems Meta (the company). The researchers releasing this stuff have almost nothing to do with Meta other than being bankrolled by the slaughterhouse. You aren't the customer, you are the pawn in big tech's game of thrones. Your good will is a commodity to be traded, almost literally. It will be used against you the moment it's convenient. This is open weights because Meta couldn't monetize it in any other way than to cloud developer's judgement of their reputation. But I guess most people just don't care. I'm glad it's open. It does not make me think any better of Meta.
- darig 2mo ago[dead]
- captainbland 2mo agoTo be honest the main issue with meta has never been around open/closed software. They've also done react, Cassandra and some other bits. But this, like their open weights is like a feather pressing down on the scale compared to things like promoting genocide in Myanmar, enabling Cambridge analytica, creating a huge closed ecosystem which dominate(s/d) local community communication, mandating doxxed communication, trying to replace actual community communication with algorithmic nonsense etc.
- larodi 2mo agoMeta and its products, as a whole, is a threat to your kids, your mental health, your community's health and the planet as a whole. It is just sad and very repulsive everyone fell so easily addicted to their social drug. Yes - it is a drug, and it is hard to get off from. Nothing redeems them at this point of time, they are doing exactly ZERO to redeem. Tossing open weight models (not opensource!!) is not a basis for redemption, and does not constitute remorse in any way. Trying to portray it as such is complicity to META's crimes against humanity.
- Grombobulous 2mo agoI think it’s also worth pointing out that that there are numerous less evil options to choose from. Perhaps none of the AI companies are shining examples of high ethics, but basically all of them have ethical high ground over Meta. At least Anthropic isn’t sending private videos from pervert glasses to contract workers in Africa. It’s a low bar but it’s a bar nonetheless.
- OsamaJaber 2mo agoThe comparison set is Gemma4-31B and Qwen3.6-27B, not the current Qwen Fair on size, but the headline numbers are against a model a generation back
- Zambyte 2mo agoWhat more recent open weight Qwen release is there?
- NorwegianDude 2mo agoThat is the most recent Qwen and Google models, there is no newer version, yet. Qwen3.8 27B might come in a couple of days tho, if it's launched alongside the large one when the Qwen3.8 countdown reaches zero.
- korykaai 2mo ago[flagged]
- moron4hire 2mo ago"Meta Muse" immediately made me think of Metamucil. Product teams really need to hire at least one or two people with a 12-year-old's sense is humor. They need to winnow all the potential stupid jokes out of their product namings.
- eugene3306 2mo agowill it run on 2x 5060Ti with 16GB each?
- skohan 2mo agoIt should - the kquant-dynamic variant is targeted towards 32GB. Downloading it now to give it a try.
- leansensei 2mo agoIt does, beautifully. Now let's wait for an NVFP4 GGUF!
- BoredomIsFun 2mo agoyes. you can even parallelize two cards and get 1.7 times the speed.
- mirekrusin 2mo agoGreat to see Meta back, looks like really strong, local model, can't wait for llama.cpp support.
- jakswa 2mo agosome support already merged, and I verified in a local build that it runs (cannot get MTP params working tho, about ~40 tok/s on my beefy 800GB/s 7900XT w/ 20GB VRAM). https://github.com/ggml-org/llama.cpp/pull/26841 https://github.com/ggml-org/llama.cpp/pull/26841
- bwfan123 2mo agoJust tested muse-glimmer:30b-mlx on my laptop. Works great although a bit slow.
- jakswa 2mo agoAnother candidate for the 7900XT (20GB VRAM) I got sitting around. I pulled latest llama.cpp (targeting vulkan during build) after seeing a muse PR merged a few hours ago, and unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL runs on my 7900XT barely (and with no MTP). Sits at 19GB VRAM w/ 4 parallel 113k context slots, all layers on GPU, and at 700 tok/s prompt, and ~36 tok/s generation. Waiting on Q3 to download to check speed + do my usual anecdotes. I generate beefy code snippets and poems, and also ingest my HOA declaration and answer nuanced questions. edit: i should've prefaced this somewhere with: This card ballparks at 800GB/s IO, which I can't seem to find easily on the market anymore. Kinda the ideal card for this model, if I just had a _little_ more VRAM (XTX is 24GB). edit2: not mtp, this is dflash model (param in child comment). I'm up to ~60 tok/s generation and sitting at 19GB VRAM (i added --no-mmproj (makes it text-only i believe) because I'm used to speculative decoding wanting more VRAM and I'm already close to the limit :sweat_smile:)
- jakswa 2mo agoQ3 results: unsloth/Muse-Glimmer-30B-GGUF:UD-Q3_K_XL gets down to 15.6GB VRAM and full context (131k) on the 4 parallel slots. Prompt/generation speeds about the same. Overall feeling like a nicer-fitting Qwen 3.6 27B, but want to test out MTP generation speeds once I can. edit: My favorite bit of reasoning I saw go by in my "generate me a beautiful code snippet" anecdote: 'Could give a snippet of beautiful code: the "hello world" in brainfuck? No.' edit2: my first dflash speculative model! no mtp. I'm up to ~60 tok/s on empty context with `--spec-type draft-dflash`
- wyzer 2mo agoHow are you handling the tradeoff between quantization for device fit and accuracy loss on tool calling? That's where local agents typically break down in production.
- jckahn 2mo agoWhere is the pelican??
- spaqin 2mo agoThat's a bit amusing - not that I have the hardware to run it, but officially it's not available in Hong Kong. Not that getting it would be much of a problem with a help of a VPN either, but I'll assume mainland China is also restricted. Certainly not a competition for Chinese open weight models... in China.
- dhchun1203 2mo ago[flagged]
- GodelNumbering 2mo agohttps://xcancel.com/finkd/status/2086755195535413696 https://xcancel.com/finkd/status/2086755195535413696 "... Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model..." This is bigger news - good for self hosting enthusiasts and a strategically sound move for Meta. Any push towards 'anti Chinese' models will directly benefit Meta as the competition on the frontier open-weights American models is almost non-existent. Meta will have no problem being #1.
- JLO64 2mo agoIt wouldn’t surprise me if Meta does become the #1 American open weights provider, but I doubt it’ll be easy. Thinking Machines has a good amount of talent behind them as I understand it and their Inkling model was decent (admittedly not great though). I think Meta’s biggest problem is going to be internal as there’s be a bunch of headlines posted here on their talent retention issues.
- azinman2 2mo agoWhat about Inkling? It's a quite large model that for some reason isn't discussed much.
- kingo55 2mo agoPoolside Laguna was quite good too (if you look beyond some of the teething issues). Had Deepseek V4 Flash 0731 not launched, their latest Laguna release was really intelligent at non-coding tasks and it would have been my go-to model for my local workloads.
- kevincox 2mo agoFor me Laguna frequently slightly corrupted text then it would be unable to notice the difference and get stuck making the dumbest conclusions. Thinks like typoed directory or function names. It was a great model other than that, but I ended up just going back to Qwen3.6
- alfiedotwtf 2mo ago
- hndhyc0bdt 2mo agoRefreshingly practical
- HardCodedBias 2mo agoLOL the mogging of GDM is hilarious. I don't know why MSL released this, but it is very nice that they did.
- ed 2mo ago[dead]
- mmaunder 2mo agoRemember when we needed 200 servers for an enterprise website because Apache used one process or thread per connection - and Nginx collapsed that into a single box overnight? That moment for LLMs is near. It’s going to move us from the big iron era of AI to small portable brains. Nature has already proved it’s possible with 20 watts and very little heat generation. And I think the data center buildout will end in carnage.
- plutokras 2mo agoWhat specific technical signals make you think we're close to a shift like that?
- mmaunder 2mo agoThe researchers who published Attention Is All You Need didn’t have the benefit of the LLMs they birthed. Take a look at the prompt that solved the Cycle Double Cover conjecture, and which has been adapted to achieve breakthroughs in cybersecurity. The field is entering a feedback loop that is leading to exponential innovation. We’re at the beginning of the curve. And right now the big iron data center approach is brute forcing the problem.
- phkahler 2mo agoI dont think its exponential innovation. Rapid incremental innovation is happening very fast with some occasional bigger bumps.
- nhecker 2mo agoBecause I wasn't familiar with it and others might be curious too: that prompt is available at https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_prompt.pdf https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... and is just below 5 KiB of text.
- therealdrag0 2mo agoDo we know how much token dollars it took to solve?
- deleted 2mo ago[deleted]
- TommyLe999 2mo ago[dead]
- TommyLe999 2mo ago[dead]
- mytailorisrich 2mo agoRandom question: Would you be able to run this model on a Macbook Air M5 (latest)?
- albrewer 2mo agoIt it has less than 64gb then probably not
- mytailorisrich 2mo agoThanks. Ah yes, I skipped over their own figures in the article. K-Quant-17GB seems possible, though as they state 24GB.
- mark_l_watson 2mo agoMeta is rocking AI. As of last week I have been using their excellent muse coding harness with their model Muse Spark 1.2. Starting this morning I am running their new local 30B model muse-glimmer on my old MacMini 32G using Ollama (remember to increase the context size!) and pi coding harness. I am getting good results with muse-glimmer running locally, with the caveat that everything runs slowly (e.g., give it a task and then go walk outside or do Qi Gong exercises for a while).
- spaceywilly 2mo agoNewb question but I’m curious what would help it to run faster? Would it need more vRAM or just system memory?
- spmurrayzzz 2mo agoThe biggest gain you'll get is faster memory, provided you have enough capacity to load all the weight into vram. The DGX sparks and Apple silicon memory bandwidth (and also memory access latency) drag down the decode speed quite a bit. I have two GPU rigs both with 2x RTX Pro 6000, can get ~250 tk/s decode with deepseek-v4-flash in native mixed precision. For context, in antirez's dwarfstar project he only gets ~20-40 tk/s on the same model @ 2bpw on M5 Max. The latter is for sure usable if it's your only option, but it's really hard for me to personally go back to speeds like that when I've experienced the former. (Also worth noting dwarfstar only has experimental support for dspark spec dec, when that lands it will definitely give a big boost at higher acceptance rates)
- wincy 2mo agoIt runs very quickly on my RTX 5090 fwiw. Whole thing is loading entirely into vRAM with a ~130k context size (the max) fitting as well.
- codazoda 2mo agoThat's a $5k 32GB card for anyone who doesn't know all these off the top of their heads (like myself).
- androiddrew 2mo agoI'd really like to see a 45B-ish dense model ready for a dual GPU setup. Something with a little more intelligence while still within the range of some higher end local setups.
- tgtweak 2mo agoThere is definitely an under-served target memory size of 48GB - almost everything aims for: 12, 16, 24, 32, 64, ...) But most dual-gpu setups, 3090/4090 (and some mac configs afaik) have 48GB, and most 64GB systems would do well with the extra 16gb of overhead saved. 48GB is also moderately common in PC memory configurations since 24gb DIMMs are a thing.
- treksis 2mo agothank you zuck.
- brumbelow 2mo agoand now the recent Meta model 'security issue' begins to make sense
- soupspaces 2mo agowhat's the catch?
- soupspaces 2mo agoStay tuned!
- hn97o8vvbt 2mo agoQuietly the best thing in the thread
- golly_ned 2mo agoHaving just bought a 5070 Ti (16GB) instead of a 5090 (24GB), I am sad.
- floturcocantsee 2mo ago5090 has 32GB of RAM.
- mpaepper 2mo agoWhy did you decide for the 5070 Ti? You will always suffer compared to the 5090?
- bhelkey 2mo agoI assume due to price. The 5070 Ti costs ~$1k, the 5090 costs ~$3.5k.
- BoredomIsFun 2mo agoThrow in 5060ti. By the way 5090 is 32 GiB.
- andy99 2mo agoThe gguf is up and works, I don’t know if it’s them or unsloth that’s facilitated this but it’s nice because e.g. Inkling still doesn’t appear to have support in llama.cpp which makes it irrelevant to a class of user. Unfortunately I don’t have enough experience with Qwen 27B to immediately compare, but I do it’s Qwen 3.6 35B A3. It’s much slower obviously but it seems to be way more efficient with its thinking to the point that using it might actually be faster. I find Qwen and some others rehash the same things over and over when thinking without getting anywhere, in mg limited checks here Muse is much better.
- dofm 2mo agoI don't really use the Qwen 3.6 27B though I do test the variants (Bonsai, ThinkingCap). I really like the 3.6 35B A3B for experiments, and it seems OK, but as you say, it spins round in thinking loops more than say the 26B Gemma 4 does. If Muse doesn't actually-wait itself as much it will be very interesting. I am just downloading it to run my small tests.
- jedbrooke 2mo agothe “Actually… But wait!” style responses are so annoying, even Claude opus struggles with this so I’d be interested if meta has done something to cut down on that while still giving good responses
- dofm 2mo agoIt's not perfect but it is very terse! Better than BottleCap managed to do with post-training Qwen in ThinkingCap. I suspect it will help a lot with enabling preserve-reasoning, because the biggest apparent limitation of this model is the 128K context window. Though the practical issue I am seeing on my M1 Max MBP is that performance suddenly drops off a cliff if I have DFlash enabled.
- hadlock 2mo ago128k context window is a complete non-started for us. We need to optimize our most needy agentic jobs, but our average context is well above that
- aand16 2mo ago[dead]
- swrrt 2mo agoJust asking, what is the recommended models for M3 MacBook with 18G memory? Seems modern local models are not available.
- qaz_plm 2mo agoYou can try this site, toggle your computer specs at the top for a refined list of models and tokens/sec. https://www.canirun.ai https://www.canirun.ai
- bwfan123 2mo agoNext step: Burn the weights of these local models into an asic that ships cheap on a laptop (AMD/taalas looking at you), and I will be a happy camper. Make it pluggable so I can select a model I want. I use qwen3.5-coder currently on my laptop, and while it works well enough for me, it is somewhat slow processing tokens. I would hazard a guess that fast small models with a smart agent harness can do quite well compared to large models which cant be run locally.
- spwa4 2mo agoFrom twitter Alexandr Wang > 3/ muse glimmer was developed with its own architecture and recipe, optimized for its size and agentic performance requirements. This means we're in the endgame does it not? If the architecture was NOT optimized for intelligence ...
- kburman 2mo ago[dead]
- m00dy 2mo agowhat I can tell is that Meta is just starting and it is so underrated.
- harisamin 2mo agoLet’s give thanks to all those meta engineers who have been ripped for my heir teams (while sitting right by them) working on manually tagging data. I guess the morale dip paid off in some way? I wish you all well and hope you find some happiness … IYKYK
- Aurornis 2mo agoUnsloth has quantized versions uploaded: https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF The quantized releases often change in the weeks following release as new improvements are discovered, so either use a tool that checks HuggingFace for new versions or manually check back in a few days or weeks to check for improved versions. Initial reports are good. It hasn't been out long enough for anyone to really test thoroughly, but the people I know who have stable non-public test cases are reporting impressive results compared to even Qwen3.6 27B. That's a good sign that this might not be benchmaxxed (trained to excel at public benchmarks with less impressive performance on general tasks) which has been becoming common with recent releases. www.reddit.com/r/localllama is a good place to keep up with the details from people who are actually using it. It feels strange to recommend a subreddit over Hacker News, but on this topic the /r/localllama threads are much more on topic right now if you're looking for information about the model. There are some initial reports that even the 2-bit quantization is looking somewhat usable. That might make it small enough to squeeze into 16GB GPUs. I'd take those reports with a grain of salt because early tests are often optimistic and I've yet to see good results from anything 3-bit or less, but it should be fun to experiment with.
- brrrrrm 2mo agoare these all uniform quantization? or mixed and matched by layer (can't tell from the naming scheme)
- simonw 2mo agoPelican, rendered by Muse Glimmer on my Mac running LM Studio (with this model release: https://lmstudio.ai/models/muse-glimmer https://lmstudio.ai/models/muse-glimmer): https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff20d4cd0ea7596990f7910ead616493e https://tools.simonwillison.net/markdown-svg-renderer#url=ht... It has all of the components of a pelican riding a bicycle, though not exactly arranged in the right order! (For comparison, here are the pelicans I got from Muse Spark 1, 1.1, and 1.2: https://bsky.app/profile/simonwillison.net/post/3mseqv5z4qk2z https://bsky.app/profile/simonwillison.net/post/3mseqv5z4qk2... )
- tarruda 2mo ago> It has all of the components of a pelican riding a bicycle, though not exactly arranged in the right order! Maybe a sign that they didn't have SVG pelicans in the dataset
- BoredomIsFun 2mo agoIt is very bad with any svgs.
- cpfohl 2mo agoPicasso's Pelican
- wasabi359 2mo ago[flagged]
- ThouYS 2mo agoQwen 3.6 27B is still such a beast!
- nirbendavid 2mo agoMany companies are stressed about token cost, as we are moving to a consumption based charge. In the meantime - new open source models, such as DeepSeek V4 Flash and GLM5.2 reduced the price to about 13x chepaer. Also OpenAI had reduced its price for considerably. Now Meta is back in this game. The upcoming months are going to be interesting (GoT)...
- jhgik798 2mo agoHow many data using in Polish Language?
- Kassandraripley 2mo ago[flagged]
- reilly3000 2mo agoPSA: Fast RAM isn't going to be getting cheaper anytime soon. Acquiring inference hardware is a really good way to own an appreciating hard asset. Learning how to use it and cool it is a hacker's journey worth taking. My 4090 I bought in late 2022 for $1600 is selling for a cool $3,489.95 right now, and going strong under nominal use. My DRR5 has tripled in value, my nvmes almost doubled. I grabbed a 128GB M5 Max MacBook Pro when they were still available and told all my friends to buy at least one. With that and a base M4 Studio 36GB, HuggingFace rates that hardware as: > Amazing! You have a total of 128.94 TFLOPS of computing power. 71.3% percentile on scale of "GPU Poor" to "GPU Rich" The way I see it, these are amazing machines that the richest folks are hovering up. I think they should be in the hands of regular people as much as possible. They depend on an incredibly global, increasingly fragile supply chain. If the become impossible to produce, their value would increase tremendously. I think they will become really valuable to you to use the tokens directly, but if that isn't the case, they can be rented out or resold. Please don't just buy any hold. Let's try to get as many people that can use them for decent things that help humans. For example: https://spectrum.ieee.org/small-language-models-ai-pharmaceuticals https://spectrum.ieee.org/small-language-models-ai-pharmaceu...
- hnx0rqy49u 2mo agoClear, useful, done
- brcmthrowaway 2mo agoAny MLX results?
- ddzzz 2mo agohttps://pytorch.org/blog/fast-ondevice-agentic-ai-with-executorch/ https://pytorch.org/blog/fast-ondevice-agentic-ai-with-execu...
- noodleweb 2mo agoHappy to see meta back in the game, it's like after llama nothing came out that was comparable to mainstream open models.
- kyledrake 2mo agoThe post suggests that you need an rtx 5090 use it, which is currently selling for around $5,000 USD. I wouldn't exactly call that "my device", since my device costs about 25% of that for the entire computer. For the same cost, you could run on a frontier model on a pro plan for two years. The economics dont make a lot of sense for this to me, so I would love some input on why people want to do this instead (privacy, for fun, etc).
- mayank 2mo agoIf you’re doing breakeven math on subscriptions, consider that your own rig can run 24/7 whereas you will get a fraction of that with sub rate limits. Even if you factor in PG&E residential rates, the breakeven is a lot closer to months for overnight long-running agentic coding a couple times a week. And in terms of interesting use cases: recently pointed an agent at Blender and gave it vision. That setup can essentially iterate on a scene forever.
- biesnecker 2mo agoIt seems exceedingly unlikely that the current Pro plan costs will hold for the next two years. The subsidization train is going to end eventually.
- spelk 2mo agoIs the assumption here that inference costs will stay roughly static, or that frontier models will keep getting more expensive quickly enough to offset efficiency gains? Because I don’t think “the subsidization train is going to end” necessarily means current pricing becomes impossible. If capital keeps pouring into frontier AI, companies still have an incentive to subsidize access while competing for users and market share. And if that subsidization starts drying up, there’s even more incentive to bring inference costs down by making smaller and cheaper models catch up to today’s frontier capabilities. So either way, I’m not sure you can extrapolate from the cost of serving current frontier models to what equivalent capability will cost two years from now.
- coder543 2mo agoThe post does not imply the 5090 is needed, that is just a common reference point. A single six year old RTX 3090 works great: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_glimmer_actually_fits_on_a_single_rtx_3090/ https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl... I fully expect Meta will release other, smaller Muse models in the near future too. The 5090 is also supposed to be a $2000 GPU, not a $5000 one. The entire market is utterly distorted right now, which will impact cloud inference more and more over time too. They are not immune to the absurdly high RAM prices, so their prices will have to go up over time too until the RAM supply chain goes back to normal.
- nezhar 2mo agoI tried to run it with lemonade by installing it via hf but did not succeed, it gets some weird 500 errors. I also see that ollama has currently only an mlx version available. Anybody here succeed to run this on AMD?
- PuPi 2mo ago[dead]
- koof 2mo agokind of a nonspecific complaint, but i haven’t yet had much luck with anything under ~120b, feels like models released on that order is coming to a trickle. the last few qwen models didn’t seem to go that high, and i got worse results than qwen3.5-122b
- TormentNexusAI 2mo agoThe combo that makes agents reliable: progressive tool routing, persistent memory, and multi-model failover.
- folienumero 2mo agoIn my experience it's faster (10tk/s vs 35tk/s) and better than qwen3.6 series.
- ignitioncar 2mo ago[flagged]
- dofm 2mo agoI haven't really got that far in, but it writes in a sort of clipped, geeky note form in the reasoning traces without too obvious claudeisms, it seems to have been trained to have a level of wit, almost. Like, in the car wash test, this was in the thinking traces: “Walking won't get the car washed.” and: “Perhaps answer: Walk if you want to wash yourself? No” Which made me laugh out loud. Even in the final answer: - - - You have to drive it. Walking 50m won't get the car clean, it'll just get you to the car wash. If you mean you going to the car wash to check prices / pay / get a brush, then yeah, just walk the 50m. It's about 30 seconds on foot and you save the cold-start emissions of firing up the engine for a distance you could roll. If you mean the car itself getting washed, the car needs to be at the car wash. You can push it 50m for a workout, but driving it 50m is the practical way. - - - The emphasis on "you" was from the model. I mean I write like this so I can't judge its tone harshly :-) ETA: The knowledge cutoff is January this year, so it didn't encounter car wash discourse in the scraped training set, though I suppose you can't rule out some kind of fine tuning to deal with this scenario. Still made me chuckle. ETA 2: obviously I wrote this before you added your last paragraph. WTF dude.
- alexeiz 2mo agoI got Muse Glimmer to say "Drive the car, walk yourself." The logic is unbeatable.
- heysagnik 2mo agoeven 30B model is too large to large on local device (low end). meta should provide free hosted model api to use it.
- Schlagbohrer 2mo agoMeanwhile those of us with 128GB RAM plus some VRAM don't have any good modern (last 8 months) open weights models to make use of all that. I don't care if it would run 5 tok/s, I want a smarter model than Qwen3.6 which avoids loops and can handle more context than 80k before crashing.
- heysagnik 2mo agowhy don't you use the quantized version of kimi-k3
- Schlagbohrer 2mo agoI do not see a quantized version of kimi-k3 on huggingface that can fit into 128GB. The smallest Unsloth q version is 594GB https://huggingface.co/unsloth/Kimi-K3-GGUF https://huggingface.co/unsloth/Kimi-K3-GGUF
- Godsend69 2mo ago[dead]
- hypfer 2mo agoHaving played around with this model a bit, I am fairly confident that it is not competing in the coding space. It can do that, but its actual selling point appears to be a different take on guardrails and safety alignment. Either that or the only new training data left was industrial quantities of dark romance literature and Wattpad. Clever business move. 131k context is more than enough for that use case, and due to that small K/V footprint, you can probably have a bunch of characters on the same GPU. Or it's just a happy little accident. We will never know. ___ I was informed that normal people use LLMs for mundane tasks like asking for a pancake recipie. That it apparently can also do decently. Unfortunately, it is also very confident, regardless of whether it is actually correct. So maybe it should actually stay the smut engine and nothing else.
- jawiggins 2mo ago> Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It’s small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation. The next iteration in LLM products is a 24/7 thinking loop where the claude-code like thing gets input continuously from your wearable, notifications, and newsfeeds and is constantly preparing things for you.
- lukaslalinsky 2mo agoThis is is already possible with Claude Code. I use a setup where I have one instance monitoring a local queue, I have a web app for receiving webhooks from various sources and pushing them to the queue. Plus email for things that don't have webhooks. That instance then decides what to do with each input, sometimes it can spawn additional agent to investigate/prepare, sometimes it creates a ticket assigned to me and then waits for me input. All of that just uses the monitoring tools built into CC. The dispatcher loop doesn't need to be extremely smart, so I might experiment replacing it with a local model like this.
- Computer0 2mo agoIf you are willing and not too busy, What model do you use and what is your cost? (If using subscription would you be able to check with 'npx ccusage').
- arkmm 2mo agoI'd also be really curious about the cost to run something like this, and what things you think it's particularly helpful for?
- lukaslalinsky 2mo agoI run this on a side of the Claude Pro subscription that I use for other purposes. My main motivation was root cause analysis of production issues. I have a solo project and unfortunately my mental state has been degrading over the last years. I would avoid looking at production issues, because I didn't have the energy to focus on the investigation. So I automated this, setup the loop, setup metrics/logs access for Claude to use and now whenever something goes bad, I have a single report that I can act on easily, and if I don't, it will ping me in a way that's not spammy like automated alerts. But I'm finding more uses for it.
- wxw 2mo agoMeta's clearly changing strategies back towards their original "frontier open source", but this time around they have a lot more competition from leading Chinese labs. I'm all for it though, and I think Glimmer is a fantastic bet on locally-hostable models. I for one would love to self-host as much as I can.
- catoc 2mo agoPersonally I would never trust a coding agent or agent harness from Meta. I agree with their open-source model approach, but actually trusting Meta… to protect my privacy and my data… when it’s running on my personal hardware… Not . In . A . Million . Years - that ship has sailed
- ionwake 2mo agoSorry I dont know if this is the right place but... 2000AD The Glimmer Rats , was the best drawn comic strip story by far in that publication.
- hougaard 2mo agoTried it (the full version, using 120 GB RAM), wasn't impressed, gave it some defective code, and asked it to fix all errors. It kept looping around and around and digging itself deeper and deeper into a rabbit hole; eventually, it got into a "reasoning" discussion about whether a custom compiler was used that supported the wrong syntax...
- aussieguy1234 2mo agoThe SWE bench verified score is similar to Opus from not so long ago. Sure, you can get better performance from cloud models. But most software, not just AI, will be faster and more reliable in the cloud. The question is do we need that additional power and cost. If the answer is no, then just like other software, people will run AI locally.
- jmspyderbot 2mo ago[dead]
- TokenLat 2mo ago[flagged]
- realaaa 2mo agoand immediately followed up with Manifesto from the man himself - what / how are they going to make of it longer term? I guess for FOSS and self hosted it is good - but I am still wondering how are they going to Meta-stasize it ;)
- myshapeprotocol 2mo ago[flagged]
- shubhamsinghani 2mo ago[flagged]
- coryadams 2mo agoGood candidate to test breaking down subtasks in agent flows to cut token spend. I'll be interested to see how it performs with a LORA driven adaptor model stitched on to drive agentic security remediation tasks.
- soldhtml 2mo agowould be fun to see side- by-side
- omidggmh 2mo ago[dead]
- haobing0304 2mo ago[flagged]