3 ms·
There was a 3.5 122B 10A release - https://huggingface.co/Qwen/Qwen3.5-122B-A10B https://huggingface.co/Qwen/Qwen3.5-122B-A10B
by mattnewton 2mo ago
There was a 3.5 122B 10A release -
https://huggingface.co/Qwen/Qwen3.5-122B-A10B https://huggingface.co/Qwen/Qwen3.5-122B-A10B
- kanemcgrath 2mo agoI tried it for a bit, and It was not really worth its size. It got swept up in all the other AI news recently, but laguna s 2.1 I think is the best ~100B moe model right now
- nozzlegear 2mo agoI didn't mention it above, but Laguna S is my other favorite model. I use Qwen a lot more, it's smaller and faster, but I like to switch to Laguna when I feel like I need a "heavy hitter" for certain huge or complex tasks.
- tommica 2mo agoWhat on earth hardwares do you guys have to be able to run 100gb models locally?! That's crazy! I'm here struggling to even get 27b models to run in somewhat usable way
- nozzlegear 2mo agoHaha I'm on an Mac Studio with an M1 Ultra, 64gb ram. I bought it when it first came out, it just happens to be good for local LLMs. I have to use a smaller quant of Laguna S though (I think 4-bit? Not at my machine to check), as 8-bit and full size definitely don't fit in the 64gb I have.
- lcnPylGDnU4H9OF 2mo agoYeah, a good rule of thumb is that the weights take up ~100% of the size of the model, so 100B bytes (8-bit quant) would be, well, 100GB and a 4-bit quant would be half that.
- tommica 2mo agoOh, that is a useful rule to know! Thanks!
- cyanydeez 2mo agoThat's not exactly the math. Theres also vram needed for context. I operate several 72-128 GB machines and the larger the context the slower they go. And the context takes space +kv cache. KV cache drives usefulness as your context grows, it needs to pull the kv cache. Simplified, the context has to be run on every turn, so the KV cache supplies the processed tokens, so it just needs the new inpute.
- CyberDildonics 2mo agoI don't know if that's, well, a rule of thumb, it might be, well, straight multiplication.
- lcnPylGDnU4H9OF 2mo agoWell, yeah, the straight multiplication, er, well, "follows" from the rule of thumb. Hope that helps!
- CyberDildonics 2mo agoSaying, well a rule of thumb is, well, 100 billion bytes is a 100 gigabytes, is well, not a rule of thumb. It is, well, just the common definition.
- lcnPylGDnU4H9OF 2mo agoRight. The rule of thumb is that the overhead size of the model that's not the weights is so vastly outweighed by the actual number of weights that it can be disregarded. My shorthand for that was to write "the weights take up ~100% of the size of the model". What then "follows", both in the sense that the explanation is written after the rule as well as that it logically follows, is that, well, 100 billion bytes is, you know, 100 GB. I don't see why you're, like, paying so much attention to this?
- tommica 2mo agobrb, going to see if 2nd hand mac studios are available!
- colordrops 2mo agoMoE models can use system memory along with a GPU.
- formvoltron 2mo agoand get high token bandwidth?
- colordrops 2mo agoSimilar to a spark, which isn't blazing fast but usable.
- numpad0 2mo agoNot badly so because MoE models(identifiable by "CoolName-xxxB-AxxB" naming scheme) have bunch of branches in the middle that only one out of all gets non-zero values. Each of branches aka "Experts" as well as top/bottom parts are significantly smaller than the whole, and so CPU emulation of CUDA operations mixed with GPU taking as much as possible become not so out of question, unlike for dense models("CoolName-xxxB" without "-AxxB")
- sarjann 2mo agodgx spark, nvfp4 so I have spare room for KV cache (context)
- sznio 2mo agoquantized + offload I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s
- TonyStr 2mo agoIQ4 qwen 122b-a10b would mean 61GB total size and 5GB active, so about 5GB of the model loaded into GPURAM plus any generated context, and 61GB of weights loaded into system RAM? I don't know if that math is correct, but does that run well? Wouldn't that only leave 3GB of system RAM?
- numpad0 2mo agoYeah, here I am sitting deeply deeply deeply regretting not buying couple CMP 170HX at $200 or $350, knowing I could just flip them ethically at purchase price if nothing came of it... I could have just casually built a 128GB dual A100 local AI monster
- jimmySixDOF 2mo agoI'm working with a lab that has a few Ampere GPUs on infiniband and they are just not compatible with the latest quants and vLLM updates. FP8 is about as low as you can go.
- numpad0 2mo agoBut they're reportedly a soft nerfed GA100 64GB/40GB at $1200, that's not more expensive and certainly can't be slower than a Mac Studio.
- cyanydeez 2mo agoUsually theyre quantized. Also, there was a window where AMD 395+ W/128GB was just a high end $2500 hardware with unified gpu memory.
- pettijohn 2mo agoStrix Halo, 128GB RAM. I got a refurbished Corsair AI Workstation for a smoking price ($2100) about two months ago. Lucky timing that it was in stock.
- AbsurdCensor 2mo agoStrix Halo as well. Bought it for $1,800 new on sale and shoved an extra 4tb drive into it. Been amazing for local AI. Maybe not the absolute fastest thing (usually around 30t/s depending on the task) but has been awesome for a local AI box that I can solar power.
- SummSolutions 2mo agoNice. Mind sharing the solar side of your setup?
- AbsurdCensor 2mo agoCouple of rack mount batteries and roughly 5kw of solar panels. Feeds into a subpanel so I can flip it when I want a couple rooms of solar on the house, or hook a generator up if needed. Can't power the entire house, but works well for thinks like computers, lighting, etc. And if I want to expand, just throw on more panels, or realistically, just throw on more batteries to store the juice.
- SummSolutions 2mo agoThanks - that’s what I want to do!
- ndom91 2mo agoUsing qwen 3.6 27b for local coding as well and downloaded Laguna s 2.1 but haven't had time to give it a full spin yet. Curious for any more experiences
- mattnewton 2mo agoI agree. 27b dense really did seem like the sweet spot.