4 ms·
Would love to see a Qwen 3.5 release in the range of 80-110B which would be perfect for 128GB devices. While Qwen3-Next is 80b, it unfortunately doesn't have a
by tarruda 8mo ago
Would love to see a Qwen 3.5 release in the range of 80-110B which would be perfect for 128GB devices. While Qwen3-Next is 80b, it unfortunately doesn't have a vision encoder.
- bytesandbits 8mo agomaybe a deepseek v4 distill. give it a few days
- PlatoIsADisease 8mo agoWhy 128GB? At 80B, you could do 2 A6000s. What device is 128gb?
- vladovskiy 8mo agoGuess, it is mac m series
- the_pwner224 8mo agoAMD Strix Halo / Ryzen AI Max+ (in the Asus Flow Z13 13 inch "gaming" tablet as well as the Framework Desktop) has 128 GB of shared APU memory.
- hedgehog 8mo agoKeep in mind most of the Strix Halo machines are limited to 10Gbe networking at best.
- paulsmal 8mo agoyou can use separate network adapter with RoCEv2/RDMA support like Intel E810
- hedgehog 8mo agoMost Ryzen 395 machines don't have a PCI-e slot for that so you're looking at an extension from an m.2 slot or Thunderbolt (not sure how well that will work, possibly ok at 10Gb). Minisforum has a couple newly announced products, and I think the Framework desktop's motherboard can do it if you put it in a different case, that's about it. Hopefully the next generation has Gen5 PCIe and a few more lanes.
- scoopdewoop 8mo agoNot quite. They have 128GB of ram that can be allocated in the BIOS, up to 96GB to the GPU.
- khimaros 8mo agoallocation is irrelevant. as an owner of one of these you can absolutely use the full 128GB (minus OS overhead) for inference workloads
- EasyMark 8mo agoCare to go into a bit more on machine specs? I am interested in picking up a rig to do some LLM stuff and not sure where to get started. I also just need a new machine, mine is 8y-o (with some gaming gpu upgrades) at this point and It's That Time Again. No biggie tho, just curious what a good modern machine might look like.
- breisa 8mo agoThose Ryzen AI Max+ 395 systems are all more or less the same. For inference you want the one with 128GB soldered RAM. There are ones from Framework, Gmktec, Minisforum etc. Gmktec used to be the cheapest but with the rising RAM prices its Framework noe i think. You cant really upgrade/configure them. For benchmarks look into r/localllama - there are plenty.
- aruggirello 8mo agoMinisforum, Gmktec also have Ryzen AI HX 370 mini PCs with 128Gb (2x64Gb) max LPDDR5. It's dirt cheap, you can get one barebone with ~€750 on Amazon (the 395 similarly retails for ~€1k)... It should be fully supported in Ubuntu 25.04 or 25.10 with ROCm for iGPU inference (NPU isn't available ATM AFAIK), which is what I'd use it for. But I just don't know how the HX 370 compares to eg. the 395, iGPU-wise. I was thinking of getting one to run Lemonade, Qwen3-coder-next FP8, BTW... but I don't know how much RAM should I equip it with - shouldn't 96Gb be enough? Suggestions welcome!
- lm28469 8mo agoThat's the maximum you can get for $3k-$4k with ryzen max+ 395 and apple studio Ms. They're cheaper than dedicated GPUs by far.
- tarruda 8mo agoMac Studios or Strix Halo. GPT-OSS 120b, Qwen3-Next, Step 3.5-Flash all work great on a M1 Ultra.
- sowbug 8mo agoAll the GB10-based devices -- DGX Spark, Dell Pro Max, etc.
- tgtweak 8mo agoSpark DGX and any A10 devices, strix halo with max memory config, several mac mini/mac studio configs, HP ZBook Ultra G1a, most servers If you're targeting end user devices then a more reasonable target is 20GB VRAM since there are quite a lot of gpu/ram/APU combinations in that range. (orders of magnitude more than 128GB).
- kristianp 8mo agoBy A6000, do you mean the older Ampere generation model? 48 GB ddr6, released 2020 [1]. Can you even buy those new still? [1] https://www.techpowerup.com/gpu-specs/rtx-a6000.c3686 https://www.techpowerup.com/gpu-specs/rtx-a6000.c3686
- Tepix 8mo agoHave you thought about getting a second 128GB device? Open weights models are rapidly increasing in size, unfortunately.
- tarruda 8mo agoConsidered getting a 512G mac studio, but I don't like Apple devices due to the closed software stack. I would never have gotten this Mac Studio if Strix Halo existed mid 2024. For now I will just wait for AMD or Intel to release a x86 platform with 256G of unified memory, which would allow me to run larger models and stick to Linux as the inference platform.
- kylehotchkiss 8mo agoI aspire to casually ponder whether I need a $9,500 computer to run the latest Qwen model
- amelius 8mo agoYou'll need more since RAM prices are up thanks to AI.
- 3abiton 8mo agoGiven the shortage of wafers, the wait might be long. I am however working on a bridging solution. Sime already showed Strix Halo clustering, I am working on something similar but with some pp boost. Unfortunately, AMD dumped a great device with unfinished software stack, and the community is rolling with it, compared to the DGX Spark, which I think is more cluster friendly.