3 ms·
It is in a weird middle ground. It is much worse ram and much slower hardware (factor 4 or so IIRC) and only an option if you need more than 24 but less than 20
by diffeomorphism 2y ago
It is in a weird middle ground. It is much worse ram and much slower hardware (factor 4 or so IIRC) and only an option if you need more than 24 but less than 200. Also, only if you think that 9000€ is pocket change but 20000€ is cost prohibitive. Also, you need all that power but no server, no ECC,... And you only ever want to do inference but no training or tuning.
If you want a Mac anyway, sure. But if you don't care, this seems like a very, very specific Venn diagram.
- int_19h 2y agoIt's not much worse RAM, though. RTX 4090 has memory bandwidth of 1050 Gb/s. M2 Ultra is 800 Gb/s. And you can get a Mac Studio with Ultra and 128Gb of RAM for $3K or less. It's great for 70-150B models. You're correct that it's only good for inference, but most people running local LLMs only do inference.
- keybits 2y agoThat config Mac Studio costs 5,800 Euros (minimum). Where do you get it for 3,000 USD?
- int_19h 2y agoYou don't need an M2 for this, an M1 will do just fine. Nor do you need the one with maxed-out SSD, which jacks up the price considerably. Finding a brand new M1 Ultra for $4K or less right now is pretty easy. When I got mine, about a year ago, $3K was the best deal I could find.