4 ms·
$4800?
by jononomo 2y ago
$4800?
- andreygrehov 2y ago$4,599.00 + tax = $4,920 The price is insane, but I needed a solid dev machine, so that I could run multiple IDEs (GoLand, IntelliJ) + Figma. Additionally, I wanted to dive deeper into AI/ML. Apple's MLX easily competes with Nvidia RTX 4090 [1]. Everything is just instant, and I'm extremely happy with the purchase. [1] https://appleinsider.com/articles/23/12/13/apple-silicon-m3-pro-blows-away-nvidia-rtx-4090-gpu-in-ai-benchmark https://appleinsider.com/articles/23/12/13/apple-silicon-m3-...
- ein0p 2y agoNah man. A 4090 is way faster, especially per dollar spent. I have nothing against Apple and own an M3 Max MBP as well, but lets be real here.
- enceladus06 2y agoFor today’s ai-accelerated workloads, will you use a 14th gen intel + 4090 or Apple M3? The software tooling is not available for Apple, there is no CUDA support basically. This is not an emotional judgment, if an Apple supported cuda on Apple Silicon I’d use it, but it doesn’t.
- ein0p 2y agoI use both. Apple’s higher end hardware works well for local LLM inference thanks to its unusually high memory bandwidth and frugal energy consumption. But for serious work you need CUDA, no ifs or buts.
- andreygrehov 2y agoMy apologies. The link I provided is not reliable by any means. I actually knew about this, but somehow forgot. Apple can compete with 4090 only for a certain types of tasks (LLMs?), but overall you are correct, 4090 is way faster.
- makeitdouble 2y agoI bit the bullet and went to windows with a discrete GPU. The power envelope is a subpar , but raw performance is definitely there. With a mobile class 4070 my benchmarks were above the Mr max's and as it's nvidia, tool and framework compatibility are no issue. And price was less than 3K, extended support included.
- paulmd 2y agohow do you handle models that are bigger than 16GB on that? my m1 max 64gb whips through dolphin-llama3:70b just fine, that's around 42gb usage. I get partial offload on dolphin-mixtral:8x7b, that's significantly slower but still reasonably usable, whereas even the quantized models chug on my work m1 max 32gb. everything i've seen is that once you seriously cross the threshold of VRAM usage/gpu offload, the performance collapses. and sure, if you stay within that the 4090 (or even 4070 mobile) is a lot faster, but the most popular current families of models don't generally fit on a 4070 or even 4090. on a 41.5gb model, would a 4070 mobile/4060 ti 16GB really beat a 64gb m1 max? it seems doubtful but idk. if there's viable ways to split models across multi-gpus i'm all ears, but afaik there's not (outside NVLink). Other than buying or renting A100s, the only serious option would be bergamo with avx-512 and a bunch of memory, much slower but avx-512 does have inference instructions now, and that gets you into the class of memory you (currently) need for things like snowflake-arctic-instruct:128x3B. (and of course training is a whole different kettle of fish, but everyone knows that at this point.)
- makeitdouble 2y ago> how do you handle models that are bigger than 16GB on that I don't, my setup won't compare in term of memory affected to the GPU part. I actually don't know what the best workaround would be. My priorities were portability, versatility and VR capacities above all, so yes it wont hold 40G in GPU RAM.