4 ms·
An off-topic question, are Apple's M-series chips any good at current AI/ML work? How does it compare with dedicated GPUs?
by BossingAround 2y ago
An off-topic question, are Apple's M-series chips any good at current AI/ML work? How does it compare with dedicated GPUs?
- theodric 2y agoMid-tier gaming GPU performance, but (potentially) access to gobs of memory (if running on a host with gobs of memory) owing to the unified memory design. For certain use cases which require loading huge datasets but don't necessarily require massive compute (i.e. inference on large models) they can be cost-competitive relative to something like an H100.
- yieldcrv 2y agoIf apple offers these in a data center that’s open access it’s game over for NVDA
- Keyframe 2y agoBased on which fantasy premise?
- yieldcrv 2y agoThe premise where these are readily available for mass purchase, have a hardware and software stack that already works reliably, and have a lower energy footprint than other offerings and somewhat competitive on cost, but that wont be the main selling point, just availability
- hu3 2y agoThey aren't even on the same league, computing-wise.
- coldtea 2y agoFor the purposes we're discussing, they're nowhere near competitive with Nvidia.
- Keyframe 2y agolower energy but at least they're slow? You need to add up to the same performance level and then consider the cost and running cost. I bet it's not close on that. Availability maybe, but as you've noted - zero availability for data center environments. Those volumes would also then fall onto TSMC/Samsung/Whatever where Nvidia is stuck as well.
- yieldcrv 2y agoI’m aware Apple is also beholden to TSMC’s capacity too sadly
- saberience 2y agoI have an M3 chip in my laptop, it has more memory than my 4090 but it's still way slower when inferencing. So as long as the model fits in memory, Nvidia GPUs are going to be way faster just because they have more/faster compute cores. Of course, if the model fits in memory on your M chip and doesn't in your Nvidia chip, the M chip wins by default. However, I would say, if you load a 70B model in your M chip, while it WILL work, the tokens/sec will be slow as fuck... so it kinda doesn't matter anyway.
- throw_nbvc1234 2y agoHow does the inference for LLMs impact battery life? For SD it can be 5% battery per image at times.
- NBJack 2y agoThe latest Nvidia drivers offer an option to start using system memory when VRAM is insufficient. It certainly slows things down, but it does work. It's not perfect in my experience, but it is an alternative for large models.
- wiradikusuma 2y agoSounds great! Do you have source maybe a wiki?
- NBJack 2y agoThis might be the best place to start: https://nvidia.custhelp.com/app/answers/detail/a_id/5490/~/system-memory-fallback-for-stable-diffusion https://nvidia.custhelp.com/app/answers/detail/a_id/5490/~/s...
- mrweasel 2y agoI've only tried it on my M1, running Llama-3 via Ollama. It works, but it's slow to the point where it's not really usable. Maybe there are smaller models you can run that will perform better.
- gkfasdfasdf 2y agoWhat size model did you try and how much memory does your M1 have? See my other comment, my experience has been that llama3 was very fast on an M1.
- mrweasel 2y agoI was just running: ollama run llama3, So that would be 8B parameters, on a 8GB M1 Air. Maybe it's just my expectations, but it seems rather slow to process queries. Depending on the prompt somewhere between the 10 - 40 tokens per second, but that very much depends on the prompt. My complaint is the time between the prompt is entered and output starts.
- gkfasdfasdf 2y agoCompared to a GPU like a 4090 with equal vram it probably won't fare well as others point out, however it far outperforms any CPU. On an M1 Ultra MacBook Pro I was seeing like 40 tokens/second with llama3:7B vs 9 tokens/second on various Intel servers/desktops with sufficient ram.