4 ms·
Can you expand on this? What is the current co tender for LLM inference that is ~on par with Apples offerings?
by cced 3y ago
Can you expand on this? What is the current co tender for LLM inference that is ~on par with Apples offerings?
- smoldesu 3y agoIt entirely depends on workload. This blog has a diligent mixed-model comparison of the M3 laptops lined up against the Titan RTX (2019) and Tesla V100 (2017): https://www.mrdbourke.com/apple-m3-machine-learning-test/ https://www.mrdbourke.com/apple-m3-machine-learning-test/
- rbanffy 3y agoWould be interesting to see results for the beefier Macs with more RAM. Curious as to what one could do with larger AI models. I'm aiming at 64GB or more for my next Mac, but not because of AI but to ensure longevity. Macs in this house tend to have very long and productive lives.
- seec 3y agoExpect the number of cases where it would give a real advantage would be pretty small. In all the cases presented, Apple gets completely destroyed even with their best chip at around 10x worse; that is before even talking about the cost, after which you can only conclude that Apple hardware makes absolutely no sense for this type of task. Even if you found enough tasks where the larger memory available would make a difference to average mixed performance closer to a NVidia solution, they would seriously need to lower their RAM price for this argument to hold up... With their current commercial offering, Apple isn't competitive for anything, even though the computers and chips themselves are not bad at all.
- treprinum 3y agoM3 Max 128GB RAM inference on LLaMA2 70B is much faster than on Zephyrus G14 64GB with 4090/16GB because one can fit the whole model into the RAM and share it with GPU. For smaller models that can fit into 16GB G14 smokes M3 Max.
- cyanydeez 3y agohttps://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inferen... should give you consumer comparison
- rbanffy 3y agoApple was, IIRC, the first company to make the fact they could have 192GB AI models fully accessible to the GPU a selling point. I remember an AMD chip that could have up to half the memory accessible to the GPU (but not shared with the CPU, IIRC), and read something similar about Intel ones, but no explicit references to sharing the full RAM pool with the GPU at full bus speed.
- smoldesu 3y ago> the first company to make the fact [...] a selling point. I don't think that's as impressive of an accomplishment as you think it is...