3 ms·
Well, Metal can only allocate a smaller portion of “VRAM” to the GPU — about 70% or so, see; https://developer.apple.com/videos/play/tech-talks/10580 https://de
by sunpazed 3y ago
Well, Metal can only allocate a smaller portion of “VRAM” to the GPU — about 70% or so, see; https://developer.apple.com/videos/play/tech-talks/10580 https://developer.apple.com/videos/play/tech-talks/10580
If you want to run larger models, then CPU inference is your only choice.
- __loam 3y agoAren't these things supposed to have cores dedicated to ml?
- brucethemoose2 3y agoThey are not as fast as the GPU (but much lower power). Also, not many implementations can even use it.
- azinman2 3y agoYou’re thinking of the neural engine. I’m not sure that llama.cpp makes use of this. They’d have to turn it into a CoreML model to do so.