5 ms·
Our investigations indicate that it might not be possible to achieve ANE performance improvement over CPU for LLM Decoder inference with batch size of 1 [0]. Ju
by ggerganov 4y ago
Our investigations indicate that it might not be possible to achieve ANE performance improvement over CPU for LLM Decoder inference with batch size of 1 [0]. Just to make it clear - I'm no expert in Core ML / ANE, so these conclusions could be totally wrong.
[0] https://github.com/ggerganov/whisper.cpp/discussions/548#discussioncomment-5265849 https://github.com/ggerganov/whisper.cpp/discussions/548#dis...
- fwlr 4y agoDon’t sell yourself short! (And you have my apologies in advance if my excited comment above has created any extra work for you)
- DennisAleynikov 4y agoNeural Engine across the M1 and M2 series is also sadly very limited. I bought one thinking I could exploit it for StableDiffusion and other tasks but found that most libraries say to use GPU for faster generation. What I found is not only is the engine the same on m2 pro (meaning I upgraded for no reason from my m1 basemodel) but it also doesn't scale at all except in the m1 Ultra where it's doubled simply because it's using two dies bridged. Neural Engine can generate 512x512 images pretty easily but takes a while even compared to using the GPU on a basemodel m1 Mac Mini. It's kinda crazy. Looking into ways to improve it and take advantage of the neural engine in the future but the current situation is very limited. Even apples official implementation and coreML libraries seem to prefer you run them on Metal