3 ms·
You're right there is no way to specifically target the neural engine. You have to use it via CoreML which abstracts away the execution. If you use Metal / GPU
by babl-yc 1y ago
You're right there is no way to specifically target the neural engine. You have to use it via CoreML which abstracts away the execution.
If you use Metal / GPU compute shaders it's going to run exclusively on GPU. Some inference libraries like TensorFlow/LiteRT with backend = .gpu use this.
- scosman 1y agoExactly. And most folks are using a framework like llama.cpp which does control where it’s run.