3 ms·
I can’t imagine apple has chip production scaled to the point it could support all the inference queries coming from all the iPhones
by ninjha01 2y ago
I can’t imagine apple has chip production scaled to the point it could support all the inference queries coming from all the iPhones
- c1sc0 2y agoThey don’t need to if they can process a significant chunk of the queries on-device. Llama3-level inference works fine on M2-level chips today and the M4 is already in a mobile device.