3 ms·
I wonder when/if we will start to see Apple include a silicon encoded model into their chips. Similar to Taalas build Llama 3.1 silicon with 17,000 tok/s infere
by Humphrey 1mo ago
I wonder when/if we will start to see Apple include a silicon encoded model into their chips. Similar to Taalas build Llama 3.1 silicon with 17,000 tok/s inference.
So could the M7 actually include an AFM 3B model, alongside a generic neural engine?
- wmf 1mo agoThat model would be larger and more expensive than the entire M7 chip.
- tyre 1mo agoWhy would a local model for a consumer device need 17k tok/s? Apple is better off building chips with generalizable TPUs (or equivalent) so they can upgrade/patch models.
- MaxikCZ 1mo agoI cant shake the feeling of "640KB is enough for everybody". Imagine not one AI answering over 1 minute but a team of 100+ agents in hieararchical structure taking care of your request in seconds, checking each other.
- jeffybefffy519 1mo agoIt would open heaps of use cases, you could almost pass it over frames of images the camera sees in real time for example...
- deleted 1mo ago[deleted]
- dyauspitr 1mo agoIt would make zero sense. These things are improving by leaps and bounds every week, we are not at the point where you can burn weights into silicon and put it on one of the largest consumer devices on the planet yet.
- aqfamnzc 1mo agoOn the other hand, models these days are getting to the point where even if all development halted permanently, they would continue to be useful long into the future. (At least until their knowledge base or linguistics become too outdated.)
- deleted 1mo ago[deleted]