4 ms·
> With this set of optimizations, on iPhone 15 Pro we are able to reach time-to-first-token latency of about 0.6 millisecond per prompt token, and a generation
by revscat 2y ago
> With this set of optimizations, on iPhone 15 Pro we are able to reach time-to-first-token latency of about 0.6 millisecond per prompt token, and a generation rate of 30 tokens per second. Notably, this performance is attained before employing token speculation techniques, from which we see further enhancement on the token generation rate.
This seems impressive. Is it, really? I don’t know enough about the subject to judge.
- bastawhiz 2y agoFor a phone running locally, that's pretty fast. The bigger question is how good the output is. Fast garbage isn't useful, so we'll have to wait to see what it actually ends up looking like outside of demos.