3 ms·
The article seems to be based on misinterpreted information: > Apple’s model has extremely low latency (0.6 milliseconds to first token), outperforms similar s
by red2awn 2y ago
The article seems to be based on misinterpreted information:
> Apple’s model has extremely low latency (0.6 milliseconds to first token), outperforms similar sized Phi and Gemini models from Microsoft and Google.
From [1] it is 0.6 millisecond per prompt token so unless the prompt is one token the latency would be higher.
> Apple’s server-side models running in the Private Cloud Compute are apparently quite near the GPT-4o in terms of quality
The benchmarks from [1] never mentioned GPT-4o, the best model they compared to for the server model is GPT-4-0125. If their model almost matches GPT-4o they wouldn't need to integrate with ChatGPT.
[1]: https://machinelearning.apple.com/research/introducing-apple-foundation-models https://machinelearning.apple.com/research/introducing-apple...
- deleted 2y ago[deleted]
- elicksaur 2y agoIf you stream the answer, the first token time is roughly the per token time.
- lostmsu 2y agoNo, you have to feed the entire prompt token-by-token before getting response.
- elicksaur 2y agoOh, I see, I misread. Thought it meant 0.6ms per output token. Now I get that it’s saying “prompt token”, so if your prompt is 100 tokens, that’s 60ms. That seems pretty fast. 1.6k tokens for a 1s time. Do other models compare to that? I’m not sure what the current top ranking for this metric looks like.
- lostmsu 2y agoLatency here is weird. You are using the number as bandwidth in this calculation. Perhaps reporter doesn't really know what he's talking about.