4 ms·
It depends on what you are looking at. Time to 1st token is faster on the M5 because of HW accelerators helping the prompt interpretation (and it is CPU-bound)
by harrouet 2mo ago
It depends on what you are looking at.
Time to 1st token is faster on the M5 because of HW accelerators helping the prompt interpretation (and it is CPU-bound).
Token generation after that is GPU-bound and will profit from the higher bandwidth of the M4 Max.