4 ms·
Isn't there something to be said for owning your own hardware though?
by sanderjd 2mo ago
Isn't there something to be said for owning your own hardware though?
- bigyabai 2mo agoNot if it's 5-10x slower than a remote inference server. Mac prefill latency is exhausting.
- sanderjd 1mo agoOh tell me more about prefill latency.
- jamiek88 1mo agoM5 changed that a lot though - it could still be better but 4x improvement made it cross the frustratingly slow barrier for me.
- spwa4 1mo agoCurrently m5 max has a prefill rate for Qwen 3.8 27B of 400+ tok/s, then generate at ~60 tok/s. (And it can be improved further, the software is not yet at the level it is for CUDA) That means for context under ~5k or so ttft (time to first token) it's going to respond faster than Claude. If the answer is less than ~1k I think the request finishes sooner. And it's ~claude 4.5 or 4.6 level intelligence. I've had it work for more than a day on a pi.dev "loop engineering" project involving writing software. Plus privacy. Plus offline.