2 ms·
> A model you can run on a loptop is simply not going to work as well as it's needed for programming The models you can run on a high-spec laptop today are app
by OtherShrezzing 2mo ago
> A model you can run on a loptop is simply not going to work as well as it's needed for programming
The models you can run on a high-spec laptop today are approximately where frontier models were 12-18mo ago (albeit at a lower tok/s rate). If you scan back through hn comments from that era, you’ll find plenty of people saying “this is powerful enough to massively increase my productivity”.
- anon373839 2mo ago> albeit at a lower tok/s rate Not always! I get 80-100 tok/s from Qwen 3.6 35B-A3B on a MacBook Pro thanks to MTP. With long contexts that dips to around 50-60. However, prefill is much slower than API models. So it becomes really, really, really critical to not have cache misses.
- drivebyhooting 2mo agoSlower is meaningfully dumber when you’re time bounded and need all the inference time compute you can get.