Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gitpusher42
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
61.
▲
by
gitpusher42
2mo ago
The full route changes almost every token. The cache works through partial reuse, about 40% of experts repeat on the next token and 57% within two tokens, cutting I/O from 166 to 88 ms/token on M2 Mac. The longest exact repeat we
62.
▲
by
gitpusher42
2mo ago
AFAIK it should not because it is only reading
63.
▲
by
gitpusher42
2mo ago
The process stays at around 2 GB with 16 slots and a 4K context on both the M5 and M2. But yeah, Apple might be doing some magic under the hood
64.
▲
by
gitpusher42
2mo ago
Uh, I’m afraid it is Apple only. It is written using Apple’s GPU language, Metal, and heavily relies on the Apples’s shared memory architecture Windows PCs would require a completely different approach
65.
▲
by
gitpusher42
2mo ago
Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots
66.
▲
by
gitpusher42
2mo ago
I haven't tried it but it should work! You can try it and share your results, it would be really appreciated I tried it on my wife's M1 MacBook Air 512GB and it gets 4–5 tok/s Also, it must be easy to adjust for iPhones and i
67.
▲
by
gitpusher42
2mo ago
It depends on the use case. I measured this exact model with a 4k context on the mlx engine. It runs at 75 tok/s on my M5 Mac Pro and using 14 GB of RAM. For my engine the same model uses 2 GB of RAM and produces 31–35 tok/s. The
68.
▲
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
(github.com)
919 points
by
gitpusher42
2mo ago
|
344 comments