Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stoatmagoats
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
stoatmagoats
3y ago
Friendly internet stranger’s input: - you don’t get GPU acceleration just by using unified memory. Llama.cpp still only uses the CPU on Apple Silicon chips. - the difference in tokens/sec is likely attributable to memory bandwidth. Mac