Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
AlekseiSavin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
AlekseiSavin
1y ago
yeah, tokens per second
2.
▲
by
AlekseiSavin
1y ago
it looks like llama.cpp has some performance issues with bf16
3.
▲
by
AlekseiSavin
1y ago
right now, we support AWQ but are currently working on various quantization methods in https://github.com/trymirai/lalamo
4.
▲
by
AlekseiSavin
1y ago
You're right, modern edge devices are powerful enough to run small models, so the real bottleneck for a forward pass is usually memory bandwidth, which defines the upper theoretical limit for inference speed. Right now, we've figu
5.
▲
by
AlekseiSavin
1y ago
already) https://github.com/trymirai/uzu-swift