Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
msbhogavi
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
msbhogavi
6mo ago
The hardware situation is way better than you think, and quantization is a huge part of why. Take Qwen 3.5 27B, which is a solid coding model. At FP16 it needs 54GB of VRAM. Nobody's running that on consumer hardware. At Q4_K_M quantiz
2.
▲
by
msbhogavi
6mo ago
"As much memory as possible" is right for model capacity but misses bandwidth. Apple Silicon has distinct tiers: M4 Pro at 273 GB/s, M4 Max at 546 GB/s, M4 Ultra at 819 GB/s. Bandwidth determines tok/s once the