Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ashishdhiman23
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
2 ms
·
1.
▲
by
ashishdhiman23
7mo ago
Nice results on the compression ratio. One thing worth measuring alongside file size is actual on-device inference behavior after quantization. We ran FP32 and INT8 variants of a similar-scale model on Snapdragon 8 Gen 3 through Qualcomm AI
2.
▲
by
ashishdhiman23
7mo ago
The dynamic budget negotiation idea is interesting. One thing I've found working with on-device AI models is that the "budget" isn't just about what the device can handle in theory — it shifts meaningfully between runs o