3 ms·
I’m using ONNX Runtime with 4-bit quantization on a Raspberry Pi 4. I preload the quantized model into shared memory so multiple processes can reuse it. Evict o
by byte-bolter 1y ago
I’m using ONNX Runtime with 4-bit quantization on a Raspberry Pi 4.
I preload the quantized model into shared memory so multiple processes can reuse it.
Evict old sessions by LRU when I hit a 1 GB RAM cap.
For batching, I accumulate inputs over 50 ms to boost throughput without hurting latency.
So far I get ~15 RPS on a 7 B Llama 2 model.
- tynskid2025 1y agoDo you have a repo outline of how you did this, I would be so grateful