3 ms·
Lots of ML applications don't generally care about latency, and throughput only matters per dollar (as you can just scale horizontally). To my knowledge, most
by upbeat_general 3y ago
Lots of ML applications don't generally care about latency, and throughput only matters per dollar (as you can just scale horizontally).
To my knowledge, most inferencing (at least for simpler models) happens on cheaper, slower GPUs that have better throughput/dollar.