3 ms·
Highly recommend quantizing the model (https://pytorch.org/tutorials/recipes/recipes/dynamic_quantization.html https://pytorch.org/tutorials/recipes/recipes/dyn
by crabbycarrot 4y ago
Highly recommend quantizing the model (https://pytorch.org/tutorials/recipes/recipes/dynamic_quantization.html https://pytorch.org/tutorials/recipes/recipes/dynamic_quanti...). I converted the large model to use int8, and I'm able to run it 5x real-time on CPU with pretty low RAM requirements with still very good quality.