3 ms·
I help teams run transformers in their production systems on CPU, using my product based on ONNX Runtime. This is a great article, but if you’re using somethin
by binarymax 4y ago
I help teams run transformers in their production systems on CPU, using my product based on ONNX Runtime.
This is a great article, but if you’re using something based on BERT or RoBERTa, you don’t need to do much. Distillation is usually the only step you need to take if you’re really picky, or if your scale is millions of requests per day and you’re not making enough money to support the infrastructure.
I have had mixed results with quantization and sparsification, but IMO it’s just not worth it as they can be unstable.