4 ms·
Yes, these optimizations work with existing tflite models, so long as the quantized operators they use are supported in XNNPACK.
by Marat_Dukhan 5y ago
Yes, these optimizations work with existing tflite models, so long as the quantized operators they use are supported in XNNPACK.
- elephantum 5y agoI see, in order to benefit, model has to be quantized. It is not super clear which kinds of quantization are supported. Both Fp16 and Int8?
- Marat_Dukhan 5y agoIn order to benefit from optimizations in *this blog post* the model needs to be quantized to 8-bit integers. However, XNNPACK supports floating-point inference as well (including with FP16 weights), see https://blog.tensorflow.org/2020/07/accelerating-tensorflow-lite-xnnpack-integration.html https://blog.tensorflow.org/2020/07/accelerating-tensorflow-...
- elephantum 5y agoThanks!