3 ms·
Have you ran the model in full FP16? It is possible a lot of performance is lost when running quantized versions.
by tarruda 2y ago
Have you ran the model in full FP16? It is possible a lot of performance is lost when running quantized versions.