2 ms·
Nice results on the compression ratio. One thing worth measuring alongside file size is actual on-device inference behavior after quantization. We ran FP32 and
by ashishdhiman23 7mo ago
Nice results on the compression ratio. One thing worth measuring alongside file size is actual on-device inference behavior after quantization.
We ran FP32 and INT8 variants of a similar-scale model on Snapdragon 8 Gen 3 through Qualcomm AI Hub. Inference barely changed (0.176ms vs 0.187ms) but peak memory actually went up slightly (121MB vs 124MB) which was counterintuitive. The file was 70% smaller but runtime characteristics didn't follow the same pattern.
Would be interesting to see TFLite inference timings on actual mobile hardware alongside the size numbers. Compression without on-device profiling can be misleading.