4 ms·
Thanks! Haven't seen that paper, I'll check it out. I think quantization only helps with inference speed if the network is running on CPU with negligible gains
by laibert 9y ago
Thanks! Haven't seen that paper, I'll check it out. I think quantization only helps with inference speed if the network is running on CPU with negligible gains on GPU (Tensorflow only supported CPU on mobile last I looked which was a while ago). However your app is already super fast so don't I think anyone would notice if it was marginally faster at this point!
- timanglade 9y agoYeah I was surprised it became so fast once I started using small networks. I actually toyed with the idea of slowing down the transition to results artificially to provide better UX lol