3 ms·
what is the special hardware/setup that achieves mid-single digit number of microseconds latency for deep learning inference you referred to?
by jstrong 5y ago
what is the special hardware/setup that achieves mid-single digit number of microseconds latency for deep learning inference you referred to?
- dcolkitt 5y agoMy understanding is that O(5 uS) is achievable on optimized FPGAs with reasonably large networks. Because of the parallelization, large networks don’t add that much more latency as long as you have enough gates. But I have little experience on FPGA stacks, so can’t say for sure. Even in software, I’ve been able to hit O(15 uS) using optimized FANN libraries. But the nets are far smaller than deep, and pretty ruthlessly pruned and compressed. Another trick that helps is pre-differentiating across all the variables you don’t expect to change on a latency critical event. E.g. if you’re running a liquidity take strategy, you can pre-differentiate assuming the opposite touch size and deep book stays constant, because you’re only gonna act following on an aggressor trade at the touch.
- fighterpilot 5y ago> My understanding is that O(5 uS) is achievable on optimized FPGAs with reasonably large networks. Putting aside whether it's technically possible, do you know if any groups are actually having good success with this approach (NNs on microstructure) in live trading?