4 ms·
That depends on the model architecture and how it was written since that informs the size of the search space. The typical range is 10 mins to 10 hours. It won
by jakestevens2 1y ago
That depends on the model architecture and how it was written since that informs the size of the search space.
The typical range is 10 mins to 10 hours. It won't be fast but you only have to do it once and then those optimizations are set for every forward pass.
- sitkack 1y agoDo you learn the capabilities of the underlying hardware relative to the kernel src? You should be able to start predicting perf using learned static profiling.
- jakestevens2 1y agoNot today but we will implement memoization of kernels for each hardware backend, yes.