3 ms·
Yes. You can see this affecting perf benchmarks as well. Usually the cheapest inference providers either use approximations like tanh instead of sigmoid, nvfp4
by augment_me 22d ago
Yes. You can see this affecting perf benchmarks as well. Usually the cheapest inference providers either use approximations like tanh instead of sigmoid, nvfp4 quantizarion, etc.
There was a post here the other day highlighting this by showing the benchmark perf of different I defence providers, it's a fantastic area to cheap out in, because you can never really tell if a model is 75% good or 83% good on some specific benchmark when you use it to build your own stuff