3 ms·
How did AITemplate performance to state of art inference engine like tvm or onnx runtime ? Did AITemplate optimize/quantify network? Edit: link for TVM https:/
by Narew 4y ago
How did AITemplate performance to state of art inference engine like tvm or onnx runtime ? Did AITemplate optimize/quantify network?
Edit: link for TVM https://tvm.apache.org/ https://tvm.apache.org/
- davidatbu 4y agoI'd love to hear about this too: especially after running the model through an onnx optimizer, like this one [0]. [0] https://github.com/daquexian/onnx-simplifier https://github.com/daquexian/onnx-simplifier
- ipiszy 4y agoAITemplate only supports fp16 data types with fp16 or fp32 accumulation right now. We are working on supporting more data types and quantization. We don't have an official comparison between AITemplate and tvm / onnx for now, but we do have perf numbers like https://github.com/facebookincubator/AITemplate/tree/main/examples/03_bert https://github.com/facebookincubator/AITemplate/tree/main/ex..., https://github.com/facebookincubator/AITemplate/tree/main/examples/05_stable_diffusion https://github.com/facebookincubator/AITemplate/tree/main/ex.... Feel free to run these examples on other frameworks and compare perf.