3 ms·
Can someone please share the current state of deploying Pytorch models to productions? TensorFlow has TF serving which is excellent and scalable. Last I checked
by cweill 5y ago
Can someone please share the current state of deploying Pytorch models to productions? TensorFlow has TF serving which is excellent and scalable. Last I checked there wasn't a PyTorch equivalent.
I'm curious how these charts look for companies that are serving ML in production, not just research. Research is biased towards flexibility and ease of use, not necessarily scalability or having a production ecosystem.
- TheGuyWhoCodes 5y agoHonestly today there are way too many options. There is TorchServe but I haven't used it so I'm not sure how production ready it is. You have Nvidia's triton server which support cpu and gpu with tf1,tf2,pytorch,onnx and tensorRT. You have onnx runtime which can run on cpu and gpu and there are convertors from tf and pytorch to onnx. Then you have cloud based solutions like AWS sagemaker, elastic inference endpoints and even Inf1 instances that use AWS Inferentia chips which you would run with the Neuron SDK, they even have TensorFlow serving containers with built it support for Inferentia. End of the day it really depends on your model, size, latency, inference runtime and the cost obviously. And that's before optimizations like FP16, BFLOAT16, TF32, INT8, pruning, layers rewrite, getting rid of batch normalization etc. Then you have up and coming solutions like Neural Magic (not associated) deepsparse to create sparse models for inference. And that's just for cloud if you are talking about edge ml it's even more down the rabbit hole..