3 ms·
Why would I use this over deploying the model to a lambda function aside from lack of GPU? (not trying to be confrontational, genuinely don't know) Won't lambda
by cameronfraser 6y ago
Why would I use this over deploying the model to a lambda function aside from lack of GPU? (not trying to be confrontational, genuinely don't know) Won't lambda functions scale as needed? How does this compare cost wise?
- calebkaiser 6y agoGreat question. We actually experimented with Lambda before ever building Cortex. We ran into several issues, the three easiest to list are: 1. Size limits. Lambda limits deployment packages to 250 mb uncompressed, and puts an upper bound on memory of 3,008 mb. That's not nearly big enough for a lot of models, particularly bigger deep learning models. 2. As you mentioned, GPU inference is supported on Lambda, and for many models, GPUs are necessary for serving with acceptable latency. 3. Lambda instances can only serve one request at a time. With how slow ML inference can be—especially if you need to call another API or preform some IO request—it's easy to lock up Lambda instances for full seconds just to serve one prediction. The TL;DR is that while Lambda works for some use-cases, it in general lacks the flexibility and customizability needed for most inference use-cases.