3 ms·
I think it’s been bog standard practice to run flask via uwsgi or gunicorn with async workers and use multiple process based workers per deployed server unit (e
by mlthoughts2018 6y ago
I think it’s been bog standard practice to run flask via uwsgi or gunicorn with async workers and use multiple process based workers per deployed server unit (eg per pod in Kubernetes).
What matters is that the cumulative latency & throughput solve your problem, not how fast you can make one singular async worker thread.
I figure most people running complex web services in production would just do an eye roll at this post. Nobody's going to switch to PyPy for any of this.
My team at work runs several complex ML workloads, and we use the exact same container pattern for every service running gunicorn to spawn X async workers per pod and then scale pods per service to meet throughput requirements. Sometimes we also just post complex image processing workloads to a queue and batch them to GPU processor workers. In all these use cases, super low effort “just toss it in gunicorn running flask” has worked without issue for services supporting up to peak load of thousands to hundreds of thousands of requests per second.
- 7kmph 6y agoCould you share your company’s website?
- mlthoughts2018 6y agoNo, I can’t speak on their behalf on Hacker News, so it is important to me to stay disconnected from my employer.