3 ms·
> Queueing in front of ML models is important This sounds so clean! A super fast NoSQL to hold incoming requests. Could even throttle your free users and prio
by _akhe 2y ago
> Queueing in front of ML models is important
This sounds so clean! A super fast NoSQL to hold incoming requests.
Could even throttle your free users and prioritize the paying customers.
Btw, is the project online now?
- mitjam 2y agoHere is a queueing api server for self hosted inference backends: https://github.com/aime-team/aime-api-server https://github.com/aime-team/aime-api-server from a friend of mine. Very light weight and easy to use. You can even serve models from Jupyter Notebooks with it without needing to worry about overwhelming the server. It just gets slower the more load you send to it.
- _akhe 2y agoReally cool! I like that they have live demos to prove it out. Thanks for sharing
- itake 2y agoI use ML for content moderation at Grab and my side project. Both are online. Both use queues, but Grab's traffic is more organic and real time than the side project's.