3 ms·
> While FastAPI does support async calls at the web request level, there is no way to call model predictions in an async manner. This confuses me. How is that
by anderskaseorg 4y ago
> While FastAPI does support async calls at the web request level, there is no way to call model predictions in an async manner.
This confuses me. How is that FastAPI’s fault? Can’t you just asynchronously delegate them to a concurrent.futures.ThreadPoolExecutor or concurrent.futures.ProcessPoolExecutor? What does Starlette provide here that FastAPI doesn’t? If the FastAPI limitations are due to ASGI, shouldn’t Starlette have the same limitations?
- timliu99 4y ago> While there are several different methods to use your own executor pools or potentially use shared memory for a large model, all of these solutions are not first-class solutions for ML use cases... Definitely not FastAPI's fault and yes Starlette has the same limitations. BentoML builds additional ML features/abstractions on top of Starlette. We introduced a "runner" concept which automatically creates separate processes for models to run in.
- anderskaseorg 4y agoGreat—but then one ought to be able to performantly use that same runner equally well from FastAPI, Starlette, Quart (ASGI port of Flask), or any other ASGI framework. You’ve decided to build a convenient integration with Starlette instead of the others, but it’s weird to frame this as an argument that other frameworks are ill suited for this domain.
- chaoyu_ 4y agoThat’s is true, it would be more accurate to say FastAPI (or any ASGI) framework along is not enough for ML model serving, you need batching and runner for performance. And some additional Ml focused features for integration with other ML tools