3 ms·
What's the best way to run inference on the 70B model as an API? Most of the hosted APIs including HuggingFace seem to not work out of the box for models that
by OkGoDoIt 3y ago
What's the best way to run inference on the 70B model as an API? Most of the hosted APIs including HuggingFace seem to not work out of the box for models that large, and I'd rather not have to manage my own GPU server.