3 ms·
I'm curious about this. Can you point me to, e.g. some example code for setting up an inference endpoint with a base llama2 model on modal.com?
by easygenes 3y ago
I'm curious about this. Can you point me to, e.g. some example code for setting up an inference endpoint with a base llama2 model on modal.com?
- SparkyMcUnicorn 3y agoHere's one if their tutorials using vLLM, and they have a few other guides and example repos as well. https://modal.com/docs/guide/ex/vllm_inference https://modal.com/docs/guide/ex/vllm_inference https://github.com/modal-labs https://github.com/modal-labs Alternatively, Runpod is fairly cheap and easy to get stuff running in a few minutes and can be point/click only using their templates. https://www.runpod.io/console/gpu-secure-cloud?template=f1pf20op0z https://www.runpod.io/console/gpu-secure-cloud?template=f1pf... ("serverless" example) https://github.com/ashleykleynhans/runpod-worker-oobabooga https://github.com/ashleykleynhans/runpod-worker-oobabooga
- easygenes 3y agoThanks for that. I've used RunPod GPU cloud to setup vLLM as an Open-AI API compatible endpoint before, but haven't tried any of the serverless options yet.