3 ms·
TLDR you’ll probably serve it on gpus If it’s a small model you might be able to host it on a regular server with CPU inference (see llama.cpp) Or a big model
by tikkun 2y ago
TLDR you’ll probably serve it on gpus
If it’s a small model you might be able to host it on a regular server with CPU inference (see llama.cpp)
Or a big model on cpu but really slowly
But realistically you’ll probably want to use gpu inference
Either running on gpus all the time (no cold start times) or on serverless gpus (but then the downside is the instances need to start up when needed, which might take 10 seconds)