3 ms·
The hardware you need and amount of time you run the model really depends on which model you've fine-tuned and what your model will be doing. If you don't need
by batch12 2y ago
The hardware you need and amount of time you run the model really depends on which model you've fine-tuned and what your model will be doing. If you don't need quick responses, you could use a cpu via llama.cpp. If you are only doing something like summarization you could do the inference in batches and start and stop your GPU resources between them. If you're doing chat around something predictable, like a product, you could cache common questions and responses and use smaller model to pick between them.