4 ms·
If you care about inference time, then you'll do two things. 1. Train a student model from your fine-tuned model. (Known as "knowledge distillation"). 2. Quan
by solresol 2y ago
If you care about inference time, then you'll do two things.
1. Train a student model from your fine-tuned model. (Known as "knowledge distillation").
2. Quantize the student model so that it uses integers.
You might also prune the model to get rid of some close-to-zero weights.
This will get you a smaller model that can probably run OK on a CPU, but will also be much more efficient on GPU.
Next: architect your code so that the inference step sits behind a queue. You do not generally want to have the user interface waiting on a inference event because you can't guarantee latency or resource availability, and your model's inference processing will be the biggest slowest thing in your stack, so you can't afford to overprovision.
So have a queue of "things to infer", and having your inference process run in the background chomping through the backlog, storing the results in your database. When it infers something, somehow notify your front-end clients that it's ready in the database for them to retrieve. In this model, you can potentially run your model somewhere cheaper than AWS (e.g. a cheaper provider, a machine under your desk).
Or, for the genius move: compile the model to ONNX and run it in a background thread in the users' browser, and then you don't have to worry about provisioning; users will wonder why their computer runs so slowly though.
- FezzikTheGiant 2y agoThis is really helpful. I'm really new to this, what's the best way to educate myself on fine-tuning and deploying a LLM in the most cost and time efficient way?
- solresol 2y agoI've got video recordings of my lectures. I haven't edited them properly yet, but I can share them with you privately. Email me (it's in my profile) and I'll get them to you.