4 ms·
You could suspend it to ram, and only wake it up on request, it takes 2 seconds on my box.
by zoobab 3mo ago
You could suspend it to ram, and only wake it up on request, it takes 2 seconds on my box.
- Aurornis 3mo agoIt’s not a cost savings relative to paying API prices even if you’re suspending it. This is an option if you must run local inference, you’re not sensitive to speed, and the budget is low. It’s not going to be cheaper than paying API prices for the model though.