4 ms·
If I understand this correctly, I should be able to stand up an API running arbitrary models (e.g. from Hugging Face), and it’s not quite charged by the token b
by lemming 1y ago
If I understand this correctly, I should be able to stand up an API running arbitrary models (e.g. from Hugging Face), and it’s not quite charged by the token but should be very cheap if my usage is sporadic. Is that correct? Seems pretty huge if so, most of the providers I looked at required a monthly fee to run a custom model.
- BoredPositron 1y agoRunpod, vast, coreweave, replicate... just a bunch of alternatives that let you run serverless GPU inference.
- _zoltan_ 1y agoyou can't just sign up for coreweave, can you?
- BoredPositron 1y agoWe wrote them an email.
- lexandstuff 1y agoYes, that's basically correct. Except be warned that the cold start times can be huge (30-60 seconds). So scaling to 0 doesn't really work in practice, unless your users are happy to wait from time to time. Also, you also have to pay a small monthly fee for container storage (and a few other charges iirc).