6 ms·
We charge from the time you boot a machine until it stops. There's no enforced minimum, but in general it's difficult to get much out of a machine in less than
by mrkurt 3y ago
We charge from the time you boot a machine until it stops. There's no enforced minimum, but in general it's difficult to get much out of a machine in less than 5 seconds. For GPU machines, depending on data size for whatever is going into GPU memory, it could need 30s of runtime to be useful.
- andes314 3y agoDo you offer some sort of keep_warm parameter that removes this latency (for a greater cost)?
- mrkurt 3y agoYou control machine lifecycles. To scale down, you just set the appropriate restart policy, then exit(0). You can also opt to let our proxy stop machines for you, but the most granular option is to just do it in code. So yes, kind of. You just wait before you exit.
- Aeolun 3y agoSo just to confirm, for these workloads, it’d start a machine when the request comes in, and then shut it down immediately after the request is finished (with some 30-60s in between I suppose)? Is there some way to keep it up if additional requests are in the queue? Edit: Found my answer elsewhere (yes).
- sodality2 3y agoHow long does model loading take? Loading 19GB into a machine can't be instantaneous (especially if the model is a network share).
- loloquwowndueo 3y agoThere are no “network shares”. The typical way to store model data would be in a volume, which is basically local nvme storage.
- xena 3y agoWellllllll, technically there is LSVD which would let you store model weights in S3. God that's a horrible idea. Blog time!
- carl_dr 3y agoIt takes about 7s to load a 9GB model on Beam (they claim, and tested as about right), I imagine it is similar with Fly - I've not had any performance issues with Fly.
- bbkane 3y agoI see the whisper transcription article. Is there an easy way to limit it to, say $100 worth of transcription a month and then stop till next month? I want to transcribe a bunch of speeches but I want to spread the cost over time
- IanCal 3y agoProbably available elsewhere but you could setup an account with a monthly spend limit with openai and use their API until you hit errors. $100/mo is about 10 days of speeches a month, how much data do you have? edit - if the pricing seems reasonable, you can just limit how many minutes you send. AssemblyAI is another provider at about the same cost.
- bbkane 3y agoThanks! Maybe 50hr of speeches. It's a hobby idea so I'll check these out when I get some time