4 ms·
AWS has V100's available so most Universities with a decent budget should be able to swing this.
by deeeeeplearning 6y ago
AWS has V100's available so most Universities with a decent budget should be able to swing this.
- belval 6y ago24$/hour for 8 V100 to be exact.
- sudosysgen 6y agoIt's interesting how renting 8 GPUs costing 7000$ each costs about the same as renting the services of the average US worker.
- smabie 6y ago24/hour comes to a salary of $49,920, working 8 hours a day, 5 days a week, no vacation, no holidays. Meanwhile the GPUs cost $56,000 total. So the numbers aren't really very far off.
- sudosysgen 6y agoThey kind of are, these GPUs last multiple years. Anyways, it's just an interesting observation.
- smabie 6y agoI mean you also have to take into account electricity, which I assume is pretty expensive.
- sudosysgen 6y agoIt's only about 2.5kW, so around 1-2$ per hour for the full system
- notsuoh 6y agoNot even that, spot pricing on an 8 GPU instance (the 16xl I believe, the larger one has the same number of GPUs but more memory per GPU) is something like $6/hr. I use this for personal projects sometimes, get all data in S3, a good launch template, and then spin up a spot instance and be super efficient about training quickly. I've even run evals on a separate, cheaper, machine so the 16xl can spend all its time training. It's still not "cheap", but $50 for 8 hours of training on a machine like that with $64k of GPUs on board is really not bad.
- bravura 6y agoIf you wrote a blog post about how to streamline this, I would read it and upvote it.
- fxtentacle 6y agoExcept that cloud-ified V100s are significantly less powerful than if you have direct access to the hardware. Last time I checked, in AWS they're actually external devices mapped in over GBit ethernet, which is significantly slower than the 8GB/s that PCIe x4 has.
- p1esk 6y agoI routinely switch between AWS 8x V100 instances and on-premise 8x V100 servers and I observe no difference in speed (time per epoch).
- Reelin 6y agoPresumably that depends on maximum PCIe bandwidth consumption before your workload bottlenecks elsewhere? A 2018 benchmark (https://www.pugetsystems.com/labs/hpc/PCIe-X16-vs-X8-with-4-x-Titan-V-GPUs-for-Machine-Learning-1167/ https://www.pugetsystems.com/labs/hpc/PCIe-X16-vs-X8-with-4-...) seems to indicate that x8 isn't generally a bottleneck for common (at the time) workloads. x8 is a far cry from the claimed gigabit ethernet though!
- p1esk 6y agoAWS is tricky in terms of how storage is provisioned - I don't remember details, but it's easy to put your datasets on storage that is connected to your GPU servers over 1Gb link. That could easily become a bottleneck. Datasets should live on Elastic Block Storage or something like that, over high speed links. Again, it's been a while since I looked into that, so I don't remember the details.
- Reelin 6y agoThe earlier comment claimed that the GPUs (!!!) were located elsewhere on the network; I suspect that the scenario you describe is what they intended to refer to. (IIRC AWS offers compute optimized instances with a volume that's guaranteed to be backed by blocks on a local NVMe drive.)