9 ms·
Introducing Amazon EC2 P3 Instances
- dharma1 9y agoPrice: p3.2xlarge - $3/hr, p3.8xlarge - $12/hr, p3.16xlarge - $25/hr These look very good for half precision training
- moonbug22 9y agoCome on, no one with any sense pays the on demand price for these things. Watch the spots.
- dx034 9y agoThere are enough companies out there with deep pockets that want to do some ML. They'll pay pay those prices, no questions asked.
- IanCal 9y agoPer the marketing material it's up to a PFLOP of mixed precision (is that the same as just saying "half precision"? or is it 8 bit?) for $25/hour. I can easily see people paying full price for that. Still, spot price is currently $2.40.
- Beltiras 9y agoIt must be a misplaced comma. 15.7 single precision can never translate to 125 mixed.
- dharma1 9y agoThe reason Nvidia quote 120 TLFOPS mixed precision on V100 is because of the new tensor core. https://devblogs.nvidia.com/parallelforall/inside-volta/ https://devblogs.nvidia.com/parallelforall/inside-volta/
- julien_c 9y agoAlso all startups with (basically) unlimited aws credits
- maffydub 9y agoYes, p3.2xlarge in us-east-1b is currently sitting at $0.3204 spot. That's only marginally more than p2.xlarge (at $0.2259 in us-east-1e). I'm sure this will change with demand, though. :(
- deleted 9y ago[deleted]
- dharma1 9y agoYep agreed. Didn't want to post the spot pricing since it changes all the time :)
- deleted 9y ago[deleted]
- plantain 9y agoBut where are the C5 instances? It's been 11 months since Amazon announced Skylake C5's and we're still waiting! https://aws.amazon.com/about-aws/whats-new/2016/11/coming-soon-amazon-ec2-c5-instances-the-next-generation-of-compute-optimized-instances/ https://aws.amazon.com/about-aws/whats-new/2016/11/coming-so...
- STRML 9y agoWaiting for them as well. Most of all, we really need fast-CPU instances with the ENA, not the Intel NIC.
- jsolson 9y agoOut of professional curiosity, what are you looking for from ENA? (I'm an engineer on Google Compute Engine with a deep interest in customer networking use stories, particularly heavy utilization customers, even if they're not my customers :)
- moconnor 9y agoAn exaflop of mixed-precision compute for $250M over 3 years. That’s ballpark what the HPC community is paying for their exaflop-class machines. You’d still build your own for that money, I think, but it’s an interesting datapoint.
- dx034 9y agoHow long if you build it your own incl electricity prices? If margins are similar to other EC2 instances, you'd probably break-even after 6 months or so. Which makes EC2 uneconomical for any lab/company that can utilise the cluster 24/7. Still nice if you quickly need to get some model results though.
- dharma1 9y agoIf you're going to be running it for 24x7 for 3 years, I think it'd be worth doing the apples-to-apples comparison of buying your own V100s vs renting them from AWS. The DGX Station with 4 V100s is $70k
- dexterdog 9y agoIt wouldn't be $250K over 3 years. It would be $250K up-front to get the lowest current pricing.
- againa 9y agoUse reserve instances or use spot. The price decrease substantially. Then when you don’t need it... you don’t pay it... it’s a good deal
- jerianasmith 9y agoyaah
- Smerity 9y agoThe P3 instances are the first widely and easily accessible machines that use the NVIDIA Tesla V100 GPUs. These GPUs are straight up scary in terms of firepower. To give an understanding of the speed-up compared to the P2 instances for a research project of mine: + P2 (K80) with single GPU: ~95 seconds per epoch + P3 (V100) with single GPU: ~20 seconds per epoch Admittedly this isn't exactly fair for either GPU - the K80 cards are straight up ancient now and the Volta isn't sitting at 100% GPU utilization as it burns through the data too quickly ([CUDA kernel, Python] overhead suddenly become major bottlenecks). This gives you an indication of what a leap this is if you're using GPUs on AWS however. Oh, and the V100 comes with 16GB of (faster) RAM compared to the K80's 12GB of RAM, so you win there too. For anyone using the standard set of frameworks (Tensorflow, Keras, PyTorch, Chainer, MXNet, DyNet, DeepLearning4j, ...) this type of speed-up will likely require you to do nothing - except throw more money at the P3 instance :) If you really want to get into the black magic of speed-ups, these cards also feature full FP16 support, which means you can double your TFLOPS by dropping to FP16 from FP32. You'll run into a million problems during training due to the lower precision but these aren't insurmountable and may well be worth the pain for the additional speed-up / better RAM usage. - Good overview of Volta's advantages compared to event the recent P100: https://devblogs.nvidia.com/parallelforall/inside-volta/ https://devblogs.nvidia.com/parallelforall/inside-volta/ - Simple table comparing V100 / P100 / K40 / M40: https://www.anandtech.com/show/11367/nvidia-volta-unveiled-gv100-gpu-and-tesla-v100-accelerator-announced https://www.anandtech.com/show/11367/nvidia-volta-unveiled-g... - NVIDIA's V100 GPU architecture white paper: http://www.nvidia.com/object/volta-architecture-whitepaper.html http://www.nvidia.com/object/volta-architecture-whitepaper.h... - The numbers above were using my PyTorch code at https://github.com/salesforce/awd-lstm-lm https://github.com/salesforce/awd-lstm-lm and the Quasi-Recurrent Neural Network (QRNN) at https://github.com/salesforce/pytorch-qrnn https://github.com/salesforce/pytorch-qrnn which features a custom CUDA kernel for speed
- agibsonccc 9y agoGreat write up as usual! Could you elaborate more on the python overhead a bit? We have fp16 support running in dl4j but I don't think we've really done much with volta yet beyond get it working. In practice, (especially when we do multi gpu async back round loading of data) we find gpus being data starved. I would love to compare support for what you're seeing with pytorch.
- science404 9y agoWhy Ireland and not the UK? I can imagine a lot of startups/banks in London could use this... Brexit fears?
- maffydub 9y agoI wouldn't read too much into this - Amazon's Ireland region was deployed earlier (2008?) than London (2016?) and seems to receive updates earlier too.
- remus 9y agoLondon only came online relatively recently, maybe there's some operational stuff getting in the way of deploying? Or perhaps London has relatively few users at the moment, so the number of clients who will be able to take advantage of more specialised instances is also relatively low?
- moonbug22 9y agoLondon region is tiny and not the region to go to unless you have specific geolocation requirements.
- tt293 9y agoBecause Ireland is so close to the UK that the latency will not matter?
- maxehmookau 9y agoThe London Region is really new. Amazon are using it as a way in to UK-only projects (specifically healthcare due to NHS regulations). I suspect their datacentres are considerably smaller than that of Ireland along with their client-base for the moment. I wouldn't read too much in to it, brexit-wise.
- puzzle 9y agoThe Dublin data center is probably much larger than the one in London, for starters.
- rsynnott 9y agoIreland is one of the older and larger regions, and generally gets new stuff at the same time as Virginia or very shortly afterwards. London and Frankfurt tend to be delayed a bit. See https://aws.amazon.com/about-aws/global-infrastructure/regional-product-services/ https://aws.amazon.com/about-aws/global-infrastructure/regio...
- eggie5 9y agoHere's my results: Testing new Tesla V100 on AWS. Fine-tuning VGG on DeepSent dataset for 10 epochs. GRID 520K (4GB) (baseline): * 780s/epoch @ minibatch 8 (GPU saturated) V100(16Gb): * 30s/epoch @ minibatch 8 (GPU not saturated) * 6s/epoch @ minibatch 32 (GPU more saturated) * 6s/epoch @ minibatch 256 (GPU saturated)
- dharma1 9y agoThanks! Curious how this would scale on the 8x or 16x instances
- sethgecko 9y agoIs there an AMI that comes with Tensorflow/keras with GPU support preinstalled or you have to do it yourself?
- Smerity 9y agoAmazon offer an official AMI which comes preloaded with various deep learning frameworks: MXNet, TensorFlow, CNTK, Caffe/2, Theano, Torch and Keras. For the P3 (Volta V100) instances you'll want to ensure you use an AMI preloaded with CUDA 9, though not all DL frameworks are happy with that yet. https://aws.amazon.com/amazon-ai/amis/ https://aws.amazon.com/amazon-ai/amis/
- mv4 9y agoDidn't know they had an AMI like that. Thank you.
- sipherhex 9y agoBe careful with the non-CUDA 9 AMIs. CUDA 8 programs will run, but terribly slowly as they JIT their GPU code without optimization for Volta. You want the CUDA 9 AMI version (https://aws.amazon.com/marketplace/pp/B076TGJHY1?qid=1509090457271 https://aws.amazon.com/marketplace/pp/B076TGJHY1?qid=1509090...), but it currently only has MXNet and TF. If you need other frameworks there's the NVIDIA AMI (https://aws.amazon.com/marketplace/pp/B076K31M1S?qid=1509090567000 https://aws.amazon.com/marketplace/pp/B076K31M1S?qid=1509090...) and Volta optimized containers for NVCaffe, Caffe2, CNTK, Digits, MXNet, PyTorch, TensorFlow, Theano, Torch, CUDA 9/CuDNN7/NCCL.
- jeffbarr 9y agoMore details in my blog post at https://aws.amazon.com/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/ https://aws.amazon.com/blogs/aws/new-amazon-ec2-instances-wi...
- mcherm 9y agoI thought the comparison of 1 second of computation today to the lifetime computation of older computers (since they were released) was clever.
- ZeroCool2u 9y agoThanks Jeff, I forwarded your blog post to our Chief Scientist.
- SloopJon 9y agoThis post states, "In order to take full advantage of the NVIDIA Tesla V100 GPUs and the Tensor cores, you will need to use CUDA 9 and cuDNN7." What version of TensorFlow does it use? From what I can tell, TensorFlow doesn't fully support the latest versions yet.
- puzzle 9y agoTF 1.4 does, but you need to build it yourself. RC1 is out: https://github.com/tensorflow/tensorflow/releases/tag/v1.4.0-rc1 https://github.com/tensorflow/tensorflow/releases/tag/v1.4.0... All our prebuilt binaries have been built with CUDA 8 and cuDNN 6. We anticipate releasing TensorFlow 1.5 with CUDA 9 and cuDNN 7.
- sumt 9y agoYou can use the new AWS Deep Learning AMI which has a version of TensorFlow enhanced for CUDA 9 and Volta support https://aws.amazon.com/blogs/ai/announcing-new-aws-deep-learning-ami-for-amazon-ec2-p3-instances/ https://aws.amazon.com/blogs/ai/announcing-new-aws-deep-lear...
- sipherhex 9y agoChris from NV here. You can also get a full compliment of DL framework containers, as well as CUDA 9/CuDNN 7/NCCL 2 base container, optimized for Volta by NVIDIA via this AMI https://aws.amazon.com/marketplace/pp/B076K31M1S?qid=1509089922554 https://aws.amazon.com/marketplace/pp/B076K31M1S?qid=1509089...
- bprasanna 9y ago...advanced workloads such as machine learning (ML), high performance computing (HPC), data compression, and crypto__________.
- Yuioup 9y agoHow many bitcoins can you mine out of this on max power and would it be profitable? I'm sure that Amazon has done the math on this but I'm still curious.
- geofft 9y agoIt's not just that Amazon has done the math, it's that sufficiently liquid cryptocurrencies will, by the efficient market hypothesis, quickly gain enough value to make mining on whatever Amazon offers no longer profitable. As soon as you're able to profitably mine without an up-front capital investment, people will take advantage of the arbitrage opportunity until the market adjusts its price, and if the currency is designed at least somewhat competently and has enough of a working market (both of which are definitely true of Bitcoin), that won't take very long. Cryptocurrencies are the invisible robot hand of the market. (Which is, I think, not a claim about whether they're good, but certainly a claim about whether they are to be feared. If you squint hard enough, the giant Bitcoin mines in China are the work of an unfriendly AI employing people to make paperclips.)
- corford 9y agoHmm just tried to spool up a p3.2xlarge in Ireland but hit an instance limit check (it's set at 0), went to request a service limit increase but P3 instances are not listed in the drop down box :(
- avvakum 9y agoSame problem here and it does not seem to be zone specific. I wonder how others worked around this ...
- g105b 9y agoBitcoin?
- pyvpx 9y agobitcoin mining at any hope of profitability comes with custom, specific ASICs.
- psychometry 9y agoRandom question: Why are we still using mostly GPUs for computation rather than CPUs custom-designed for ML tasks?
- fooker 9y agoThe definition of an 'ML task' tends to change.
- dmoy 9y agoFor some definition of "we", we are not.
- Diederich 9y agoCan you expand on that?
- zolthrowaway 9y agoGPUs are quite good at doing arithmetic in parallel. A large part of machine learning is doing arithmetic on large data sets. It makes sense to do these operations in parallel. For example, implementing k-nearest neighbors on a GPU is almost 2 orders of magnitude faster than on a CPU[0]. GPUs just work very well when you have a a lot of data and you are able to run the operations on the data set in parallel. Machine learning seems to fit this model quite well which is why you see many GPUs used in this field. Other things that take advantage of parallelism would be graphics and crypto-currency mining. [0] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.159.9386&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.159...
- scott_karana 9y agoIf you want to offer PaaS with FPGAs or ASICs, by all means. I'm sure there'd be some interest :) ML might be a bit of a moving target though.
- oh-kumudo 9y agoWhy not?
- jerianasmith 9y agoP3 instances no doubt provide a powerful platform and is going to be useful for data compression.
- arnon 9y agoAnd GPU databases that use compression will gain another big advantage
- mamon 9y agoSlightly off-topic but I'm curious: Nvidia Volta is advertised as having "tensor cores" - what does it take for a programmer to use them? Will typical Tensorflow or Cafe code take advantage of it? Or should we wait for some new optimized version of ML frameworks?
- exDM69 9y ago> Will typical Tensorflow or Cafe code take advantage of it? Yes, the support should already be there for both frameworks.
- JeanMarcS 9y agoIf ever you've got password hash to decrypt :)
- DTE 9y agoHi guys, Dillon here from Paperspace (https://www.paperspace.com https://www.paperspace.com). We are a cloud that specializes in GPU infrastructure and software. We launched V100 instances a few days ago in our NY and CA regions and its much less expensive than AWS. Think of us as the DigitalOcean for GPUs with a simple, transparent pricing and effortless setup & configuration: AWS: $3.06/hr V100* Paperspace: $2.30 /hr or $980/month for dedicated (effective hourly is only $1.3/hr) Learn more here: https://www.paperspace.com/pricing https://www.paperspace.com/pricing [Disclosure: I am one of the founders]
- sspiff 9y agoThis is great! I'm looking for a way to run serverless (Amazon Lambda style) GPU operations (preferably using OpenCL). Are there any plans for such a service in your platform?
- DTE 9y agoWe have definitely been thinking a lot about what that would look like (i.e. is it more of a job architecture, an API, clustering, etc). Would love to hear your thoughts on what GPU Lambda might look like. Feel free to hit me up directly dillon [@] paperspace [dot] com if you want to continue the conversation :)
- dylanz 9y agoWe're adding sync support to Worker (which has GPU support) at Iron.io soon! This will allow you to run long running background jobs (current behavior) as well as sync serverless/faas Lambda-like functions within a single API.
- jnbiche 9y agoDigitalOcean for GPUs, awesome! For someone wanting to play around learning more about machine learning, would one of your Standard GPU units be ideal? If so, which one would you recommend? (Or do you think I'd need a dedicated GPU unit?)
- DTE 9y ago
- kshnell 9y agoLooks like Paperspace announced Volta support yesterday: https://blog.paperspace.com/tesla-v100-available-today/ https://blog.paperspace.com/tesla-v100-available-today/ One nice thing here is you can do monthly plans instead of reserved on AWS which is a minimum $8-17k upfront. Really great to see the cloud providers adopting modern GPUs.