5 ms·
I’m a self taught ML engineer, and when I was starting on my ML journey I was incredibly frustrated by cloud services like AWS/GCP. a)they were super expensive,
by ilmoi 6y ago
I’m a self taught ML engineer, and when I was starting on my ML journey I was incredibly frustrated by cloud services like AWS/GCP. a)they were super expensive, b)it took me longer to setup a working GPU instance than to learn to build my first model!
So I built https://gpu.land/ https://gpu.land/.
It’s a simple service that only does one thing: rents out Tesla V100s in the cloud.
Why is it awesome?
- It’s dirt-cheap. You get a Tesla V100 for $0.99/hr, which is 1/3 the cost of AWS/GCP/Azure/[insert big cloud name].
- It’s dead simple. It takes 2mins from registration to a launched instance. Instances come pre-installed with everything you need for Deep Learning, including a 1-click Jupyter server.
- It sports a retro, MS-DOS-like look. Because why not:)
The most common question I get is - how is this so cheap? The answer is because AWS/GCP are charging you a huge markup and I’m not. In fact I’m charging just enough to break even, and built this project really to give back to community (and to learn some of the tech in the process).
HN special: email me a few lines about yourself and what you’re working on and get $10 in free credit. I’m at hi@gpu.land.
Otherwise I’m around for any questions!
- taf2 6y agoBecause a lot of the data could be sensitive do you sign baa’s for HIPAA?
- ilmoi 6y agoSorry I missed this. Yes, more than happy to. Email me at hi@gpu.land.
- chovybizzass 6y agowas going to mine but you say not to.
- ilmoi 6y agothanks for reading the FAQ! Most people don't seem to bother:)
- Sanzig 6y agoIs general CUDA development permitted, or can I only use this for deep learning? There's a simulation problem that I encounter often at work which should parallelize very well, and I was considering getting my feet wet with CUDA by writing a solver for it.
- ilmoi 6y agoCUDA develpment is totally fine. The service was designed with DL in mind, but that shouldn't stop CUDA dev in any way.
- ogiberstein 6y agoWow, sounds really useful! Will check it out
- liuliu 6y agoThank you! This looks indeed cheap. Can you share more stories on data persistence / checkpointing? If I have a job that requires 8 V100 with 3 days, what types of reliability am I looking at?
- ilmoi 6y agoExactly the same as you would be had you been running an EC2 instance at AWS. The machines are hosted, maintained and managed in exactly the same way in a whitelabel DC.
- sacheendra 6y agoThis I highly unlikely. Hyperscalers like AWS have custom power delivery, rack organization, redundant networking, etc. which make their instances reliable. Long running jobs often use checkpoints which require high speed networking and storage, which I don't see an option for. Eg, I cam get EC2 instances with 100gbps networking. Great job starting the service! But, I think you have ways to go before reaching Hyperscalers level reliability.
- shrubble 6y agoIt's highly likely that the reliability of this service will be as good or better than AWS. I have servers hosted in a similar quality datacenter with literally 9 years of uptime. It would be 10 but the customer shut down his services... By the way, unless something has changed, you can't get more than 10gbps of throughput of a single tcp stream in those 100gbps setups...
- nixgeek 6y agoWhy do you believe an individual stream is limited to 10Gbps?
- shrubble 6y agoThey have stated as much in the documentation. See various speeds and feeds depending on instance size and type, and whether inside or outside the VPC at e.g. here: https://d1.awsstatic.com/events/reinvent/2019/REPEAT_2_Deep-dive_into_100G_networking_&_Elastic_Fabric_Adapter_on_Amazon_EC2_CMP334-R2.pdf https://d1.awsstatic.com/events/reinvent/2019/REPEAT_2_Deep-...
- ashish01 6y agoNice! I really like the top up feature. Really gives me peace of mind that I will not blow up my budget accidentally. I have been using https://datacrunch.io/ https://datacrunch.io/ which has almost same feature set.
- etaioinshrdlu 6y agoCan you use GeForce cards and then make us pinky swear to only use it for blockchain processing applications? This is what LeaderGPU does (https://www.leadergpu.com/#chose-best https://www.leadergpu.com/#chose-best), and they had even lower pricing. Although lately they have had demand outstrip supply. Hetzner used to have GTX 1080 instances for about $100 a month, no longer though, and I'm lucky to be grandfathered in to about 8 of them. I told myself a couple years ago that compute will only get cheaper over time. In reality, compute has gotten MORE expensive over time! I cannot match or scale the compute price I locked in a couple years ago with Hetzner. Some of that is the crypto market raising GPU prices, but it is also NVIDIA's licensing making their cheaper cards unavailable in cloud servers...
- ghoomketu 6y agoFor somebody who was 0 knowledge of ML but has always been wary of the huge costs involved this sure looks like a great offering! But I still don't know what it would cost to get something useful out of it. Can you (or anybody who knows about ML) tell me a very very ballpark amount of what costs it incurs to train a model like www.remove.bg that automatically removes background from photos? I'm not trying to build a clone but curios what sort ot financial investment it takes to make such things.
- perennate 6y agoUnless you're using unsupervised learning or can find a good dataset, most of the cost will be in labeling the data. I'm not too familiar with background removal, there may be some self-supervised/contrastive learning approaches, but generally they don't work as well as supervised learning. Even for a week of training, the compute cost is only $200. Edit: maybe you can get some decent results just with COCO segmentation labels: https://towardsdatascience.com/background-removal-with-deep-learning-c4f2104b3157 https://towardsdatascience.com/background-removal-with-deep-...
- ghoomketu 6y agoDoes it make sense to add random backgrounds to already transparent pngs and then use it to train the models? P.S. Really useful link btw. Thank you. P.P.S. One more question if you don't mind. After I've trained the model what kind of hardware does it require to use that model to convert a png image to transparent? I'm guessing that requirements are different when you already have a trained model? Sorry for all the n00b questions. I'm not trying to create a background removal site btw, just trying to grasp what kind of costs it takes to create and run something like it.
- typon 6y agoHey ilmoi, Do you support multi-node training? Particularly, can I reserve for example a 64-GPU instance via 8 nodes and perform distributed training?
- Proven 6y agoNot much on high speed networking in the answer, which sounds suspicious (it's the most important part)
- ilmoi 6y agoYou can! All machines within a single account are inside a private VLAN, which means they all talk to each other (out of the box, no setup required), but nobody else on gpu.land sees them. See the "Are instances within my account connected?" item in the FAQ - https://gpu.land/faq https://gpu.land/faq If you're going to need 64 GPUs I'll have to increase your account limit (currently 16 GPUs per account). Email me at hi@gpu.land
- teruakohatu 6y agoIs there any persistent storage options? I looked through the faq and comparisons and couldn't find any mention of it.
- ilmoi 6y agoWhen you stop the instances (instead of killing it), you persist data. It costs $0.02/gb/mo, so $4/mo for a 200gb harddrive. I will add to FAQ, thanks for pointing out.
- teruakohatu 6y agoThat sounds very reasonable thank you. I will send you an email soon before I give it a try.
- mraza007 6y agoHey I loved the product you made. As a student who is always concerned about the spending on cloud services as they can be pricey but I’m just curious I have used google Colab in the past and it charges $10/Month so I wanted to know how is it different when compared to google colab I would love to hear your thoughts on this
- ilmoi 6y agoHey! So the way I would describe your progression as a student (at least from my experience): - Traditional ML -> probably can run on your laptop - Simple DL -> colab is great. Cheapest there is. - SOTA DL -> you're probably looking at training times well into 10s of hours / days, so you'll need something that can last longer than 12h of colab time. Plus at this point you're probably sophisticated enough that you want to setup your instance once and start/stop it rather than going through setup with colab every time. That's where https://gpu.land/ https://gpu.land/ fits it. So tl;dr; - absolutely use colab first (it's cheapest), but when you outgrow it, consider gpu.land.
- mraza007 6y agoperfect and it totally makes sense thanks for clarifying
- benohear 6y agoThis is all very cool, and FWIW I also love the retro design. Sorry if I missed it, but one thing I couldn't find on the website is what legal entity is behind the service. It would be important for me and the organisations I work with to know who we are trusting with our data. It might even be a legal obligation to have that info on the site depending on which country you are located in (Germany for sure, not sure about others).
- ilmoi 6y agoHey - for sure. Because I'm the only team member, the entity behind the business is registered as a "sole trader". If you email me on hi@gpu.land I'm happy to share the details.