5 ms·
So basically, your calculations for deciding whether purchasing physical GPUs or using cloud GPUs is more cost effective should add $5,000 (50 hours at $100 an
by illumin8 8y ago
So basically, your calculations for deciding whether purchasing physical GPUs or using cloud GPUs is more cost effective should add $5,000 (50 hours at $100 an hour seems reasonable) onto the cost of any physical rig, and assume you have a spare 50 hours of opportunity cost to set it up.
All the ML researchers lusting for bare metal need to decide if they want to be "Gentoo Ricers" - https://funroll-loops.teurasporsaat.org/ https://funroll-loops.teurasporsaat.org/ - or if they just want to train models and get real work done.
- _Wintermute 8y ago50 hours is definitely an outlier. I'm pretty clueless when it comes to linux and drivers, and I managed to get an nvidia GPU working on ubuntu in about an hour or two of swearing and a dozen reboots.
- tntn 8y agoOr assume you aren't an outlier and can do it in 0.5-3 hours.
- swebs 8y agoWhy spend 50 hours figuring it out from scratch again if the guy in the article already told you exactly what you need to do?
- danieldk 8y agoExcept that I set up CUDA and Tensorflow in less than an hour. Admittedly, this is on Ubuntu (my colleagues don't like wild distro experiments), which may be easier. I compile every Tensorflow version (to compile with march=native). It is usually uneventful and takes a couple of minutes of my time. By the way, a Tesla card also costs > 5000 Euro and then add the cost of a machine with enough fast cores, memory, and disk space and you are north of 25000.
- tntn 8y agoThe author says they chose Manjaro because they couldn't get Ubuntu to work.
- p1esk 8y agoI've just installed it all from scratch on Ubuntu 16 in about 20 minutes. Just followed the instructions on TF website exactly. Didn't compile TF though, don't see the need.
- tntn 8y agoI know; see my other comments on the thread. I was responding directly to "Ubuntu may be easier" to indicate that the distro difference is not the cause of the authors problems. They've clearly got a messed up system or install procedure.
- perfinion 8y agoThere's definitely a need its waaay faster with -march=native. https://blog.perfinion.com/2018/07/tensorflow-cpu-supports-instructions/ https://blog.perfinion.com/2018/07/tensorflow-cpu-supports-i...
- p1esk 8y agoDid I miss it, or have you not compared GPU TF binary to GPU TF source? I thought that was the goal of your blog post as declared in the first paragraph...
- scottlegrand2 8y agoIn my experience, the same mindset that cannot figure out clear directions on how to set up a GPU machine is similarly incompetent at cleaning up data and understanding machine and deep learning models sufficiently to produce useful results. For bonus points, many of them seem to look down on those that can as "the help" or as "Ops." That said, configuring a GPU machine is (still) ridiculously complicated IMO. It's just that I've been doing it for almost 20 years.
- leblancfg 8y agoAuthor here. Good point, but 50 hours seems like an outlier for sure: I'm no expert whatsoever when it comes to debugging hardware in Linux. Assuming setup time across all first-time builds is Poisson-distributed, feels like the peak of that curve is more like 2-3 hours. Certainly something to keep in mind if considering doing it, though. Thoughts: * If I was setting up multiple identical machines, I'd only need to solve the issue once. * Details matter: Switching from VGA to HDMI connector made a world of difference. Me not noticing that earlier cost me ~2 days of trial-and-error to go to waste.