3 ms·
K80s and K40s have a K. That's two generations old. We are currently on P and about to be V. If you think K80s are good, you are far far behind the times in mac
by chronic940 9y ago
K80s and K40s have a K. That's two generations old. We are currently on P and about to be V. If you think K80s are good, you are far far behind the times in machine learning. A single 1080Ti even in a 4U server outperforms 2 K80s.
- sgt101 9y agoOk - p100's? The thing is that a K80 costs.... £4k what is it that people are paying for?
- dragandj 9y agoECC memory (not that everyone really need that)...
- sgt101 9y agoIn a tensorflow job I find it hard to imagine that a memory error would make a jot of difference... wouldn't the error simply be removed by the network training? If the error was in the execution then I imagine that the epoch would fail and it would then just be a matter of restarting.
- arnon 9y agoIf you're running a rackmount server, you need the Tesla series. We found that the GeForces tend to burn out when under heavy load, whereas we've not had a single Tesla series ever burn out.
- maksimum 9y agoOoh ooh tell us more. Which GeForces and which server chassis? Adequate power supply?
- sgt101 9y agoYes - more details please. We've been using Titan-X and 1080 Maxwells in some Broadberry 4u chassis for the last year/18mths and we've had no burnouts so far. I'm buying replacement pascals, and I can't justify teslas when I can get 8 geforces for the price of 1...
- arnon 9y agoDell R720 / R730s, dual GPU (typically with K40m or K80) in there with 1100W dual redundant PSUs and the GPU enablement kit. We also set our fans to constant 70% min, to keep the airflow good. On some servers we introduced the GTX 1080, either along-side a K40/K80 or two per chassis (see http://arnon.dk/how-does-the-nvidia-gtx-1080-stack-up-against-the-nvidia-tesla-k40/ http://arnon.dk/how-does-the-nvidia-gtx-1080-stack-up-agains...). They actually work 15% faster on average compared to the Tesla K series (Remember it's a 5 year old card), but they just stop working after a few months, or return inconsistent results for some operations. Now, we're not doing graphics with them. We have a GPU database called SQream DB - and we depend on the results to be correct. In the end, they didn't make a lot of sense for us to deploy in a production environment, so back to the Tesla series we went.