13 ms·
Nvidia Digits DevBox
- afsina 11y agoI think Boxx Apexx-5 boxes are already on par with these (even more powerful). http://www.boxxtech.com/products/apexx-5 http://www.boxxtech.com/products/apexx-5
- choppaface 11y agoThat unit with only one K40 Tesla appears to be $12,000. If the nVidia box is really $15,000 (with 4 Titans, 9TB SATA, etc), then the nVidia box looks like a much better value (and probably more powerful).
- afsina 11y agoKeep in mind this has a dual socket motherboard. My colleagues bought 2 of those boxes with 4 titan-X and they are actually cheaper than the price on the web site.
- modeless 11y agoNvidia owns deep learning. They are alone at the top. Intel and AMD aren't even in the picture. I think this could end up being a bigger business than graphics accelerators. There's a huge opportunity here for the first company to put out a specialized deep learning chip that can beat GPUs (which is definitely possible; probably by 10x or more).
- Numberwang 11y agoCould you elaborate as to why this is a huge opportunity business wise?
- mryan 11y agoIn the coming years we will see a lot more applications powered by deep learning. If someone releases a chip that provides 10x performance compared to the current GPU-based method, they will sell a lot of chips.
- x0x0 11y agoNo doubt -- the problem is the assumption that 10x performance is possible. Though if anyone can deliver it will (imo) be Intel: it will be a process war, and they seem to be able to do more transistors than anyone else.
- azinman2 11y agoOr IBM: http://research.ibm.com/cognitive-computing/neurosynaptic-chips.shtml#fbid=YUVZuhPyPsg http://research.ibm.com/cognitive-computing/neurosynaptic-ch...
- maaku 11y agoIntel may have to find a different niche. Their Xeon Phi approach to parallelism is very interesting, speaking as someone dabbling in AGI, but is not comparably well suited for deep learning.
- azeirah 11y agoHow well suit are fpga's for deep learning?
- amund 11y agohttps://gigaom.com/2015/02/23/microsoft-is-building-fast-low-power-neural-networks-with-fpgas/ https://gigaom.com/2015/02/23/microsoft-is-building-fast-low...
- Thimothy 11y agoThe answer to the question "How well suited are FPGAs for -insert field where GPUs or vanilla processors do no excel at-?" is always "Far better than the GPUs or proccessors but with an abysmal power usage". In this kind of new developments FPGAs are usually used as a test before moving to ASICs.
- maaku 11y agoBasically everything you interact with on a daily basis involves (deep) machine learning. Everything from advertising to logistics to drug discovery to credit evaluations to packet routing on the internet either involves deep learning in its operation, or is presently implemented with some optimal strategy that was discovered via deep learning. Machine learning runs the world. As for the business opportunity, if you are in one of these industries then one of the prime ways you differentiate from your competitors is how well your algorithm works. And how well it works depends a great deal on how much training data you are able to feed into it, which in this era of information saturation is basically a hardware limitation. Give me a 10x more powerful machine learning system, and I'll give you a few basis points advantage over the competition, and that'll make you, not them, the dominant player.
- hueving 11y agoPacket routing? If there is an example of this, it is very far from the norm. Routing on the Internet is driven first and foremost by peering arrangements.
- maaku 11y agoNow ask yourself how decisions are made regarding new infrastructure deployments. But what I actually had in mind was QoS.
- modeless 11y agoDeep learning has already completely taken over speech recognition, face recognition, and object recognition. Going forward deep learning is the technology that will solve sensing and perception for machines. You're going to want sensing and perception in your phone and every computer you use, but even more than that: every drone, every self-driving car, every kind of robot is going to need multiple deep learning chips. Each of those chips is going to need vastly more FLOPS and more memory bandwidth than the biggest CPUs and GPUs of today. Beyond sensing and perception, I believe that deep learning will also be the technology that solves planning, natural language, reasoning, and creativity for machines. This is much more speculative, but you can see the beginnings of planning in the DeepMind Atari work, the beginnings of natural language processing and reasoning in machine translation and various question answering systems, and the beginnings of creativity in Deep Dream and other generative models people have done. Of course once all those pieces are solved then AI is solved. The market for a technology that solves AI is practically unlimited. Intel is constantly searching for the next big application that will require more processing power so people need to buy faster CPUs and they can justify spending $X billion on their next fab. In recent years they've been struggling to find it. Deep learning is it. The appetite for FLOPS and memory bandwidth is unbounded for the foreseeable future. Unfortunately for Intel, CPUs are weak on both compared to GPUs. Maybe Xeon Phi can morph into a deep learning system?
- lucidrains 11y agoI am similarly very bullish on this field, even for what I once thought were incredible claims that it can eventually understand the semantics behind documents, do question and answering, and even reason. What blew my mind was this lecture, given by Hinton https://drive.google.com/file/d/0B8i61jl8OE3XdHRCSkV1VFNqTWc/view https://drive.google.com/file/d/0B8i61jl8OE3XdHRCSkV1VFNqTWc... The main idea is that reasoning is just a sequence of 'thought vectors' that can be encoded within a recurrent neural network. Sounded almost outlandish to me until I watched Richard Socher's lectures and started to understand that words can be represented as vectors, and that these vectors can be then encoded into new vectors of even higher representation. 'Thought vectors' may not be so outlandish after all.
- 11y ago
- raverbashing 11y agoEspecially in the (possibly very lucrative market) of Deep Learning powered devices, like self-driving cars nVidia has already demoed hardware on that area
- Retr0spectrum 11y agoIs deep learning suitable for self-driving cars? I would expect that cars, like other safety critical systems, need to be formally proven to work as intended, which AFAIK can't really be done with neural networks.
- detaro 11y agoI don't think it would be possible to formally prove that e.g. any image recognition system will work correctly given all realistic real world inputs, so that requirement doesn't work out anyways. It's more likely that they are required to perform extensive real world tests + maybe beforehand tests against video material.
- shorodei 11y agoThere are certain subsystems of self driving (visual recognition) that are inherently learning/pattern based.
- seanmcdirmid 11y agoWhere did you get this idea from? Formal proofs even in avionics are very limited, and planes aren't exactly dropping out of the sky. Safety critical systems need a degree of auditing and testing, to be sure, but formal proofs have never been a requirement since they aren't that practical.
- cscurmudgeon 11y agoThis just FUD. Formal proofs are quite prevalent. Look at ACL2 for a start.
- 11y ago
- michaelt 11y agoI'm not an expert in these things, but I was under the impression that when things like this make it to consumers, they've run a big training data set, taken a snapshot of the trained neural network, then it's used as a static snapshot. Then, if outliers/problems are encountered in the real world, rather than training the end user's neural network alone, the outliers are fed back to the centralised system, and later a new snapshot might be sent out. After all, people want their self-driving cars and voice recognition phones to work out of the box. And if there's a misclassification you correct on your phone, you want it to propagate to your tablet, PC and smart fridge :) If things are done that way, I would have thought the market would end up with only a handful of customers - there won't be a need for a deep learning chip in every phone. There might be a world market for maybe five deep learning systems :)
- dave_sullivan 11y agoThe world market is bigger than 5 systems. I'd argue it looks more like relational databases in ~1980. Most people to this day do not know what a relational database is, but they probably use one multiple times a day. Over the years, databases have gotten bigger, faster, more complex, more powerful, easier to use, and more open. Deep learning--and machine learning more generally, because that's what we're really talking about here--is going to go through the same process. NVIDIA is positioning themselves to be the vertically integrated supplier in this market. The really interesting part is that no one is sure yet what kinds of applications these tools will open up -- it is likely to be an enabling platform/toolset in much the same way relational databases, app stores, or IaaSs were. AMD and Intel are woefully behind. Intel because it's not yet a big enough market ($137B market cap compared to NVIDIA's $12B, Intel needs an opportunity to be literally 10x as interesting to be interested). Why AMD is not all over this I'm not sure... If it's interesting for NVIDIA, it should be very interesting to a $1.6B operation like AMD.
- ris 11y ago"Why AMD is not all over this I'm not sure..." NVidia have been very good at encouraging the use of their proprietary CUDA. They also have a slight FLOPS advantage I'm told (where AMD have an integer advantage)
- trsohmers 11y agoShameless self promotion... my startup (http://rexcomputing.com http://rexcomputing.com) is producing a standalone chip capable of 64 GFLOPs/watt double precision (128 GFLOPs/watt single precision), compared to NVIDIA's next generation chips only hitting 20GFLOPs/watt single precision... and that is before you take into account the power waste of the CPU controlling the NVIDIA GPU. Our biggest plus factors compared to a GPU is that we are a fully standalone/independent chip that does not need to have a CPU with your main system memory attached to it. Large machine learning data sets are getting into terabytes in size, and the biggest bottleneck with GPUs is the PCIe link limiting them to 16GB/s and adds a whole lot of additional latency. In our case, we have direct connection to DRAM (We've been looking at DDR4 and HMC). In addition, we have designed the architecture to allow massive scalability with up to 384 GB/s of aggregate chip-to-chip bandwidth... NVIDIA's NVLink is aiming for 80GB/s in the 2018/2019 timeframe and will still need a connected CPU to issue jobs. EDIT: I should also mention that our chip is fully general purpose, but we perform really well when it comes to dense matrix math (Most deep learning), with a 10 to 15x efficiency advantage over GPU. Our real killer app is FFTs, which GPUs do abysmally on, and our current benchmarks are showing a 25x efficiency advantage over the best DSPs and FPGAs built for large constellation FFTs.
- modeless 11y agoSingle precision isn't low enough. You want half precision or maybe even lower. You also want to throw IEEE 754 out the window. Save on power and area: no denormals, no infinities, no NaNs, relaxed precision requirements. It may even be worth looking at exotic things like logarithmic number systems or analog logic (deep learning should tolerate noise extremely well). You're also going to need vast amounts of memory bandwidth, which means on-package memory, and probably specialized caches and compute units for convolution. A truly specialized deep learning chip probably wouldn't be useful for much else, but it would be a monster at deep learning. And the thing about deep learning is it scales really well. If you have a 10x faster machine you're almost certain to set world records on any machine learning benchmark you try.
- trsohmers 11y agoWhile I personally dislike IEEE Float, we decided to remain compliant for our first chip, as that is a checkbox for a lot of businesses that we want to sell into. We are looking at a new variable precision floating point format, called Unum, created by one of our advisers (And HPC industry legend) John Gustafson. Unum would be fantastic for deep learning, asyou would only use the precision actually required, thus bringing the program size and memory bandwidth numbers down and total energy efficiency up at least ~30%-50% over the same system with IEEE float... You can check out a previous HN discussion on it here: (https://news.ycombinator.com/item?id=9943589 https://news.ycombinator.com/item?id=9943589 We still have the option of including a 16 bit (half precision float) packed SIMD mode into our FPUs, which would add a bit of complexity (bringing our efficiency numbers down a bit for the double precision float, which we like to talk about as it is over 10x better than anything out there), but if there is enough customer interest we may decide to include it.
- fpgaminer 11y agoIs the field of machine learning really stable enough to warrant an ASIC? Seems to me a set of new techniques is released every year now. The lead-time on ASICs is ~1 year. So I don't see how a Deep Learning ASIC could keep pace.
- modeless 11y agoNew training techniques are released on a monthly basis, but the basic structure of convolutional neural nets has actually been around since the '90s and hasn't materially changed. I think that's a stable enough target to build an ASIC for. Even if you're unable to implement the year's hottest training technique for some reason, if you're 10x bigger/faster you'll likely still set world records using last year's techniques.
- pilooch 11y agoOpenCL is slowly being supported by some of the main deep learning packages. The annoying part of Nvidia's GPUs are drivers and closed souce CUDA. Their CuDNN extension is a proprietary blob (with bugs) that boost up perfs but remains out of touch from scientists and developers. My hunch is that OpenCL will become the standard for deep learning, thus opening up to a larger set of hardware options.
- varelse 11y agoYes, NVIDIA owns Deep Learning. And they will continue to own it for at least the next 2 years. But I would be remiss not to point out that you can build one of these things yourself for less than half the price at which they're selling it. And while I think one can build a deep learning ASIC, a 10x better ASIC seems like a tough bet to me. I mean on the surface it sounds good, but the devil is in the details here and by the time you've built something flexible enough to run every reasonable variant of a neural network both forwards and backwards, you start making the sort of decisions that make your processor look more and more like a GPU and that magical perf delta drops. Also if you start today, you need to target tomorrow's GPUs, not the $1000 consumer model you can buy on Amazon now. That said, I'm looking forward to Altera's new hardcoded floating point-enhanced FPGAs. Too bad I have no idea how much they cost.
- dharma1 11y agoEven less if you use AMD GPU's -https://www.reddit.com/r/linux/comments/2zgpj8/15000_nvidia_developer_system/cpjjff5 https://www.reddit.com/r/linux/comments/2zgpj8/15000_nvidia_... I wonder why the popular deep learning frameworks are using mainly CUDA instead of OpenCL. Is because of better Linux GPU drivers? Wondering why AMD isn't jumping on deep learning The Altera FP capable FPGA's sound real interesting too. 10 TFLOPS, OpenCL support? http://www.slideshare.net/embeddedvision/a04-altera-singh http://www.slideshare.net/embeddedvision/a04-altera-singh Looks like they're about to be bought by Intel? http://www.electronicsweekly.com/news/business/altera-important-intel-2015-06/ http://www.electronicsweekly.com/news/business/altera-import... Does this mean FPGA co-processors in the future from Intel?
- varelse 11y agoI would love to see AMD jump in the ring, and there's even an OpenCL port of Caffe in progress: https://github.com/BVLC/caffe/pull/2610 https://github.com/BVLC/caffe/pull/2610 But its performance is less than half that of a GTX 980 running CUDA. Still, AMD is silly not to try and improve on this IMO.
- dharma1 11y ago
- sklogic 11y agoPeople are doing deep learning even on mobile GPUs. Even on VC4 on Raspberry Pi.
- Twirrim 11y agoTitans are $1.5k each, so that's $6k down before you even account for the rest of the hardware to run it. Ouch.
- raja 11y agoThe boxes will be $15,000 USD. Lead time is 8-10 weeks.
- semi-extrinsic 11y agoPoster above says the GPUs cost $6k total, and cases that can fit 4GPUs aren't that expensive nor elusive. So I guess the rest of the system is made from unobtainium, or they figure deep learning is the kind of red hot topic that attracts enough suckers with money that this will fly.
- jfb 11y agoOr people for whom the money value of their time exceeds the margin over a custom build that nVidia is asking.
- semi-extrinsic 11y agoI can spec out a system like that in less than 1h, but let's say it takes 2h, ending up with a total hardware cost of ~$9k plus $200 to have it assembled. Are you really valuing your time at more than $2000 an hour?
- jfb 11y agoIf someone's willing to pay, sure! But more realistically, a ML researcher doesn't want to know the differences between Haswell-E and Skylake, or which DIMMs are rated at which speed, or what's the optimal Linux driver situation, still less to know how to fix any of the myriad things that can go wrong. When Caffe starts screwing up, it's really nice to be able to call nVidia and make your overheating Titan their problem.
- mobileexpert 11y agoNVidia should also market this for people who want to do molecular dynamics and other gpu enabled physics sim locally.
- madengr 11y agoI have one of these for electromagnetics sim: http://www.microway.com/product/whisperstation-tesla/ http://www.microway.com/product/whisperstation-tesla/
- sandGorgon 11y agoinstalled standard Ubuntu 14.04 w/ Caffe, Torch, Theano, BIDMach, cuDNN v2, and CUDA 7.0 whoa - are you telling me that the nVidia drivers on Linux are so stable that they are building a commercial deep learning system on top of that. Is this the same thing as normal graphics drivers ?
- wmf 11y agoAFAIK Linux with Nvidia GPUs is used in plenty of supercomputers and Hollywood special effects companies, so the drivers must be stable enough.
- somerandomone 11y agoThe difference is between "given this particular hardware and OS setup, the driver will work correctly, guaranteed." vs. "on your (discontinued) Sony laptop with a strange hardware interface running the beta Slackware release the driver will probably work"
- wtallis 11y agoIt's also the difference between "using our tools implementing our API on our hardware" vs. "trying to figure out the right thing when every component has a slightly different take on the API spec and the applications using the API make mistakes that we have to try to correct for with unreliable heuristics". Developers in the scientific computing world care about correctness a lot more than game developers working under unrealistic deadlines and with no commitment to long-term maintenance.
- sandGorgon 11y agothat doesnt make sense to me. If the driver works and does not crash the operating system (kernel panic FTW), that's good enough at this point. nVidia uses CUDA or OpenGL - so its not quite the question of proprietary API. At this point, I'm not worried about "framerate on my linux box isnt as good as windows".. its more "it works...".
- happycube 11y agoThe Pascal cards are going to be much better, with HBM2 memory and possibly even actual double-precision performance (which isn't a problem for deep learning, but still...)
- ris 11y agoNVidia in particular are very good at selling the future - I'll warn you that much.
- happycube 11y agoTrue 'dat, but combined with the first GPU die shrink in years, there's a decent chance a generational jump is actually coming. Whether the first version will be bug-free is much more questionable...
- Malician 11y agoIf a move from 28nm to 14nm FINFET plus HBM at the same time doesn't drastically increase performance, something is very wrong.
- alricb 11y agoFWIW, the case is a Corsair A540 with hard drive sleds in the two 5.25" bays: http://www.corsair.com/en-us/carbide-series-air-540-high-airflow-atx-cube-case http://www.corsair.com/en-us/carbide-series-air-540-high-air... Makes sense to me, since you want the best airflow possible getting to the cards in a multi-GPU setup, and unlike in conventional cases, the A540 doesn't have a drive cage between the front fans and the video cards.
- sxp 11y agoThe price for a custom build with these specs is ~$8k: http://pcpartpicker.com/p/NP4MNG http://pcpartpicker.com/p/NP4MNG Spending that money on an EC2 GPU instance would be a better use of money unless you really need a local workstation.
- ericjang 11y agoI'd recommend a local custom build over EC2. EC2 GPU instances are virtualized via a Hypervisor, which dramatically reduces performance on multi-GPU networks. and this doesn't take into account the large amounts of disk space needed for the training set.
- sandGorgon 11y agodoes anyone know which hypervisor do they use ? Can I build a local EC2 GPU instance with these GPUs ? I'm quite amazed that they are able to get the drivers, etc working with these GPUs on top of a hypervisor
- UK-AL 11y agoThere are dedicated instances specifically for this purpose.
- sandGorgon 11y agoagreed - but which hypervisor do they use ?
- UK-AL 11y agoNormally xen?
- ericjang 11y agoMy guess is that they use http://www.nvidia.com/object/virtual-gpus.html http://www.nvidia.com/object/virtual-gpus.html
- kfor 11y agoI wonder how Nvidia building their own machines goes over with the many, many third party partners building similar rigs. On the Supercomputing 2014 showroom floor it seemed like half the booths were selling something like this and were covered in Nvidia branding.
- Sanddancer 11y agoThis is a developers platform, in a rather inconvenient form factor for any sort of scale deployment. The partners honestly probably love it because it means they won't be hit with support requests because a driver's acting up, and will just be getting the sales to handle the finished product.
- nextos 11y agoI'm working on probabilistic programming. Hierarchical models are very close to deep learning. PyMC3 has a Theano backend, so this kind of setup is very exciting. Anyone else with the same thought/interests?
- bagels 11y agoWhy would I buy this, vs. renting a cluster of ec2 gpu nodes?
- jfb 11y agoData locality and performance.
- modeless 11y agoThese GPUs are more than twice as fast as Amazon's GPUs. They also have 4x the memory and there is probably more GPU to GPU bandwidth as well. Doing deep learning on clusters is not practical yet due to bandwidth issues. If you want to train state-of-the-art models you need the biggest single machine you can get, and this is it.
- Sanddancer 11y agoLatency, support, specifications. The EC2 GPU nodes have less, slower, memory, graphics cards with half the performance and a third of the memory, and less drive space. Additionally, said compute resources are not next to you, which means certain things, like deep learning combined with AR are not possible. If you're just doing number crunching, you may be fine with EC2 instances, but any sort of realtime development, you probably want a supported platform right next to you.
- z3t4 11y agoWhy not get a server tower case and motherboard while you're at it? Supermicro has some good ones.
- Sanddancer 11y agoA machine like this, you're buying the support, with the hardware as an add-on. This isn't made for the deployments, it's made for the development and debugging, where being able to call nvidia at any hour and get a decent engineer is worth it.
- seiji 11y agonewegg has been selling quad 12GB Titan X GPU combo packs for a while. single-click add-to-cart for 18 components: http://www.newegg.com/Product/ComboBundleDetails.aspx?ItemList=Combo.2349536 http://www.newegg.com/Product/ComboBundleDetails.aspx?ItemLi...
- bobjordan 11y agoWe built our own quad-titan devbox a few months ago, same general components as this, except used Core i7-5960X and threw in a few 1TB samsung SSD's in Raid, came in just at $9,000 USD hardware cost, which I think Nvidia was charging about $15,000. Still, I'm sure they aren't making a ton of money, and you get hardware guarantee with configuration (but config wasn't so bad..).
- viklas 11y agoAgree - we crunched the numbers and came up with the same figures to do-it-yourself (~USD$9K). Although, I remember from the day this was announced (few months back) that Nvidia were loud-and-proud that they weren't going to make money from this. Each box was hand built and tested, so not deemed to be a large-scale device - they recognized that it's a niche market. $15K is probably OK(ish) if you figure in your own time for the DIY build...probably a few days. Plus you get some vendor support, warranty on the whole package, certified working stack, future test-bed for CUDA updgrades (will work first), etc as you say. In wild agreement. Save maybe 30% doing a custom build...so they aren't adding a huge mark-up, as they would for a gaming machine. Apparently...someone at Nvidia is looking a bit further into the future than just the short-term revenue.
- erikj 11y agoIt looks like the NeXTcube: https://upload.wikimedia.org/wikipedia/commons/2/27/NeXTcube.jpg https://upload.wikimedia.org/wikipedia/commons/2/27/NeXTcube...