12 ms·
Volta: Advanced Data Center GPU
- bmiranda 9y ago815 mm^2 die size! That's at the reticle limit of TSMC, a truly absurd chip.
- kurthr 9y agoI agree... there's not much more they can do to scale since off die is still slow. Unless they stitch across the exposure boundary! However, they have been at the reticle limit since they were in 28nm. GM200 (980 Ti and Titan X) was 601 mm^2 at TSMC... the maximum possible at the time.
- tostitos1979 9y agoI've seen some huge mainframe die back in the day. What is reticle limit exactly? Thanks for educating a SW guy :)
- Terribledactyl 9y agoPart of the chipmaking process is burning layers into wafers covered in photoreceptive material. Photomasks/reticles used to cover entire wafers making many units at once, but now the processes are so small they have to compress the image (4-10 times is typical), burn a couple units, step over repeat on the same wafer. This GPU is so large, they can only fit 1 of them in a single burn step.
- deepnotderp 9y ago193i immersion steppers,a la ASML have 32x26 as the reticle limit
- tbrownaw 9y agoIt's something along the lines of the film size for the super fancy camera they use in one of the steps. (The silicon wafer would be the equivalent of the entire roll of film.)
- gigatexal 9y agoThese tensor cores sound exotic: "Each Tensor Core performs 64 floating point FMA mixed-precision operations per clock (FP16 multiply and FP32 accumulate) and 8 Tensor Cores in an SM perform a total of 1024 floating point operations per clock. This is a dramatic 8X increase in throughput for deep learning applications per SM compared to Pascal GP100 using standard FP32 operations, resulting in a total 12X increase in throughput for the Volta V100 GPU compared to the Pascal P100 GPU. Tensor Cores operate on FP16 input data with FP32 accumulation. The FP16 multiply results in a full precision result that is accumulated in FP32 operations with the other products in a given dot product for a 4x4x4 matrix multiply," Curious to see how the ML groups and others take to this. Certainly ML and other GPGPU usage has helped Nvidia climb in value. I wonder if Nvidia saw the writing on the wall so to speak with Google releasing their specialty hardware called the Tensor hardware that Nvidia decided to use it in their branding as well.
- bmiranda 9y agoGoogle's hardware is for inference, not training.
- gigatexal 9y agothanks for clarifying.
- josephpmay 9y agoVolta is for both inferencing and training, but has an emphasis on inferencing
- JustFinishedBSG 9y agoIt doesn't matter, operations are the same in forward and backward mode. "Made for inference" just means "too slow for training" if you are pessimistic or "optimized for power efficiency" if you are optimistic. Otherwise training and inference are basically the same
- david-gpu 9y ago
- hesdeadjim 9y agoI find it so cool that technology created to make games like Quake look pretty has ended up becoming a core foundation of high performance computing and AI.
- pgodzin 9y agoMatrix multiplication is important for graphics and important for finding the weights of a neural network
- hesdeadjim 9y agoYep, hard to imagine though that the original creators of the Nvidia TNT or Voodoo had any idea that GPUs would become fully programmable computing hardware used for non-graphical applications.
- abritinthebay 9y agoI don't think they'd have been super surprised. Just pleasantly happy. AI Accelerators have been a thing for decades - DSPs were used as neural network accelerators in the early 90s - and Cell processors were a thing by 2001. GPUs just became vastly more accessible to general purpose program in the last decade. People were doing it back in the 90s but it was seriously hard. We finally hit a tipping point where it's just kinda hard.
- cr0sh 9y agoThere were also the various custom "systolic array" processor designs in the 1980s (the ALVINN vehicle, and earlier projects which led to it, used these for early neural-network based self-driving experiments).
- scott_s 9y agoI remember back in 2004 when I heard a fellow grad student was working on using GPUs as a co-processor for scientific computing, I though "Wow, that's esoteric and niche."
- tobyhinloopen 9y agoTime to play some games on it
- mtgx 9y agoI have a feeling eventually Nvidia will, like Intel, de-prioritize the consumer market in favor of the much more profitable server/machine learning market.
- hatsunearu 9y agoPerhaps, but the desktop gaming market is still growing and is a huge part of NVIDIA's income.
- kanwisher 9y agoI'ved thought that but the per unit volume is huge. Every game console, phone, tablet, PC needs a GPU. Even low-end devices are expected to run games. Thats billions of units, albeit at lower margins
- zokier 9y agoAnd practically every game console, phone, tablet and vast majority of PCs are running integrated GPUs. Integrated GPUs that are not nvidia. Unless NV gets into the licensing market, the growth potential for them seems somewhat limited.
- mattnewton 9y agoWow, this is just Nvidia running laps around themselves at this point. Xenon Phi still not competitive, AMD focused on the consumer space, looks like the future of training hardware (and maybe even inferencing) belongs to Nvidia. (Disclosure: I am and have been long Nvidia since I found out cudnn existed and how far ahead it was)
- coldtea 9y ago>Xenon Phi still not competitive, AMD focused on the consumer space, looks like the future of training hardware (and maybe even inferencing) belongs to Nvidia. Assuming there's a big future to training hardware and inferencing. Many of those "new paradigms" / "silver bullet technologies" have come and gone in the last decades.
- mattnewton 9y agoThat's true, but there is reason to believe this time is different™, with killer applications in medical image understanding, natural language understanding, and self driving cars, all of which could drive demand of these chips by themselves. It is possible we will discover new dominant architectures that don't use this hardware well but I am putting my money on us coming up with even more applications that do use this hardware well.
- deepnotderp 9y agoThere's something coming for them: deep learning processors. I'm biased, since I'm part of one, but there's little to no modification of the software stack necessary, so it's a credible threat to nvidia.
- mattnewton 9y agoI hope so, if only because it keeps them running at this pace! Kudos for charging the 800lb gorilla head on.
- p1esk 9y agoWhat do you think about them open sourcing DLA of Xavier?
- randyrand 9y agoWhat are the silver boxes that line both sides of the card? Huge Capacitors?
- smitty1110 9y agoFerrite chokes, part of the power delivery system.
- randyrand 9y agoWhy are they needed?
- 6d6b73 9y agoTo get rid of electrical noise.
- flamedoge 9y agoim assuming chip draws yuuge power
- Keyframe 9y agoYou're not wrong. 300W, holy shit.
- qb45 9y agoFor the same reason as around any other CPU or GPU and lots and lots of other chips: buck converter, i.e. 12V 20A DC in, 1.2V 200A DC out.
- hatsunearu 9y agoIt's part of a step down voltage regulator called a buck converter. The buck converter works by putting a pulse of energy into the inductor and stretching it out to lower the voltage. This creates the core voltage.
- hatsunearu 9y agoInductor, not chokes. Part of the buck converter to create Vcore.
- caenorst 9y agoDid they communicate any release date and price during the show ?
- abhshkdz 9y agoDGX-1 with Volta — $149k, Q3; DGX Home Station with Volta — $69k, Q3
- tanderson92 9y agoAny information about when this architecture will make it onto Tesla or Quadro products available to "mass" market?
- abhshkdz 9y agoI think Jensen mentioned this would be available with OEMs Q4 onwards
- arca_vorago 9y agoMore great hardware being stuck behind proprietary CUDA when OpenCL is the thing they should be helping with. Once again proprietary lock in that will result in inflexibility and digital blow-back in the long run. Yes I understand OpenCL has some issues and CUDA tends to be a bit easier and less buggy, but that doesn't detract from the principles of my statement.
- MichaelBurge 9y agoNobody else is even bothering to compete, so standards don't really matter. Let them do their job: I'd rather have faster GPUs.
- tanderson92 9y agoStandards matter if you care about software and hardware freedom.
- Symmetry 9y agoI wonder if the individual lane PCs will pave the way for implementing some of Andy Glew's ideas for increased lane utilization in future revisions? http://parlab.eecs.berkeley.edu/sites/all/parlab/files/20090827-glew-vector.pdf http://parlab.eecs.berkeley.edu/sites/all/parlab/files/20090...
- 1024core 9y agoFTA: "GV100 supports up to 6 NVLink links at 25 GB/s for a total of 300 GB/s." The math doesn't add up.
- lowglow 9y agoI'm really happy our startup didn't go all in on Tesla (Pascal architecture) yet. These look amazing.
- mattnewton 9y agoI feel like every time I buy cards, Nividia announces the successor with absurd improvements.
- dom0 9y agoOTOH improvements in the mainstream segments seem to go slower: Mainstream cards are about twice as fast now as they were five years ago.
- lowglow 9y agoYeah, I just sprung for a Titan Xp -- waiting for it to become obsolete next month.
- mattnewton 9y agoWell, close to already if you are looking at $$/comprable performance, with the 1080ti
- mastazi 9y agoThe Titan Xp (with lowercase p, as opposed to the Titan XP) came out after the 1080 Ti so I'm sure GP took the latter into consideration before making a decision...
- lowglow 9y agoYep. I'm not sure it was worth the extra $$ for the extra specs just yet. We'll see when we SLI it. The issue though is no memory sharing with the GTX/Titan line. If that were the case, I probably just would have sprung for two 1080Tis out the gate. Definitely loving the eight 1080Tis they just fit in here though: http://www.velocitymicro.com/promagix-g480-high-performance-computing.php http://www.velocitymicro.com/promagix-g480-high-performance-...
- gwbas1c 9y agoHow long until Tesla sues for trademark infringement? "from detecting lanes on the road to teaching autonomous cars to drive" makes it sound like there is an awful lot of overlap in product function.
- cr0sh 9y agoI doubt anything like that would happen. While Tesla Motors was founded prior to the creation of the Tesla GPU architecture, there's not really any overlap - in fact, I wouldn't be surprised if Tesla Motors wasn't using something like this from NVidia: http://www.nvidia.com/object/drive-px.html http://www.nvidia.com/object/drive-px.html As far as any overlap software-wise is concerned, while it isn't super clear what Tesla Motors is doing for their self-driving systems, based on what I've seen it seems like they are using only "basic" lane-detection and identification along with some other algorithmic vision-based systems. I'm not saying that's everything they are doing, just what I have seen released publicly on their vehicle platform. NVidia, on the other hand, has been experimenting with using neural networks (deep learning CNNs specifically) to drive vehicles using only camera information: https://arxiv.org/abs/1604.07316 https://arxiv.org/abs/1604.07316 This is actually a fun CNN to implement - I (and many others) implemented variations of it in the first term on Udacity's Self-Driving Car Engineer Nanodegree. We weren't told to do it this way, but I chose to do so after reviewing the various literature, plus it seemed like a challenge (and it was for me). Udacity supplied a simulator: https://github.com/udacity/self-driving-car-sim https://github.com/udacity/self-driving-car-sim ...and we wrote code in Python (Tensorflow and Keras) to train and drive the virtual car. For my part, I had set up my home workstation with CUDA so that Tensorflow would utilize my GPU (a lowly GTX 750 TI SC - though it seems like it might have a similar GPU capability as NVidia's Drive-PX system, based on what I've researched - a Mini-ITX mobo, a PCI-E slot riser, and a GTX 750 would make a decent low-end deep-learning platform for self-driving vehicle experiments, and cost a fraction of what the Drive-PX sells for).
- sargun 9y agoTesla Motors uses Tegra chips to power their console. So, nVidia is probably okay.
- 9y ago
- deleted 9y ago[deleted]
- braindead_in 9y agoSo when are the new AWS instances are coming?
- arnon 9y agoThis is odd for NVIDIA. They usually push out revised versions in the second year, not change the entire architecture to the new one. Feels like they're feeling AMD breathing down their necks with their VEGA architecture, which should be very interesting. AMD have also stepped up their game with ROCm which might take a chunk out of CUDA.
- Robadob 9y agoAs I recall, Volta (3d memory) has been delayed multiple times due to supply and this is only a very limited release of their highest end hardware for deep learning all pegged for Q3/Q4 release. A field where they haven't really any competition. Can't imagine we will be seeing any Volta GeForce cards released till next year.
- dogma1138 9y agoVolta GeForce will come early 2018 likely with GDDR6 at this point.
- Athas 9y agoDoes this architecture improve on 64-bit integer performance? Have any of the GPU manufacturers said anything about that? At some point it becomes a necessity for address calculations on large arrays.
- sipherhex 9y ago"With independent, parallel integer and floating point datapaths, the Volta SM is also much more efficient on workloads with a mix of computation and addressing calculations" https://devblogs.nvidia.com/parallelforall/inside-volta/ https://devblogs.nvidia.com/parallelforall/inside-volta/ Under "New SM" in "Key Features" section
- jabl 9y agoBut if you read the article it seems the integer units are int32, so not capable of 64-bit computations.
- grondilu 9y agoI was wondering if this will be used in supercomputers. Apparently yes: > Summit is a supercomputer being developed by IBM for use at Oak Ridge National Laboratory.[1][2][3] The system will be powered by IBM's POWER9 CPUs and Nvidia Volta GPUs. https://en.wikipedia.org/wiki/Summit_(supercomputer) https://en.wikipedia.org/wiki/Summit_(supercomputer) Summit is supposed to be finished in 2017, though. I'm quite surprised this is possible since the Volta architecture has only just now been announced.
- Scaevolus 9y agoThe Summit contract was signed in November 2014: http://www.anandtech.com/show/8727/nvidia-ibm-supercomputers http://www.anandtech.com/show/8727/nvidia-ibm-supercomputers Supercomputers have very long planning and development cycles. So do GPUs and CPUs. The contract specified chips that didn't yet exist (Volta and POWER9) as much more than codenames on a roadmap.
- Etheryte 9y agoInteresting to note that Nvidia's stock rose about 18% (!, 102.94USD on May 9, 121.29USD on May 10) in a single day after this announcement. I expected the market to react, but this seems disproportionate.
- virtuallynathan 9y agoThey announced this the day after earnings, earnings caused the jump, this compounded (maybe).
- boulos 9y agoMy favorite outcome of Volta is that it's the first GPU they've produced that actually can claim this SIMT thing due to its separate program counters (we had a spirited debate about whether or not just doing masking but presenting the programming model meant the chip was SIMT or just that CUDA was but GPUs weren't).