7 ms·
Caffe2 adds 16 bit floating point training support on the NVIDIA Volta platform
- mappu 9y agoAre numbers available for FP16 on P100 or FP32 on V100? It would make for a more direct comparison. EDIT: Nvidia's advertised TFLOPS are: FP16 FP32 FP64 V100 30 15 8.5 P100 21.2 10.6 5.3 K40 4.29 4.29 1.43
- Tom1971 9y agoYour table doesn't include the new "tensor core" TFLOPS of V100. That's a core that does 4x4 FP16 matrix multiplication + 4x4 FP32 accumulation in one go. That's where V100 gets its boost, up to 120 TFLOPS.
- bsprings 9y agoTensor Cores: 120 TFLOP/s mixed-precision (peak). Typo in your table: V100 FP64 is 7.5 TFLOP/s.
- maldeh 9y agoFrom first looks, there is little doubt that that NVIDIA's Volta architecture is a monster and will revolutionize the AI and HPC market. But the article seems to avoid quantifying how 16-point FP operations are beneficial against 32- or 64-bit FP operations in real-world usecases, or how the Caffe2 / NVIDIA architecture provides any significant boost to FP16 in particular, especially apropos to images (or why FP16 is better for image data in general). I'm interested more in understanding why Caffe2 would outperform Theano, Tensorflow, MXNet, etc. once Volta chipsets are generally available, beyond early pre-release optimization, particularly when most of the front-runners are already leveraging / taking into account NCCL, CuDNN, NVLink, etc. When the burden of adding support for new NVIDIA primitives is so low, what gives Caffe2 an advantage beyond an ephemeral "we were partners with NVIDIA first" one-up that would last for a couple of months at most? (Apologies in advance if this post sounds overly negative, but I am constantly evaluating the current crop of frameworks for the trade-offs they enforce on the problem space, and a definitive answer would be very helpful.)
- deepnotderp 9y agoCaffe2 seems to be engineered for efficiency as opposed to flexibility. For example, Caffe2 is apparently Facebook's go-to choice for mobile edge device deployment for applications such as style transfer.
- throwaway87213 9y agoEfficient how ? In terms of memory management ?
- eleitl 9y agoTPUs are smaller in terms of Si real estate and burn less power. Also, they can be made much faster.
- sja 9y agoI mean, it uses a static graph definition instead of dynamic graphs (allows for deployment optimizations), but is there anything Caffe2-specific that would give it a leg up over Theano/TensorFlow/whatever other static graph library?
- raverbashing 9y ago16fp is faster, and the increased precision is not needed in most cases (In fact there are even some experiments with a very low number of bits - can't find the link now though)
- buildbot 9y ago16 bit FP operations seem to be more than enough for most networks - in fact for inference, google claims in their TPU paper that most network parameters can be quantized down to 8 integer bits! There is also some work on pure binary networks, xnornet and binarynet. Low precision doesn't seem to effect these networks as long as they are trained with binary weights in the first place.
- user5994461 9y ago>>> I'm interested more in understanding why Caffe2 would outperform [...] A vendor-made benchmark where the vendor is outperforming other vendors. Let's not jump to conclusion on what's really fastest.
- jakebasile 9y agoThe GPU looks like a monster, and I am sure it'll deliver more power to AI applications, but what I really want to know is when I can put one in my gaming PC. I think it's great that what started as a specialist gaming device is now being used in industry for Big Things. The development cost that Nvidia et al. have invested in new designs has undoubtedly been financed in (large) part by the gaming community. Now income and advancements for both sectors feed into the other and gamers like me are reaping the benefits with reduced price:performance across the range.
- pjmlp 9y agoMeanwhile I still remember being seated at the Games Development Conference talk in 2009 where Intel tried to convince us how Larrabee would change the world of computing. Still waiting for them to produce anything worthwhile buying instead of AMD and NVidia GPUs.
- dr_zoidberg 9y agoActually, it's now called Xeon Phi: https://en.wikipedia.org/wiki/Xeon_Phi https://en.wikipedia.org/wiki/Xeon_Phi
- pjmlp 9y agoI know, they are pretty hard to get and as such largely ignored by HPC and gamming communities.
- dr_zoidberg 9y agoTheir main sales pitch is "given that it's a bunch of x86 put together, you don't have to port your code to get Massive Paralellization by Intel (TM)". Some supercomputers use them, but it's true: I've never seen a worthwhile comparisson between that and a strong (or equivalent) GPGPU. There are some very interesting talks by Intels experts on how to use these guys with AVX and AVX512 (or however its called now) and the multiplicative improvement both multicore and AVX give when used together (the sales line goes along like this "10x for multicore, and on top of that 4x for AVX, voilá 40x speedup!"). I don't work in an HPC environment, so I don't really know if they stand up in reality to Intels claims.
- throwaway87213 9y agoI can see why they launched 1080 Ti early. Had I not seen this I'd definitely not be waiting for Volta. How are things in the red camp ? There was some HIP thing where Fiji was as good as Pascal.
- eleitl 9y agoAny idea on the timeline for consumer Volta?
- dogma1138 9y agoEarly 2018 depending on GDDR6 supply. Seems like Hynix will start mass production on in late 2017.
- distances 9y agoPhoronix said "Volta desktop graphics cards are expected to succeed Pascal in late 2017 or early 2018" in https://phoronix.com/scan.php?page=news_item&px=NVIDIA-Volta-V100 https://phoronix.com/scan.php?page=news_item&px=NVIDIA-Volta....
- v4n4d1s 9y agoGerman news portal heise.de said Q2 2018.
- DrNuke 9y agoThe GTX 1070 for laptops is destined to be the most interesting price/performance opportunity for a while imho, you can even undervolt the CPU -0.100 to -0.150 to reduce overheating.
- tanderson92 9y agoI was told Q4 2017 for non-graphics consumer (what kind of consumer are you referring to), but that is earlier than most others are saying. https://news.ycombinator.com/item?id=14310633 https://news.ycombinator.com/item?id=14310633
- dharma1 9y agoStill no good CuDNN equivalent from AMD for machine learning that would make their cards competitive for that use case
- BugsJustFindMe 9y agoI understand the pedigree of these cards, but at what point do we stop calling them GPUs and start calling them something else? MPU? TPU? I don't know, but isn't it a little bit weird to keep using the word "graphics" for something that is being made more and more specifically for other things?
- tome 9y agoNvidia rep: "Our latest MPU will train your models lots faster". Potential customer: "What's an MPU?" N: "Oh, it's the same thing we used to call a GPU" P: "Why did you change the name?" N: "Because it's a little bit weird to keep using the word 'graphics' for something that is being made more and more specifically for other things" P (puzzled, and marginally less likely to make a purchase): "Oh right"
- joelthelion 9y agoI don't see the point of adding to the confusion. They've been called GPUs for ever, everybody agrees on the term, why not continue to use it?
- DaiPlusPlus 9y agoI like the term "GPGPU" (General-Purpose Graphics Processing Unit) as it's at least explicit about its general-usefulness.
- function_seven 9y agoEh, that seems like it’s intended for general purpose graphics (i.e. 2D business applications vs “special purpose” 3D, gaming, VR, etc.)
- roel_v 9y agoNo. For many years now, gpgpu has meant 'using gpu's for non-graphics calculations'.
- iamNumber4 9y agoI'm a long time Nvidia user. However, the recent article headlines have been a word salad. The new tesla volta super flip flop at 1.21 gigawatts blah blah blah. Just saying. Also Kudos to Nvidia for the buzzword/made up word creation for their products.
- Symmetry 9y agoSurely getting people to refer to their vector lanes as "cores" was their biggest piece of marketing magic. Yes, they're more flexible than SIMD cores but not more so than an OoO core's execution ports.
- DocSavage 9y agoThe AnandTech article on the Volta has a lot more information on the new architecture: http://www.anandtech.com/show/11367/nvidia-volta-unveiled-gv100-gpu-and-tesla-v100-accelerator-announced http://www.anandtech.com/show/11367/nvidia-volta-unveiled-gv... It's interesting the speed up isn't more pronounced between Volta and Pascal considering the Tensor cores on paper give you about 6x the MFlops. The price differential looks large. From AnandTech: "By the numbers, Tesla V100 is slated to provide 15 TFLOPS of FP32 performance, 30 TFLOPS FP16, 7.5 TFLOPS FP64, and a whopping 120 TFLOPS of dedicated Tensor operations. With a peak clockspeed of 1455MHz, this marks a 42% increase in theoretical FLOPS for the CUDA cores at all size. Whereas coming from Pascal, for Tensor operations the gains will be closer to 6-12x, depending on the operation precision."
- trueSlav 9y agoNo free lunch and all that... 40% seems quite nice if you think about the transistor count (15b to 21b).
- Aliyekta 9y agocudnn 7?
- TekMol 9y agoCan the Volta architecture be used to run WebGL in a Browser?