7 ms·
GradientFlow: Training ImageNet in 1.5 Minutes on 512 GPUs
- gambler 8y ago512 GPUs on 56 Gbps network? I'd rather see researchers exploring potentially more efficient alternatives to traditional neural nets, like XOR Nets, or different architectures like ngForests or probabilistic and logistic circuits, or maybe listen to Vapnik and invest into general statistical learning efficiency.
- WrtCdEvrydy 8y agoRemember this is the regular flow of hardware. First we built awesome high power single core chips, then multi core chips and continued improving performance per dollar. A single GPU in 10 years might be able to smoke these 512 GPUs.
- gambler 8y ago>A single GPU in 10 years might be able to smoke these 512 GPUs. Firstly, it will not, because there are physical limitations to silicon-based computing and I don't think we will get a different kind in the next 10-20 years. Secondly, unlike commodity programming, AI is an area where improving algorithm/learning efficiency is infinitely more valuable than figuring out how to throw more hardware at the problem. For starters, any algorithmic improvement makes all further research and experimentation easier for everybody as opposed to a handful of agents who have access to ridiculously powerful hardware. I can list many other reasons, like global energy consumption and the need for learning in embedded devices. I find it highly annoying that deep learning enthusiasts immediately turn into uber-skeptics whenever the conversation touches on other ML approaches. Makes me wonder how difficult it is to get funding for fundamentally new AI research in this climate.
- WrtCdEvrydy 8y agoFor reference, 10 years ago, graphics cards were on a 65nm platform while today, we use 14nm graphics cards. The physical limitations may be reached, but we'll just throw more cores unto the PCB to compensate.
- JustFinishedBSG 8y agoWe can't even do that, the biggest GPUs today are at the reticle limits of foundries> They literally can't be made bigger. Of course we can still go the chiplet way but...
- CyberDildonics 8y agoBut what?
- frankchn 8y ago... but power requirements would put a limit to that as well. The RTX 2080 Ti already has a 250W TDP. Putting a couple of those on a single card and you are looking at >1000 W. Cooling and power becomes very hard as we are effectively trying to run and cool a space heater at the same time.
- MasterScrat 8y ago> AI is an area where improving algorithm/learning efficiency is infinitely more valuable than figuring out how to throw more hardware at the problem Improving algorithm/learning efficiency gets much easier when you can iterate faster.
- ousta 8y agowelcome to machine learning 101.
- ribalda 8y agoI have realised that the age of a AI is just a movement of big corporations towards a profitable monopoly. They are teaching us how to solve problems with hardware instead with algorithms, and we, as individuals, will never have access to their computing power.
- nl 8y agoThis is completely wrong. I work in the field, and the way it works is like this, in this order: 1) Someone works out how to do something 2) Someone works out how to improve the accuracy 3) The accuracy maxes out 4) People improve training efficiency. We see this over and over again. Take a look at the FastAI results on ImageNet training speed for example.
- fromthestart 8y agoYou can easily build and train commercial grade neural nets on consumer hardware. Read about, say, a state of the art image recognition net on arxiv, pull an implementation for python in tensor flow or cafe or pytorch or what have you from GitHub (tons of open source), and with nothing more than just a cursory understanding of what you're doing and some basic programming skills, you can run scripts to train and evaluate functional neural networks on your own data set. I've fully trained numerous modern architectures on a single 1080TI to perform image recognition in a matter of days. All of this is within reach of the average developer. If there's any monopoly, it's over training data, which Google and Amazon happen to specialize in. But even large datasets exist as open source. From what I can tell, machine learning, the precursor to AI, is here, and both knowlege and implementation are fully accessable to the general populace.
- IshKebab 8y agoI thought we were optimising for cost now?
- thro_awayz_days 8y agoSilicon is inferior to chemical energy. Human is upwards of thousands of orders of magnitude more efficient than today's best GPU's. However, speed != efficiency. Classifying imageNet with a human would take 1000 hours.
- CyberDildonics 8y agoThat's interesting, do you have a link to the paper this is from?
- the8472 8y agoOn the other hand GPUs are made for computing and will crunch those numbers for you until they die. If you want a human to classify things you have to consider the lifetime cost of making said human and keeping it entertained. They also need idle periods every day, during which they don't even shut down! You can't power them with PV cells either, instead they rely on carbohydrates produced via a horribly inefficient chemical photosynthesis process. And if you intend to let your human classifier run for 8 hours a day you better buy at least three of those for error correction. And I must say this comparison is still quite lenient towards the humans since we're not even comparing them to purpose made silicon entities but generalists.
- trhway 8y agoMatrix begs to difer.
- jjoonathan 8y agoHow does efficiency vary with clock speed? Can you "just" underclock chips to get high efficiency?
- heavenlyblue 8y ago>> Silicon is inferior to chemical energy. I assume if I power the silicon with batteries it's going to stop being inferior?
- RhysU 8y ago> Classifying imageNet with a human would take 1000 hours. You mean it would take 60,000 people one minute? Doable.
- frankchn 8y agoThe authors trained ResNet-50 in 7.3 minutes at 75.3% accuracy. As a comparison, a Google TPUv3 pod with 1024 chips got to 75.2% accuracy with ResNet-50 in 1.8 minutes, and 76.2% accuracy in 2.2 minutes with an optimizer change and distributed batch normalization [1]. [1]: https://arxiv.org/abs/1811.06992 https://arxiv.org/abs/1811.06992
- gpm 8y agoA TPUv3 pod is ~107 petaflops (Googles number from your paper). 512 Volta GPUs is ~64 petaflops (Nvidias number from [1]). v3 pods don't seem to be publicly available. A 256 chip 11.5 petaflop v2 pod is $384 per hour, $3.366 million per year. [2] Meanwhile Google Cloud Volta GPU prices (which are probably inflated over building your own cluster, but are hopefully close enough to a reasonable ballpark) are $1.736 per hour, would be $7.791 million per year for 512. Unless Google GPU prices are really inflated, clusters are legitimately substantially cheaper than cloud GPUs, or these researchers did a poor job, it seems like this is a good advertisement for TPUs. [1] https://www.nvidia.com/en-us/data-center/volta-gpu-architecture/ https://www.nvidia.com/en-us/data-center/volta-gpu-architect... [2] Pod availability / performance / pricing information here: https://cloud.google.com/tpu/ https://cloud.google.com/tpu/ [3] GPU pricing info: https://cloud.google.com/gpu/ https://cloud.google.com/gpu/
- bitL 8y agoOh well, this is the death of democratic AI and an end of independent researchers :-( There goes any hope of a single Titan RTX producing meaningful commercial models.
- ummonk 8y agoIs there some scaling limit that prevents people from doing the same in several hours with a couple of GPUs?
- MasterScrat 8y agoIf each experiment takes you literally 400 times longer than for a Google researcher, your chances of figuring out anything new drops dramatically. I was at an ML conference last year and I asked the panel: if I want to go forward in ML, should I rather do a PHD or work in the industry? a professor (!) answered that I should work in the industry as most research groups don't have the funding to be competitive.
- p1esk 8y agoIf independent researchers don’t like that they should come up with a way to train resnet in 1 minute on a single GPU. Unless you think the ultimate best method to train neural nets has been discovered? Somehow 30 years ago researchers managed to invent cnns and rnns using hardware million times slower than what you can buy today for a few thousand bucks.
- RhysU 8y agoWhy? What deep problems have been solved? How will this make our children better?
- nl 8y agoFor all those complaining about the cost: FastAI trained RestNet-50 to 93% accuracy in 18 minutes for $48[1] using the same code which can be run on your own GPU machine. If you want to do it cheaper and faster, you can do the same for in 9 minutes for $12 on Googles (publicaly available) TPUv2s. This isn't a monopolization of AI, it is the opposite. [1] https://dawn.cs.stanford.edu/benchmark/ https://dawn.cs.stanford.edu/benchmark/