8 ms·
It takes a lot of energy for machines to learn
- ALittleLight 6y agoThe article points out that the energy to train BERT is comparable to a flight across America. There are hundreds of thousands of flights every day - why should we be concerned about the equivalent of one extra? Especially given that BERT improves ~10% of English language Google search results, it seems like we're getting a lot in return for relatively small energy use. On top of that, Google buys 100% of their energy usage from Green sources. I think it's great to talk about other methods of training or architectures that don't require so many parameters. The point about how BERT consumes vastly more text than humans do when they learn to read is interesting. But trying to phrase this like an environmental issue just seems disingenuous and misleading.
- alisonkisk 6y agoGoogle buying "green" isn't a justification because it fails the categorical imperative. It's impossible for everyone to buy "green" energy.
- toomuchtodo 6y agoAs always, there’s nuance. You can schedule workloads, machine learning/training included, where the power is cleanest (low carbon). Google relies on electricitymap.org/Tomorrow to do this. https://blog.google/inside-google/infrastructure/data-centers-work-harder-sun-shines-wind-blows https://blog.google/inside-google/infrastructure/data-center... The more folks who elect to pick where to compute based on low electrical carbon intensity, the faster the grid turns over to clean generation. You must vote with your fiat. I encourage technologists to include this consideration in their workload scheduling requirements. Renewables are almost always cheaper than fossil generation as well.
- curiousllama 6y agoThe categorical imperative doesn't require everyone to be practically able to do the thing. Is it not moral to feed a starving man because people in China aren't able to feed him?
- rictic 6y agoThere's more than enough solar and wind, it's just a matter of how much we as a society want it. The green energy revolution has barely begun, and we're still climbing the learning curve. More demand means a faster climb means exponentially more green energy sooner.
- himinlomax 6y agoBuying more green energy creates demand for more green energy, which in turn creates more supply (wind, solar farms ...) Meanwhile, the aforementioned transatlantic flight cannot get greener, whether Bert is trained or not.
- adamsea 6y agoNot sure why folks are defensive? The article is mainly informational. TFA: "What does this mean for the future of AI research? Things may not be as bleak as they look. The cost of training might come down as more efficient training methods are invented. Similarly, while data center energy use was predicted to explode in recent years, this has not happened due to improvements in data center efficiency, more efficient hardware and cooling." Also. One flight in and of itself isn't an environemntal problem; thousands of flights are. Training this one instance of a model isn't an environmental problem; I would be curious to see some educated guesses about the number of models being trained over the next ten years. Not an expert but I know that use of ML is exploding - lots of new use-cases and thus lots of new models. So it makes sense to me think about this stuff. To
- ALittleLight 6y agoThe article suggests that the energy usage of language models is a problem. I don't think energy usage is a problem. I'm not sure how you interpret this as being defensive. There are hundreds of thousands of flights per day. Adding an additional flight, or even an additional thousand flights, to substantially improve Google doesn't seem like a big cost. Consider: Would your life be worse if one random trans-American flight was cancelled today, or if Google searches became 10% worse? Another way of thinking about this same point, is, why is the author writing about language models if they are so concerned about the environment. Surely the airline industry is a better subject as, again, they fly hundreds of thousands of flights each day. It's hard to take someone seriously when they are focusing on an infinitesimal part of a huge problem. It's also misleading because there is a difference between energy used by an airline flight, where the energy comes from burning jet fuel, and energy used in a data center. In Google's case (the people training BERT) the energy used was 100% renewable - Google reached that goal in 2017. Perhaps Open AI didn't use renewable energy to train GPT-3, but I wager they didn't power their machines by burning jet fuel either. Maybe electricity used to power to train language models will become a meaningful issue at some point in the future. I don't think that future is close at hand though.
- adamsea 6y ago
- spiznnx 6y agoBERT is much smaller than GPT-3 (500x fewer parameters), not sure why the article doesn't point that out more explicitly. The mentioned paper itself notes a 300,000x increase in compute used for training language models over the last 6 years. I think the point is, if nothing changes, the costs will be very significant very soon.
- qeternity 6y agoWe’re already at soon. Another order of magnitude increase (which at our pace is all but guaranteed in the near term) will make large scale training a significant R&D line item even for the FAANGs. It will also push efforts to be commercially viable, as you can longer do proof-of-concept nets at $500m
- PragmaticPulp 6y ago> I think the point is, if nothing changes, the costs will be very significant very soon. The energy efficiency of machine learning hardware is progressing at a rapid rate. It's not accurate to assume that the energy costs will stay the same. Just look at how much it would have taken to train something like GPT-3 five years ago.
- _jal 6y agoI think it is reasonable to assume that any advances in efficiency will be subsumed by increased use, much in the same way increased energy efficiency doesn't decrease absolute overall use.
- joe_the_user 6y agoML hardware is no doubt progressing. Whether we can expect a rapid rate is a different matter. The various Moore's Laws have been breaking down (chip speed stopped doubling a while back). GPU are the sort of chip that could most benefit from just transistors doubling (last Mooore's law standing) but even this Moore's Law is breaking down. ML model size has had it's own Moore's Law, with standard model growing exponentially in size [1]. And this implies model are going butt up against the limits even more than they have already. Whether the "bitter lesson"[2] of ML is true inherently is up for a question. That current researchers have accepted it seems given. [1] https://openai.com/blog/ai-and-compute/ https://openai.com/blog/ai-and-compute/ [2] http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- acchow 6y ago> There are hundreds of thousands of flights every day 2019 was peak flight with 38.9 million flights over the year. That averages to 106k flights per day.
- acdc4life 6y agoMy brain does nlp better than any system out there. I’m also able to ride bicycles and do motor control better than Boston Dynamics. I can also construct and prove math, do physics and code all of this in Matlab and C. My brain can handle all these wide range of tasks almost seamlessly with just 15 watts of power, that Silicon Valley’s super computers can barely do 1% of.
- omgwtfbyobbq 6y agoTo be fair, it takes years of energy to train our brains to the point where it can do all those things, and our brains aren't extensible in the same way hardware/software is. I guess there's also a lot more variation in yields. ;)
- acdc4life 6y agoWell, that’s not entirely true. Alpha go self played 29 million games of Go, which is in feasible for humans. Assuming a game of Go is 5mins, it would take a human 286 years to achieve the same, assuming this person doesn’t sleep, eat or do anything else other than just play Go. CPU time is an order of magnitude faster than real time, especially on GPU clusters.
- bumby 6y agoTo be a fair comparison wouldn’t you have to include all the energy used to train “your” brain through generations of evolutionary training? Your latest model is like taking an already trained BERT model and adding a few tweaks
- acdc4life 6y agoI see BERT as non sensical. You need to be scientific and have a mathematical theory on how humans learn language, which is a multi disciplinary task requiring physicists, mathematicians, neuroscientists, cognitive psychologists and linguists. Benchmarks are useless, theories, models, experiments and testable predictions is how science progresses. You’re making a comment on cognitive science, and trying to imply that language learning in humans isn’t learned, but pre baked. The psychological, linguistic, evolutionary biology and neuroscience evidence doesn’t seem to corroborate. The evidence points stronger to humans having general learning and problem solving abilities. For instance, there was no evolutionary pressure for humans to be good at math or programming. I was not born knowing english or calculus or probability theory, these were learned abilities. Evolution favoured brain mechanisms that lead to behaviour for success in a rapidly changing world. Had I been born in ancient Rome as a farmer, I would learn to speak Latin, and learn how to be a successful farmer, instead of the physics, math, probability, computer, driving, reading skills that I learned in my life time.
- sgt101 6y agoAlso, I have downloaded and used Bert in several other applications; and I think that 10,000's of other folks have done the same.
- jkochis 6y agoThe comparison between training BERT and a 5 year old made me wonder what the carbon footprint is to raise someone to the age of 5.
- throwaway2245 6y ago> There are hundreds of thousands of flights every day - why should we be concerned about the equivalent of one extra See also: Sorites fallacy. There is a limited desire for flights, which benefit people as far as it allows them to get things where they need to be, and no more than that; there is presumably an unlimited desire for machine learning.
- jefftk 6y ago> One recent model called Bidirectional Encoder Representations from Transformers (BERT) used 3.3 billion words from English books and Wikipedia articles. Moreover, during training BERT read this data set not once, but 40 times. To compare, an average child learning to talk might hear 45 million words by age five, 3,000 times fewer than BERT. The human brain is the output of an incredible number of generations of training, representing a vast consumption of energy. Most of the learning that informs the brain happened before this hypothetical five-year-old was even born.
- ravi-delia 6y agoThat is a matter of great debate. Although most people's brains wind up settling on basically identical structure over time, case studies with injured or disabled people often display incredible adaptability (ie the occipital lobe in a blind person). If indeed the brain learns most of what it knows from experience, than its (comparatively) low energy consumption would come from greater efficiency. That definitely seems possible; as a basic unit of machine learning the neuron is much more specialized than the transistor, and far slower.
- anaphor 6y agoMy interpretation of what the parent meant is that the cognitive processes which support language acquisition (which may be mostly domain-general) are already optimized for this task, for which there is a sensitive period during certain years where you need to be exposed to certain inputs, or else you never fully acquire language. See: https://en.wikipedia.org/wiki/Genie_(feral_child) https://en.wikipedia.org/wiki/Genie_(feral_child)
- MAXPOOL 6y ago550–600 million years of hyperparameter tuning and neural architecture search using an evolutionary algorithm is nothing to laugh at. But the actual learning process of a single brain uses very little training examples and is very energy efficient. A brain that uses only 20 watts of power.
- 6y ago
- firebaze 6y agoThis is beyond ridiculous, even if compared to the energy to train a human mind at the age of, let's say, 35. Raise a child to the age of 35 and the spent energy will surpass the mentioned energy cost by at least 2 orders of magnitude. And no, I don't ridicule AI ethics researchers: the energy cost of AI is negligible to the serious issues. The real harm stems from other areas.
- goatlover 6y agoBut that person would be learning tons of other things as well. It’s not equivalent.
- firebaze 6y agoThat person would probably be one of a billion. The AI is singular. This should cancel out.
- MAXPOOL 6y agoHuman mind (aka brain) has consumed little over 6000 kWh energy when it's 35. From the paper they cite algorithm kWh-PUE ---------- ------- ELMo 275 BERT(base) 1,507 NAS 656,347 NAS was Transformer (213M parameters) for neural architecture search 2019. They didn't have kWh numbers for GPT-2.
- sbierwagen 6y agoI don't know about that 6000 KWh number. More importantly, a human brain isn't trained from scratch at birth. It's had a few hundred million years of pretraining.
- wongarsu 6y agoThe 6000kWh number is assuming 20W over 35 years. That might be right if you only count the brain, but then you would have to revise the neural network numbers to only count CPU and memory, which would easily halve those numbers (fans take a lot of power). If we want to account for power supply, temperature control etc. we could use the entire calorie count of a human over 35 years, leaving us with around 30000kWh (assuming a health, not-overweight human). You could now argue that that's too high (the human is doing a lot of other things). But as you pointed out it's a bit of a pointless comparison anyways as humans don't start from scratch.
- pelasaco 6y agoIt still takes less energy than raise humans and educate them to do mediocre jobs.
- goatlover 6y agoPlenty of humans are already raised, and can do other things machines can’t.
- pelasaco 6y agoplenty of humans are raised but not educated. So let's educate them to do the other things that machines can't.
- joe_the_user 6y agoThere are very few "typical human tasks" that machine learning actually does better than a person at; playing game all the occurs to me. But driving, image description, language-translation, and so-forth, at the scale that a human does it, the human is generally still considered to do tremendously better. Of course, some ML programs are convenient in being able to do things at a large scale but a majority of what ML programs do is like language translation - not great and what we sometimes live with 'cause it's free. Basically, one should really not discount human ability, especially in tasks considered mundane. Such skills often are often considered "mediocre" not because they're actually easy but because all human can do them, see: https://en.wikipedia.org/wiki/Moravec%27s_paradox https://en.wikipedia.org/wiki/Moravec%27s_paradox
- mortehu 6y agoML is usually better at speed. For example I use a fine tuned GPT-2 for code autocomplete in vim, and even though humans could do better, they can't do it as fast.
- pelasaco 6y agoThe same was reported last year. ML could offload lawyers, doing many tasks done by them. Not better, but faster and for sure, spending less energy.
- cosmolev 6y agoMachine learning is essentially a brute force. No surprice it is not efficient.
- lightgreen 6y ago> Among the risks is the large carbon footprint of developing this kind of AI technology. No there are no risks of carbon footprint. Energy is cheap. Just build nuclear power plants. Or just build solar and wind and train models half of the time if one is paranoid about nuclear power plants safety.
- Der_Einzige 6y agoFine tuning BERT (what most people in practice do) is so much more efficient. You can do it in 8 hours on a 2080ti. They need to mention that the total number of people training large transformers from scratch is very very small. If wager that the total number of different, uniquely trained (not using previous weights - which reduces compute necessity by massive amounts) language models in existence is in the low hundreds I'd claim that these mass language models serving as the underlying encoding backbone behind more specific systems actually save energy and compute compared to the previous methods (needing far more data and thus more energy spent on getting it combined with less efficient representations like tf-idf causing many classifers to perform very slowly and thus burn lots of energy) Also, much of the recent research in this field is about model pruning, quantization, and any technique you can imagine to reduce training and inference time or memory requirements. All in all, big language models are a net positive for the environment. The effeciency gains in any number of fields from increasingly sophisticated NLP systems far, far outweighs the costs or of training them. Foundational research in environmental conservation will be accelerated by effective NLP semantic search and question answering systems. That's a single, tiny example of the potential for benefits from large language models. Pick a better target.
- Jabbles 6y agoWait until the author discovers bitcoin's energy costs.
- throwaway7281 6y agoThat's my #1 reason to be bearish on B$ - it's just a complete environmental cluster-fuck. And in order to not see this, you'll have to ignore a lot of facts - which in turn tells me a lot about those inside the crypto-bubble, namely that they do not care that much about facts.
- trhway 6y agoat Starship's $100/kg placing your Eth mining $1000 GPU in space will add only 10-20% when amortized over 100 ton rig - it is comparable with building/buying your own small powerplant what large crypto have done. AI and crypto are exploding and only going more so. Granted, the planet is becoming too small a confines for it. It seems that AI and crypto will be the killer apps of space in near future. One can also observe that humans have largest, among the animals of planet Earth, share of body energy consumed by brain and the future humans would probably have even higher share. In technology we observe the same - the pinnacle of technology - CPU - have practically 0 thermodynamic efficiency and the share of energy consumed by computers grows, and i think it will be only growing. Intelligence eats the world.
- mhh__ 6y agoHow are you supposed to cool it in space?
- hnuser123456 6y agoHeatsinks optimized for radiative cooling, which are kept in the shade by solar panels. Works much better when half or more of your surroundings aren't a planet at human-scale temperatures, as is the case on earth's surface.
- 6y ago
- ur-whale 6y agoSame BS that got started the whole Gebru affair at Google.
- HenryKissinger 6y agoI just want to say that the human brain doesn't need to perform billions of matrix products to recognize a dog. The best AI will always be a meaty human.
- syntaxing 6y agoIt’s still crazy to think about how much more efficient computational power is nowadays compared to even five years ago. I remembered I changed my power supply in anticipation for Nvidias then newest GTX series. But only to learn that a 1050Ti only had a 75W max draw so the new fancy 1kW PSU was completely unnecessary. Also to put into perspective, my Mac mini has a 30W max draw. 30W!!! It’s ridiculous how little energy it draws.
- trthomps 6y agoYou want your mind to really be blown, your brain manages to do everything it does on around 20W, even less than that mac mini.
- visarga 6y agoThe brain is not just more efficient, it also builds itself. The hardware necessary to run a neural net is given. Maybe the rigors of self replication are what's missing from AI.
- optimalsolver 6y agoDidn't pointing this out get that Google lady fired?
- Uhhrrr 6y agoDr. Gebru? No, she said she'd resign if [some demands] weren't met and Google said, "We can't meet those demands so we accept your resignation."
- visarga 6y agoProblem is Gebru is an activist, she's biased towards her own cause. Accepting her demands directly, without including all the other groups and voices, would not be right.
- beervirus 6y ago> This month, Google forced out a prominent AI ethics researcher after she voiced frustration with the company for making her withdraw a research paper. The paper pointed out the risks of language-processing artificial intelligence, the type used in Google Search and other text analysis products. That is... certainly one way of putting things.
- nerdponx 6y agoIs there more context for this?
- deleted 6y ago[deleted]
- TaupeRanger 6y agoAnd they aren't even learning in a sense that would make them categorically comparable to human-like learning. You could just say "it takes a lot of energy for a computer to crunch a lot of numbers".
- paulpauper 6y agoWhy does the computer need to learn when we can just program what the computer is supposed to do based on what we know what the desired output/result should be? Yeah, we can try to recreate calculus from first principles, or just use the equations that we already know. unless I am missing something obvious. There are PHP or C scripts for example can do image recpnition and break captchas (there is a program called Xrummer that solves captchas for spamming purposes, which is why captchas have become so complicated) ...this was in 2010 , long before machine learning became a 'thing'.
- visarga 6y agoAs complicated as it is, deep learning is simple compared to programming by hand all the required knowledge. This approach historically failed (expert systems).
- nerdponx 6y agoI don't know that a post this arrogant warrants a serious response. If you think you can write a PHP script that does what BERT or ImageNet does, by all means go right ahead.
- belval 6y agoI think most of the critics here has to come from people not understanding that GPT-3 was not and will not be trained more than a few times. If even then you still want to argue, aluminum consumes 17,000kWh per ton ton which is 11 BERTs. I don't see weird criticism about mining consuming energy. Doing stuff takes energy and while the sentiment against ML seems to be "it's snake oil", it is still very real research with actual real-world impacts. Finally, we are in a transition phase where ML is done on GPUs which consume more energy than dedicated ASICs. We already have Google TPUs and Amazon Inferentia which can be used today. Power consumption will go down as dedicated hardware gets better and better.