13 ms·
AI and Compute
- sandover 8y agoIt's machine learning. It's not AI. Please, all, let's try hard to use words that mean what they mean.
- blixt 8y agoI think that ship has sailed. The term "AI" for any behavior by a machine that changes based on input has been in use for at least over 60 years now. Whether it's the ghosts in Pacman, or a disembodied voice that tells you the weather and plays music when you ask it to.
- MarkMMullin 8y agoWe should do our best to get it back into port - part of the whole mess is that the name AI implies things about ML systems that simply aren't true - as a side note, we should also probably start using the word tensor more accurately, we've now enraged enough physics and math folks :-)
- blixt 8y agoThe entire English dictionary has evolved into its current state, and there's several words that used to have the opposite meaning just from stubborn ironic use by the masses. As much as I like to be correct about my use of words, I think AI has established itself as a term that will stick around for now. Besides, I really don't think all the stigma comes from the term "artificial intelligence". You don't have to ever mention the term to a child interacting with Alexa, they will nevertheless greatly overestimate "her" ability. I think because of the anthropomorphic nature of their interactions, and the black box implementation that prevents you from knowing the boundaries of what is possible. This something that video game characters have played on since their conception, to make humans imagine much more complex intents and thoughts behind their "stupid" hard coded behaviors. I'm okay with calling it AI even if it's not even close to on par with human intelligence. :)
- deleted 8y ago[deleted]
- sullyj3 8y agoHow is the word tensor misused? I thought it was just an n-dimensional array of numbers?
- chas 8y agoSimilar to how linear transforms can be represented as 2-dimensional arrays of numbers (that is to say matrices)[0], tensors are a higher dimensional analogue with a rich theory in their own right and a representation as higher-dimensional arrays of numbers. Similarly, if you look at a tensor solely as an n-dimensional array of numbers, it ignores important differences in the mathematical behavior of objects with the same representation. To give an example: Different parts of a tensor can behave differently under change of basis. [1] [0] https://www.youtube.com/watch?v=kYB8IZa5AuE https://www.youtube.com/watch?v=kYB8IZa5AuE [1] https://en.wikipedia.org/wiki/Covariance_and_contravariance_of_vectors https://en.wikipedia.org/wiki/Covariance_and_contravariance_...
- MarkMMullin 8y agoI rest my case :-)
- MikkoFinell 8y agoIs AI defined as "mysterious future thinking computer"? Anything we figure out how to do seems to suddenly fall outside of the definition.
- visarga 8y agoIt's the magically shrinking "AI of the gaps". "AI" covers only things we can't yet do in ML. https://en.wikipedia.org/wiki/God_of_the_gaps https://en.wikipedia.org/wiki/God_of_the_gaps
- dahart 8y agoSome of the fun of language is the volume of usage of a name or phrase causes it to become correct.
- lainga 8y agoReally? The article says "total compute" in the first paragraph and "total computation" in the second. Noun your verbs or don't noun them at all.
- mooneater 8y agoThis implies more centralization, as those with cheap access to vast compute gain a bigger relative edge.
- sgillen 8y agoYes, unfortunately both data and compute will probably become more and more centralized. At least the algorithmic components have a chance to becone available to everyone.
- samstave 8y agoHere is an off the cuff thought: What if there was (or maybe there already is?) a system which is distributed such as was SETI back in the day and its a massively distributed general AI that can be used - and people on a mass scale allow for slices of their compute to be part of the system?
- tlrobinson 8y agoSo Skynet, but it lets you rent parts of itself? I'm sure someone is writing an ICO whitepaper for this right now, if they haven't already...
- nostrademons 8y agoEOS: https://eos.io/ https://eos.io/ Raised $2.7B in its ICO, currently trading at a market cap of $10B. FileCoin: https://filecoin.io/ https://filecoin.io/ Raised $257M in its ICO. Tezos: https://tezos.com/ https://tezos.com/ Raised $232M in its ICO. Those are the 3 largest ICOs of all time, so yes, there is definitely a market for renting part of Skynet. The actual technology may or may not be vaporware or a scam. IMHO the way you build a decentralized P2P system is to give a single really smart programmer enough to live on for a couple years and see what he comes up with, not throw a billion dollars at a Cayman Islands corporation that may or may not use it for anything productive. Sorta like what Ethereum did.
- tw1010 8y agoTo me this just smells like there's some hidden force – not necessarily nefarious but definitely with the power to incentivize an exaggerated lens – pushing OpenAI to make these claims. Maybe it's the desire to keep AI in the limelight as the buzz is fading slightly. Maybe it is SV echo chamber effects, or investors, or a strategy to build hype in order to attract talent to the company. But to me, on a gut level, it doesn't feel completely ethically pure.
- gwern 8y agoHuh? What are you talking about? The escalating compute involved in DL is obvious to anyone reading the papers; OA is just doing the work of putting numbers on the trend.
- PeterisP 8y agoThere's lots of research on doing the same learning with much less resources (e.g. recent paper https://eng.uber.com/accelerated-neuroevolution/ https://eng.uber.com/accelerated-neuroevolution/ , or the example visible in this very article of AlphaZero having much, much less compute than AlphaGoZero and doing better anyways), and even without that simple hardware progress means that random gaming GPUs can handle datasets that were inconvenient a few years ago. I'd say it all depends on the size of datasets - some domains (e.g. unlabeled image data) have "effectively infinite" datasets where the amount of data you can use is limited only by your computing power, but in many other use cases all the data you'll ever get can be processed by a single beefy workstation. More available compute means that we tackle more difficult problems. However, for any single given task it's often not the case that the amount of compute grows. If anything, the graph is not showing the compute required for DL, but the compute available for DL - it gets used simply because it's there.
- geoffreyirving 8y agoYep, we're not claiming this is the compute required for DL, and for specific tasks we expect compute required to fall over time. But better algorithms actually mean compute is more important, not less, and would likely make the growth in available compute more important. For example, if a task is parameterized (by size or difficulty, say), then a better algorithm might change the asymptotic complexity from O(n^3) to O(n^2). A 2x compute increase for the old algorithm would take us from n -> 1.25n, but the new algorithm would go from n -> 1.41n.
- itchyjunk 8y agoSo for research, would using some standard petaflops/s-days when presenting results be useful? Like model x might be 1% more accurate then model Y but for same baseline petaflops/s-day, how does x and y perform? I'm guessing it might not make sense for all types of research though.
- alfalfasprout 8y agoOpenAI and the other research labs (FAIR, Google Brain, MS Research) are heavily focused on image and speech models, but the reality is the vast majority of models deployed in industry don't need DL and benefit more from intelligent feature engineering and simpler models with good hyperparameter tuning. It's definitely the exception that more compute automatically yields more performance.
- sanxiyn 8y agoI disagree. Well, you don't need DL, but DL will usually help. For example, it helps recommendation: https://github.com/NVIDIA/DeepRecommender https://github.com/NVIDIA/DeepRecommender
- svantana 8y agoDawnbench [1] is such an effort (you will need to work out the petaflops yourself from time x performance, but it lists cloud computing cost which probably is more relevant), and MLPerf is an upcoming one [2]. [1] https://dawn.cs.stanford.edu/benchmark/ https://dawn.cs.stanford.edu/benchmark/ [2] https://mlperf.org/ https://mlperf.org/
- ClassAndBurn 8y agoThat is a staggering rate of increase. I can see a future where this is less centralized; learning could happen in "phases" where a local device improves its model given local data and reports back something centrally that can be combined and used to train a shared model. This requires hardware to be miniaturized as non-ML compute has been and when that does happen we'll have the learnings from the current edge computing push. In the mean time I've excited to see what developments are made on both the hardware and software side.
- nschucher 8y agoThis is called federated learning[0] at least by Google. I don't know whether they've added this to more products or whether it works well. It would be interesting to see this done in open source. [0] https://ai.googleblog.com/2017/04/federated-learning-collaborative.html https://ai.googleblog.com/2017/04/federated-learning-collabo...
- ClassAndBurn 8y agoThank you! I was trying to find that before posting but forgot their naming of it.
- justhelpingout 8y agoFind a solid proof-of-work system for sharing signed data in this manner and you will change the world. Especially if you can re-combine the shared model with the local model.
- zinfour 8y agoSounds like what https://www.openmined.org/ https://www.openmined.org/ is working on.
- dmreedy 8y agoI'm going to draw up some charts about hull displacement on ships from the dawn of time up until about 1950. Then we can have a really informed conversation about the naval power of countries through the ages. I think we need to be ready for the implications of there one day being a battleship the size of the Pacific that will allow its owner to rule the world. Forgive the sarcasm, but I'm really put off by the aggressive weak-to-strong generalizations that are going on here. I'm also very excited about AI, but I don't understand how lines like, >> But at least within many current domains, more compute seems to lead predictably to better performance, and is often complementary to algorithmic advances. can be extrapolated to anything more than a fun conversation to have over drinks, or the plot for a bad sci-fi movie about AI (which, to be fair, are also quite prevalent in the current zeitgeist). We're definitely at a new tier of "kinds of problems computers can solve", but surely experience and history in this space should tell us that we need to expect massive, seemingly insurmountable plateaus before we see the next tier of growth, and that that next tier will be much more a matter of paradigm shift than of growth on a line. The systems on this graph all do different things in different ways. It's one thing to abstract over compute power via something like Moore's Law, or societal complexity via the Kardashev scale. But I think we need a much more nuanced set of metrics to provide any kind of insight in to the various AI techniques. Or an entirely different way of looking at 'intelligence'
- dbelchamber 8y agoI completely agree. Current AI is excellent (or at least super-human) at learning to do anything where the mechanics of the situation are clear and where the measurement of success is well defined. Beyond that, I'm not sure we've made any convincing strides towards anything truly general.
- deleted 8y ago[deleted]
- ddtaylor 8y agoI fiddled with an idea where I wrote unit tests and used them for a scoring function to train a model. Writing the number of tests and encoding the logic for a simple Linked List took orders of magnitude more code than coding the list itself.
- tzahola 8y agoFor some reason the word “compute” in this context causes me to throw up in my mouth. It used to be that only “coding” could elicit this reaction - nevertheless I’m quite fascinated by this new development.
- calibas 8y agoI support harsh penalties on anyone who tries to noun a verb.
- bstamour 8y agoVerbing weirds language. Respect your parts of speech!
- jamesblonde 8y agoNo noun is too proper to verbify :)
- dahart 8y agoVerbing nouns and nouning verbs is probably as old as verbs and nouns. These words are all nouned verbs: Chair, cup, divorce, drink, dress, fool, host, intern, lure, mail, medal, merge, model, mutter, pepper, salt, ship, sleep, strike, style, train, voice. (according to this, anyway: https://www.grammarly.com/blog/the-basics-of-verbing-nouns/ https://www.grammarly.com/blog/the-basics-of-verbing-nouns/) Shakespeare verbed nouns. "Compute" as a noun is at least 20 years old, according to my memory, and there are several high profile products named this way that are more than 10 years old.
- deleted 8y ago[deleted]
- zach 8y agoLooking at the trend here, you can see why many business forecasters and economists have predicted that advances in artificial intelligence will create huge new returns to capital. That future is worth reflecting on because it suggests a fundamental change of labor-capital dymamics. Take startups. Right now, many startups can compete on the same basis to hire talent as huge companies. But if companies with huge capital reserves can put their cash directly to work to train AI models, startups will be hard-pressed to compete with "smarter" products. Specialization will not even be much help. Looking at Beating the Averages (http://www.paulgraham.com/avg.html http://www.paulgraham.com/avg.html), PG enthused that, since established companies are so behind the curve on software development technology, there is always a chance for higher-productivity techniques like more productive languages to give smaller teams a real chance at a huge market. Of course, that this was in the era when Google was not creating new programming languages and there were no Facebook to widely deploy OCaml and Haskell. And now, AI looks to make the averages even harder to beat. Even today, if you round up the smartest members of a CS grad class, it is going to be quite difficult to directly compete with a machine learning model with access to huge amounts of data and computing resources. Looking further forwards, if machine learning is able to provide "good enough" alternatives to most human-created software, the software startup narrative — that a few talented and determined people can beat billions in resources — may not even be so relevant anymore.
- ddtaylor 8y agoIt's worth noting that some prominent figures in AI/ML are saying we are due for another "AI winter" since it's being oversold again. I don't know if I agree with that, since we are seeing some interesting things, but technically Google is kind of saying they can tentatively pass the Turing Test with phones and meanwhile even a car decked out with extra sensors and 360 LIDAR cannot detect a simple stop sign with mud on it.
- deleted 8y ago[deleted]
- acdha 8y ago> Google is kind of saying they can tentatively pass the Turing Test with phones Is Google really saying that or just the more breathless commenters? I thought they were pretty good at making it clear that Duplex took a lot of work to do well in very constrained conversational situations.
- westoncb 8y agoWould someone explain the purpose/origin of using 'compute' as a noun like this instead of a verb?
- deleted 8y ago[deleted]
- deleted 8y ago[deleted]
- GuiA 8y agoI think a lot of people in the industry got that word in their vocabulary from its usage in “Amazon EC2” (Elastic Cloud Compute). It’s certainly been used before, but that was one of the first times I remember hearing it in that context.
- tshadley 8y agoArchived 2012 discussion invokes Oxford English Dictionary to trace the original use back several 100 years. http://www.techwr-l.com/archives/1206/techwhirl-1206-00295.html#.WvyF1HUvxNA http://www.techwr-l.com/archives/1206/techwhirl-1206-00295.h...
- mannykannot 8y agoIn those examples, however, the meaning is, in current usage, 'calculation' or 'computation', not as a measure of computational work.
- tshadley 8y ago> In those examples, however, the meaning is, in current usage, 'calculation' or 'computation', not as a measure of computational work. So is the OP: "We’re releasing an analysis showing that since 2012, the amount of compute [amount of calculation, amount of computation] used in the largest AI training runs has been increasing exponentially with a 3.5 month-doubling time " Without loss of meaning, title could be AI and Calculation, or AI and Computation
- petters 8y agoIt's not wrong, but the unit "petaflop/s-day" made me smile.
- kbob 8y ago1 petaFlO/sec × day = 86400 petaFlO = 8.64e19 FlO.
- forcer 8y agoI don't get it. how does OpenAI knows how many resouces are thrown at AI calculations worldwide?
- visarga 8y agoThey are reporting only on a few well known papers. They don't know what people are doing in secret.
- tehsauce 8y agoI think it's important to notice that if we're using the metric of "300,000x" increase in computing power applied to ML models, the giant increase has mostly been due to parallel computing playing catchup on decades of moores law all at once. It will hit a wall and die with moores law fairly soon. Physics requires it.
- spunker540 8y agoHow is parallelism limited by physics? I thought the point of parallelism is you can throw more chips at a problem and see improved performance. Single chips are limited by physics, but true parallelism scales linearly ad infinitum. Can anyone with more knowledge than me speak to known limits of parallelism? I’d guess it’s not truly infinitely scalable.
- ychen306 8y agoYou can't scale linearly ad infinitum because eventually the communication (i.e. memory) cost gets too high. This reminds me of a thought experiment I heard from -- if memory serves -- Scott Aaronson. The gist is that the fastest super-computer will be on the edge of a black hole. If you run any faster, there will be too much energy concentrated on a given area, thus creating a black hole. Similarly, when you run so many parallel devices (on GPU, CPU, etc) together, you will want to put the devices as close to each other as possible (speed of light limits the rate of communication). You then pump too much heat into a small area, and getting so much heat out is, among other things, a physics problem.
- red75prime 8y agoThat's a very far limit, though. It will not have practical consequences for a long time. Also, if you don't squeeze as much as you can into a small space, you can scale sublinearly ad infinitum (in practical terms, which don't include heat death of the universe).
- sheeshkebab 8y agoParent is probably referring to amdahls law - which limits speedup in parallel computing systems https://en.m.wikipedia.org/wiki/Amdahl%27s_law https://en.m.wikipedia.org/wiki/Amdahl%27s_law
- jfaucett 8y ago> On the other hand, cost will eventually limit the parallelism side of the trend and physics will limit the chip efficiency side. Anyone working on chip architecture care to give their opinion on the next 10-20 years in chip design? It would really interest me to know if chip designers think Moore's law will continue, since that is probably going to be a big factor in the timeline for AGI.
- deepnotderp 8y agoNot gonna predict the future 1-2 decades out, since that's a fool's errand, but here's a grab bag of relevant points: 1. Moore's Law is undoubtedly slowing, but in the foreseeable future, it will likely continue. On the other hand, Dennnard Scaling which is already basically dead, will be the crunch you will likely feel more. Exponential transistors aren't too useful if they still consume so much power. To mitigate leakage we moved to FinFETs... Which actually made dynamic power worse. 2. You might be interested to know that data movement (predominantly memory access) costs orders of magnitude more than computation, especially relevant to AI compute which requires large amounts of access. These global wires already suck and don't seem to be getting any better in the foreseeable future. 3. Foundries have already been using (and thus expending) "scaling boosters" to reach their density goals. Most of these are one-time use effects that won't provide significant continuous scaling capability.
- p1esk 8y agoAnalog computing has a lot of yet unrealized potential for machine learning algorithms. However, currently it does not make sense to build a specialized analog chip to run specific type of ML algorithms, because algorithms are still being actively developed. I don't see GPUs being replaced by ASICs any time soon. And before you point to something like Google's TPU, the line between such ASICs and latest GPUs such as V100 is blurred.
- deepnotderp 8y agoPlease explain where analog computation has a benefit over digital that outweighs its numerous disadvantages.
- forapurpose 8y ago> Three factors drive the advance of AI: algorithmic innovation, data (which can be either supervised data or interactive environments), and the amount of compute available for training. Algorithmic innovation and data are difficult to track ... Are algorithmic innovations and improvements in data so difficult to track? Could they be measured by the cost of certain outputs? Or is it that the information about algorithms and data is not easily accessible?
- nutanc 8y agoThough this talks about current trends, I would place my bets on a more radical future where the current algorithms for AI are overhauled and we get much better and faster algorithms which can even work on generic CPUs.
- cmarschner 8y agoCherry-picking a few papers doesn‘t tell anything. If at all it shows what people have achieved who pushed the envelope to the extreme, mostly at Google where people can afford to not care about cost. 99.9% of the work is done using small numbers of GPUs, and that hasn‘t changed much in recent years, except for the improvements in GPU architectures. Draw this graph and you get a very different story.