7 ms·
The cost to train an AI system is improving at 50x the pace of Moore’s Law
- seek3r00 6y agotl;dr: Training learners is becoming cheaper every year, thanks to big tech companies pushing hardware and software.
- anonu 6y agoArk Invest are the creators of the ARKK [1] and ARKW ETFs that have become retail darlings, mainly because they're heavily invested in TSLA. They pride themselves on this type of fundamental, bottom up analysis on the market. It's fine.. I don't know if I agree with using Moore's law which is fundamentally about hardware, with the cost to run a "system" which is a combination of customized hardware and new software techniques [1] https://pages.etflogic.io/?ticker=ARKK https://pages.etflogic.io/?ticker=ARKK
- deleted 6y ago[deleted]
- gentleman11 6y agoDespite nvidia vaguely prohibiting users from using their desktop cards for machine learning in any sort of data center-like or server-like capacity. Hopefully AMDs ml support / OpenCl will continue improving
- QuixoticQuibit 6y agoLast I saw, they don’t even support ROCm on their recent Navi cards, so I’d be hesitant.
- Reelin 6y agoWow. This is really disappointing to see. (https://github.com/RadeonOpenCompute/ROCm/issues/887 https://github.com/RadeonOpenCompute/ROCm/issues/887) I guess PlaidML might be a viable option?
- m3kw9 6y agoIt was probably because very inefficient to begin with.
- techbio 6y agoIndeed nonexistent
- gchamonlive 6y agoI remember this article from 2018: https://medium.com/the-mission/why-building-your-own-deep-learning-computer-is-10x-cheaper-than-aws-b1c91b55ce8c https://medium.com/the-mission/why-building-your-own-deep-le... Hackernews discussion for the article: https://news.ycombinator.com/item?id=18063893 https://news.ycombinator.com/item?id=18063893 It really is interesting how this is changing the dynamics of neural network training. Now it is affordable to train a useful network on the cloud, whereas 2 years ago that would be reserved to companies with either bigger investments or an already consolidated product.
- qayxc 6y ago> Now it is affordable to train a useful network on the cloud I honestly don't see how anything changed significantly in past 2 years. Benchmarks indicate that a V100 is barely 2x the performance of an RTX 2080 Ti [1] and a V100 is • $2.50/h at Google [2] • $13.46/h (4xV100) at Microsoft Azure [3] • $12.24/h (4xV100) at AWS [4] • ~$2.80/h (2xV100, 1 month) at LeaderGPU [5] • ~$3.38/h (4xV100, 1 month) at Exoscale [6] Other smaller cloud providers are in a similar price range to [5] and [6] (read: GCE, Azure and AWS are way overpriced...). Using the 2x figure from [1] and adjusting the price for the build to a 2080 Ti and an AMD R9 3950X instead of the TR results in similar figures to the article you provided. Please point me to any resources that show how the content of the article doesn't apply anymore, 2 years later. I'd be very interested to learn what actually changed (if anything). NVIDIA's new A100 platform might be a game changer, but it's not yet available in public cloud offerings. [1] https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v100-vs-titan-v-vs-1080-ti-benchmark/ https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v... [2] https://cloud.google.com/compute/gpus-pricing https://cloud.google.com/compute/gpus-pricing [3] https://azure.microsoft.com/en-us/pricing/details/virtual-machines/linux/ https://azure.microsoft.com/en-us/pricing/details/virtual-ma... [4] https://aws.amazon.com/ec2/pricing/on-demand/ https://aws.amazon.com/ec2/pricing/on-demand/ [5] https://www.leadergpu.com/#chose-best https://www.leadergpu.com/#chose-best [6] https://www.exoscale.com/gpu/ https://www.exoscale.com/gpu/
- solidasparagus 6y agoYou are missing TPU and spot/preemptible pricing, which need to be considered when we are talking about training cost. The big one to me is the ability to consistently train on V100s with spot pricing, which was not possible a couple of years ago (there wasn't enough spare capacity). Also, the improvement in cloud bandwidth for DL-type instances has helped distributed training a lot.
- calebkaiser 6y agoThis is an odd framing. Training has become much more accessible, due to a variety of things (ASICs, offerings from public clouds, innovations on the data science side). Comparing it to Moore's Law doesn't make any sense to me, though. Moore's Law is an observation on the pace of increase of a tightly scoped thing, the number of transistors. The cost of training a model is not a single "thing," it's a cumulative effect of many things, including things as fluid as cloud pricing. Completely possible that I'm missing something obvious, though.
- gumby 6y agoLike many things, Moore’s law is garbled when adopted by analogy outside its domain. What does “more transistors” mean? To you, it means just what Gordon Moore means when he said it: opportunity for more function in same space/cost. The laypersons, marketing grabbed the term and said it would imply “faster”. Which then was absurdly conflated with CPU clock speed (itself an important input, though hardly the only one, determining the actual speed of A system). The use here is of the “garbled analogy” sort which surely is the dominant use today.
- bcrosby95 6y agoYes but that aspect of Moore's law for CPUs expired over a decade ago. It's the whole reason we got multicore in the first place.
- andrewprock 6y agoEven with multi-core, a CPU today is only 6x faster than a 10-year old CPU.
- reitzensteinm 6y agoThe difference might be even less. 4 Sandy Bridge cores (excluding memory controller and graphics) were not much bigger than the current 8 core Zen 2 die. Certainly the peak performance you can put in a socket is much higher, but it's got more silicon in it than it used to.
- mellosouls 6y agoIt is regrettable if an equivalent to the self-fulfilling prophecy of Moore's "Law" (originally an astute observation and forecast, but not remotely a law) became a driver/limiter in this field as well, even more so if it's a straight transplant for soundbite reasons rather than through any impartial and thoughtful analysis.
- kens 6y agoOne thing I've wondered is if Moore's Law is good or bad, in the sense of how fast should we have been able to improve IC technology. Was progress limited by business decisions or is this as fast as improvements could take place? A thought experiment: suppose we meet aliens who are remarkably similar to ourselves and have an IC industry. Would they be impressed by our Moore's law progress, or wonder why we took so long?
- NortySpock 6y agohttps://en.wikipedia.org/wiki/Moore%27s_law https://en.wikipedia.org/wiki/Moore%27s_law, third paragraph of the header, claims that Moore's Law drove targets in R&D and manufacturing, but does not cite a reference for this claim. "Moore's prediction has been used in the semiconductor industry to guide long-term planning and to set targets for research and development."
- imtringued 6y agoI'm not sure what the point of that question is. In theory you could have a government subsidize construction of fabs so that skipping nodes is feasible but why on earth would you do that when the industry is fully self sufficient and wildly profitable?
- solidasparagus 6y agoResnet-50 with DawnBench settings is a very poor choice for illustrating this trend. The main technique driving this reduction in cost-to-train has been finding arcane, fast training schedules. This sounds good until you realize its a type of sleight of hand where finding that schedule takes tens of thousands of dollars (usually more) that isn't counted in cost-to-train, but is a real-world cost you would experience if you want to train models. However, I think the overall trend this article talks about is accurate. There has been an increased focus on cost-to-train and you can see that with models like EfficientNet where NAS is used to optimize both accuracy and model size jointly.
- sdenton4 6y agoI would guess that this means DawnBench is basically working. You'll get some "overfit" training schedule optimizations, but hopefully amongst those you'll end up with some improvements you can take to other models. We also seem to be moving more towards a world where big problem-specific models are shared (BERT, GPT), so that the base time to train doesn't matter much unless you're doing model architecture research. For most end-use cases in language and perception, you'll end up picking up a 99%-trained model, and fine tuning on your particular version of the problem.
- deleted 6y ago[deleted]
- ersiees 6y agoI would really like a thorough analysis on how expensive it is to multiply large matrices, which is the most expensive part of a transformer training for example according to the profiler. Is there some Moore’s law or similar trend?
- deleted 6y ago[deleted]
- sktguha 6y agoDoes it mean that the cost to train something like gpt3 by OpenAI will reduce from 12 million dollars to less next year ? If so how much will it reduce to ?
- lukevp 6y agoWhat are some domains that a solo developer could build something commercially compelling to capture some of this $37 trillion? Are there any workflows or tools or efficiencies that could be easily realized as a commercial offering that would not require massive man hours to implement?
- Isinlor 6y agoYou need to be creative. But one example - colorizing old photos: https://twitter.com/citnaj https://twitter.com/citnaj
- jacquesm 6y agoTake any domain that requires classification work that has not yet been targeted and make a run for it. You likely will be able to adapt one of the existing nets or even use transfer learning to outperform a human. That's the low hanging fruit. For instance: quality control: abnormality detection (for instance: in medicine), agriculture (lots of movement there right now), parts inspection, assembly inspection, sorting and so on. There are more applications for this stuff than you might think at first glance, essentially if a toddler can do it and it is a job right now that's a good target.
- yelloweyes 6y agoanything that's even remotely profitable is already taken
- jacquesm 6y agoThis simply isn't true. Every year since the present day ML wave started has seen more and more domains tackled. Even something like that silly lego sorting machine I built could be the basis of a whole company pursuing sorting technology if you set your mind to it. And that's just resnet50 in disguise, likely you could do better today without any effort. Your statement reminds of 'all the good domains are taken', which I've been hearing since 1996 or so. Of course you'll need to do some work to identify a niche that doesn't have a major player in it yet. But the 'boring' niches are where a lot of money is to be made, the sexy stuff (cancer, fruit sorting) is well covered. But more obscure things are still wide open, I get decks with some regularity about new players in very interesting spaces using thinly wrapped ML to do very profitable things.
- bra-ket 6y ago"AI" is not really appropriate name for what it is
- gxx 6y agoThe cost to collect the huge amounts of needed to train meaningful models is surely not growing at this rate.