10 ms·
Author here: I'm using deep learning daily so I have a bit of an idea on what I'm talking about. 1) Not my point. Hype is doing very well. But narrative begins
by felippee 8y ago
Author here: I'm using deep learning daily so I have a bit of an idea on what I'm talking about.
1) Not my point. Hype is doing very well. But narrative begins to crack, actually indicative of a burst...
2) DL does not scale very well. It does scale better than other ML algorithm because those did not scale at all. If you want to know what scales very well, look at CFD (computational fluid dynamics). DL in nowhere near that ease in scaling.
3) self driving is the poster child of current "AI-revolution". And it is where by far most money is allocated. So if that falls, rest of DL does not matter.
4) Not that this matters, does it?
- evrydayhustling 8y agoThe scaling argument in the article doesn't make any sense. There are rhetorical queries like "does this model with 1000x as many parameters work 1000x as well?" but what it means to scale or perform are not clearly or consistently defined - let alone defined in a way that would make your point about the utility of the advances. OpenAI's graph shows new architectures being used with more parameters because people are innovating on architecture and scale at the same time. Arguing that old methods "failed to scale" is like arguing that processor development was a failure because Intel had to develop a 486 instead of making a 386 work with more transistors (or more something). And what does CFD have to do with anything, except maybe an odd attempt to argue from authority? Can you formalize from CFD a notion of "scaling well" well that anyone else agrees is useful for measuring AI research?
- prewett 8y agoCFD was merely used as an example of something that does scale well. I'm not sure it was the best example, since CFD isn't very common. But basically you have a volume mesh and each cell iterates on the Navier-Stokes equation. So if you have N processor cores, you break the mesh in N pieces, each of which get processed in parallel. Doubling the number of cores allows you process double the amount in the same time, minus communication loses (each section of the mesh needs to communicate the results on its boundary to its neighbors). I don't fully understand the graph, but it looks like his point is that Alpha Go Zero uses 1e5 times as many resources than AlexNet, but does not produce anywhere near 10,000 times better results. We saw that with CFDt 1e5 more cores resulted in 1e5 better results (= scales). The assertion is that DL's results are much less than 1e5 better, hence it does not scale. Basically the argument is: 1. CFD produces N times better results given N times more resources [this is implied, requires a knowledge of CFD]. That is, f(ax) = a f(x). Or, f(ax) = 1 a * f(x). 2. Empirically, we see that DL has used 1e5 more resources, but is not producing 1e5 times better results. [No quantitative analysis of how much better the results are is given] 3. Since DL has f(a * x) = b * a * f(x), where b < 1, DL does not scale. [Presumably b << 1 but the article did not give any specific results] This isn't a very rigorous argument and the article left out half the argument, but it is suggestive.
- felippee 8y agoThanks for that, that is essentially my point. Agree it is not very rigorous, but it gets the idea across. By scalable we'd typically think "you throw more gpu's at it and it works better by some measure". Deep learning does that only in extremely specific domains, e.g. games and self play as in alpha go. For majority of other applications it is architecture bound or data bound. You can't throw more layers, more basic DL primitives and expect better results. You need more data, and more phd students to tweak the architecture. That is not scalable.
- deleted 8y ago[deleted]
- evrydayhustling 8y agoMore compute -> more precision is just one field's definition of scalable... Saying that DNNs can't get better just by adding GPUs is like complaining that an apple isn't very orange. To generalize notions of scaling, you need to look at the economics of consumed resources and generated utility, and you haven't begun to make the argument that data acquisition and PhD student time hasn't created ROI, or that ROI on those activities hasn't grown over time. Data acquisition and labeling is getting cheaper all the time for many applications. Plus, new architectures give ways to do transfer learning or encode domain bias that let you specialize a model with less new data. There is substantial progress and already good returns on these types of scalability which (unlike returns on more GPUs) influence ML economics.
- felippee 8y agoOK, the definition of scalable is crucial here and it causes lots of trouble (this is also response to several other posts so forgive me if I don't address your points exactly). Let me try once again: an algorithm is scalable if it can process bigger instances by adding more compute power. E.g. I take a small perceptron and train it on pentium 100, and then take a perceptron with 10x parameters on Core I7 and get better output by some monotonic function of increase in instance size (it is typically a sub linear function but it is OK as long as it is not logarithmic). DL does not have that property. It requires modifying the algorithm, modifying the task at hand and so on. And it is not that it requires some tiny tweaking. It requires quite a bit of tweaking. I mean if you need a scientific paper to make a bigger instance of your algorithm this algorithm is not scalable. What many people here are talking about is whether an instance of the algorithm can be created (by a great human effort) in a very specific domain to saturate a given large compute resource. And yes, in that sense deep learning can show some success in very limited domains. Domains where there happens to be a boatload of data, particularly labeled data. But you see there is a subtle difference here, similar in some sense to difference between Amdahl's law and Gustafson's law (though not literal). The way many people (including investors) understand deep learning is that: you build a model A, show it a bunch of pictures and it understands something out of them. Then you buy 10x more GPU's, build model B that is 10x bigger, show it those same pictures and it understands 10x more from them. Look I, and many people here understand this is totally naive. But believe me, I talked to many people with big $ that have exactly that level of understanding.
- nopinsight 8y agoWhat do you think of advances like those in major DeepMind papers? They seem to represent significant shifts in capabilities once matured. Here's a recent example: Unsupervised Predictive Memory in a Goal-Directed Agent https://news.ycombinator.com/item?id=17177442 https://news.ycombinator.com/item?id=17177442
- goolulusaurs 8y agoThis paper is amazing, and exactly what I was thinking of posting in response to that part of the article. The amount of research Deepmind is putting out is astonishing, and even if you are paying attention it's hard to keep up with it all. Here are just a few papers I've been looking at from the last few months, maybe none of them are advancements on the level of AlphaGo & Zero but they still show significant progress in a wide variety of areas. https://arxiv.org/abs/1802.10542 https://arxiv.org/abs/1802.10542 https://arxiv.org/abs/1802.07740 https://arxiv.org/abs/1802.07740 https://www.nature.com/articles/s41586-018-0102-6 https://www.nature.com/articles/s41586-018-0102-6 https://arxiv.org/abs/1804.09401 https://arxiv.org/abs/1804.09401 https://arxiv.org/abs/1805.06370 https://arxiv.org/abs/1805.06370 https://arxiv.org/abs/1802.03006 https://arxiv.org/abs/1802.03006 https://arxiv.org/abs/1804.08617 https://arxiv.org/abs/1804.08617 https://arxiv.org/abs/1802.01561 https://arxiv.org/abs/1802.01561 And there are many others besides these, not to mention all the significant research being done by everyone else who isn't at Deepmind. The authors idea that interest and development of these topics is dying down or that Deepmind is running out of meaningful research to do just seems uninformed.
- oh-kumudo 8y ago2)Why do you think DL doesn't scale? I am curious. It can easily leverage thousands of GPUs, training on 300 millions of images (https://ai.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html https://ai.googleblog.com/2017/07/revisiting-unreasonable-ef...). No other methods is even close to leverage that amount of computational power. I don't really know about CFD, but at least in ML land and dealing with ML problems, DL is very scalable, maybe only next to random forests style algorithm, where they effectively share nothing. 3)It does matter. In fact most valuable startup around DL are CV based startups, they are mainly located in China though.
- jedbrown 8y agoCFD is good at using big machines "efficiently", but the cost of DNS scales as the cube of the Reynolds number which will never be tractable for most engineering problems. Apart from niche basic research on the edge of tractability, all the effort goes into modeling (RANS, DES, wall, etc.) to deliver statistically calibrated estimates of functionals of interest at feasible cost. Those methods actually don't "scale as well" (though the state of research is ahead of commercial software), but also don't need to because they can solve the problem in less time with less hardware. This situation is actually pretty similar to your DL analogy where more hardware provides diminishing returns for solving the actual problem.
- edejong 8y ago"Author here: I'm using deep learning daily so I have a bit of an idea on what I'm talking about." Very weak to appeal to authority. The only true argument I can find against DL/ML/AI atm is the continuing appeal to authority by PhDs who have zero engineering knowledge, zero business sense and zero understanding of risk assessment.