5 ms·
I think there's an interesting disconnect right now between research and practice. Cutting-edge research does feel like it's reaching a plateau - across most AI
by rococode 7y ago
I think there's an interesting disconnect right now between research and practice. Cutting-edge research does feel like it's reaching a plateau - across most AI fields even "major" breakthroughs are only gaining a couple percentage points and we're probably starting to hit the limits of what current approaches can achieve. When the state-of-the-art is 97% on a task, there's only so much room for improvement. Yoav Goldberg posted a tweet about Facebook's RoBERTa model that summed it up pretty well: "oh wow seems like this boring public hyperparameter search is going to take a while" [1]. There's a vague feeling of "What's next?" now that all the benchmarks are fairly well-solved but AI in general clearly doesn't feel solved.
However, state-of-the-art models aren't really used in production yet. I think the trend of "use AI/ML to solve X" has only started to pick up in the past 2 years, and it'll continue well into the 2020s. The process of taking research models and putting them into production is not standardized yet, and many models don't even really work in production - if your model takes a second to do an inference step that's fine for research but maybe not for a real product.
I think in the next decade, on the research side, benchmarks will be beaten less often, and instead there will be more focus on trying out radically new things, understanding weaknesses in current techniques, and finding new measurements that assess those weaknesses. On the industry side, there will still be lots of cool and exciting new achievements as already-known techniques are applied to old problems that haven't been addressed by AI yet.
As an aside, this was the first time in my life that I read the phrase "10s" referring to the 2010-2019. Kind of an odd-feeling moment!
[1] https://twitter.com/yoavgo/status/1151977499259219968 https://twitter.com/yoavgo/status/1151977499259219968
- Findeton 7y agoI have some hope that research like what Numenta is based upon will lead us closer to an AGI.
- Veedrac 7y ago> Cutting-edge research does feel like it's reaching a plateau It's really not. The second half of last year alone had MuZero and Megatron-LM, to name just a couple that most scream to me that we are actually progressing towards AGI. You say ‘When the state-of-the-art is 97% on a task’, but solved tasks are the least interesting tasks.
- tomrod 7y ago> solved tasks are the least interesting tasks. Yeah, but they also tend to be pretty profitable.
- antpls 7y agoAlso, Google released the Reformer just 2 days ago, and they claim it can ingest orders of magnitude more data than the Transformer. https://ai.googleblog.com/2020/01/reformer-efficient-transformer.html?m=1 https://ai.googleblog.com/2020/01/reformer-efficient-transfo... TPUv3 is estimated to be 12 or 16nm process node, so the performance of TPU's next versions could still double over the next years (if needed). At that pace of model research and hardware improvement, I would say we are still in the AI spring.
- Piskvorrr 7y agoDo we have any other indicators that it's actually progressing somewhere, besides screaming? Even such a triviality as "how do we recognize that we got there"? The research is still in very early phases, IMNSHO: impressive practical applications appear, but they're side effects of what appears as random flailing: "build it bigger, see if it helps. Build it sideways, see of it helps. Build it at full moon, see if it helps". That suggests that the applications are the low-hanging fruit, with far more interesting results still to be discovered.
- Veedrac 7y agoMuZero is in some sense the proto-holy grail, in that it implements learning and planning into unstructured tasks over purely internal models. While there is an obvious chasm between it and the end point, this is still something that has only recently become more than an abstract goal, at least to any effect. Being able to perform planning over ‘simple’ domains like Atari games and Go (and not even in the same trained model!) might not seem very comparable to the real thing, but evolutionary history spent the bulk of its time building up the basics—most animals fail most cognitive tasks—so I don't think this is indicative of the progress being misguided, especially given networks-on-GPUs is literally a 10 year old field. I think MuZero is a clear example of building by principles over random flailing. I get why there does also seem to be the latter, but it's certainly not the whole of it, and anyhow it worked for evolution ;P.
- jean- 7y ago> When the state-of-the-art is 97% on a task, there's only so much room for improvement. When models commonly achieve 97% on a task, it means it's time to define a harder task, as it's long stopped providing any useful signal.
- drongoking 7y agoI think when performance on a real world task can be expressed as a single percentage, we've over-simplified the hell out of it and it's time to rethink the problem.
- Piskvorrr 7y agoOr worse, overfitted - which means the solution will implode when faced with real data (RIP EH).
- tintor 7y agoDifference between 97% and 99.99% in perception is huge for autonomous driving purposes. 300x less likely to cause an accident.
- stefan_ 7y agoRemember those performances are for very constrained, some might say artificial tasks of specifically image recognition into categories. Your Tesla might be 99.99% in recognizing red lights but will continue to drive straight into paper boxes in its immediate path.
- dpflan 7y agoDo you know of any good resources to learn more about this idea of the rate of improvement of perception per percentage point?
- xyhopguy 7y agohttps://en.wikipedia.org/wiki/Odds_ratio https://en.wikipedia.org/wiki/Odds_ratio odds of crash a => 97% => 3 / 100 odds of crash for b => 99.99% => 1 / 10000 improvement (odds ratio in this case) is then 300x = odds of crash for b / odds of crash for a
- FartyMcFarter 7y agoThat may be true at a given moment, but how about time? I usually don't care about the probability that I'll crash within the next millisecond, but I do care about the probability that I'll crash over a whole trip.
- KKKKkkkk1 7y agoAndrej Karpathy showed in the Tesla autonomy day how Tesla had to retrain their DNNs such that they don't get confused by bicycles mounted on vehicles. If 97% means your models get confused by something you see on the road every day, I wouldn't be too pleased about the state of the art.
- mantap 7y agoAutonomous vehicles are clearly several steps up the difficulty ladder from bread and butter tasks such as speech recognition. The progress over the last 10 years is such that some subfields have exceeded human parity while others are only just getting started. Depending on which subfield interests you, progress may be slowing or accelerating. That's why another "AI winter" is a bogus and alarmist concept. Winter for whom?
- drongoking 7y agoI'm always bemused by the idea that AI is nothing but machine learning, and ML is nothing but predictive analytics. Equating research in AI with a "boring hyperparameter search" shows how narrow it's become; saying you've "gotten 97% on a problem" refers to, obviously, classification accuracy of a model on a set of labeled instances. "Use AI/ML to solve X" means finding a way to translate X into a prediction task over feature vectors. There's an old saying "If all you have is a hammer, all your problems start to look like nails." We may see an AI winter come about simply because we run out of things to pound with our hammer.
- rytill 7y agoIf all you have is a function, everything starts to look like a mapping between two sets.
- XorNot 7y ago97% is also 3 failures out of every 100 attempts. In a lot of day to day experience I suspect humans do much better then this still.
- solveit 7y agoIt depends. DL models have legitimately achieved superhuman accuracy on many tasks. Part of this is because deep learning is incredibly effective for a certain class of problems. But part of this is because humans are remarkably bad at some problems. Humans tend to be surprisingly bad at context-stripped tasks like "identify what object this blurry image is", and "what sequence of syllables is this short audio file?". But we have countermeasures to correct for our inaccuracy. Most importantly, we understand and use context to sanity-check the hell out of our imperfect senses, and nobody has any idea how we're going to get AI to do that.
- fredguth 7y agoThe current SoTA achieve this 97% with a high cost in the number of samples. We are a living proof that it can be better. I believe that there Will be a push in achieving same generalization with less data.