4 ms·
I’ll take the risk of hurting the groupies here. But I have a genuine question: what did you learn from this talk? Like… really… what was new? or potentially us
by _l7dh 2y ago
I’ll take the risk of hurting the groupies here. But I have a genuine question: what did you learn from this talk? Like… really… what was new? or potentially useful? or insightful perhaps?
I really don’t want to sound bad-mouthed but I‘m sick of these prophetic talks (in this case, the tone was literally prophetic—with sudden high and grandiose pitches—and the content typically religious, full of beliefs and empty statements.
- _l7dh 2y agoPrecision: « pre-training data is exhausted » everyone has been saying that for a while now. The graph plotting body mass against brain mass… what does it say exactly? (where is the link to the prior point on data?). I think we would all benefit from being more critical here and stop idealizing these figures. I believe they have no more clue that any other average ML researcher on all these questions.
- XenophileJKO 2y agoThe other thing that bugged me is the built in assumption that today's model have learned everything there is to learn from the Internet corpus. This is quite easy to disprove. Both in factual retention, but also meta cognition on the context of the content.
- jebarker 2y agoYeah, exactly. A human can learning vastly more about, say, math from a much smaller quantity of text. I doubt we're anywhere close to exhausting the knowledge extraction potential from web data.
- bbor 2y agoWhich is exactly the point he's making, I believe; that simply collecting more data isn't the next step. That we've reached a local plateau in scaling ability based on corpus size. Which was assumed by pretty much everyone outside the DL elite the whole time, AFAIU
- esperent 2y agoRight, but there's nothing new in that statement. I've been hearing that we're running out of data for training AIs for two years at least.
- cma 2y agoAlso left out the Baidu scaling laws paper from 2017, and his circle has a history of a kind of citation ring type thing leaving earlier stuff out https://research.baidu.com/Blog/index-view?id=89 https://research.baidu.com/Blog/index-view?id=89
- random3 2y agoWhat everyone could learn is to check their (and their communities') assumptions from not long ago. Who saw this, who didn't. Based on this many can confirm their beliefs and others can realize they're clueless. In either case, there's something to be learned but more to be learned when you realize you were wrong.
- fullstackwife 2y agoToday I searched for early discussions about Transformer here on HN, and my observation is that back in 2019 nobody in HN comments predicted what is going to happen. It was a niche topic most of the commenters ignored, no strong opinions. Probably what we are discussing here is not the next breakthrough...
- 29athrowaway 2y agoFrom your reaction I guess you were expecting a talk about a NeurIPS 2024 paper. This is a different situation. There's the "NeurIPS 2024 Test of Time Paper Awards" where they award a historical paper. In this case, a paper from 2014 was awarded and his talk is about that and why it passed the test of time. https://blog.neurips.cc/2024/11/27/announcing-the-neurips-2024-test-of-time-paper-awards/ https://blog.neurips.cc/2024/11/27/announcing-the-neurips-20... The title chosen for the HN submission leaves out that important context. So that's why you are disappointed now.
- p1esk 2y agoI am also disappointed, and I have not missed the context. The talk is empty for anyone who follows the field for more than two years, and especially for those who are familiar with his 2014 paper. Yes, he had amazing insight and intuition behind modern LLM breakthroughs, and yes, he probably earned the right to sound "prophetic", but he could have provided some interesting personal anecdotes about how the paper was written, or some fresh ideas in "What Comes Next" section of his presentation.
- 29athrowaway 2y agoTrue. The entire thing was basically "neurons go brrrr".
- abetusk 2y agoI'll give my take: * Before the current renaissance of neural networks (pre ~2014ish), it was unclear that scaling would work. That is, simple algorithms on lots of data. The last decade has pretty much addressed that critique and it's clear that scaling does work to a large extent, and spectacularly so. * Much of the current neural network models and research are geared towards "one-shot" algorithms, doing pattern matching and giving an immediate result. Contrast this with search which needs to do inference time compute or search. * The exponential increase in power means that neural network models are quickly sponging up as much data as they can find and we're quickly running into the limits of science, art and other data that humans have created in the last 5k years or so. * Sutskever points out, as an analogy, nature has created a better model for humans (the brain to mass ratio for animals) with hominids finding more efficient compute than other animals, even ones with much larger brains and neuron count. * Sutskever is advocating for better models, presumably focusing on inference time computer more. In some sense, we're coming a bit full circle where people who were advocating for pure scaling (simple algorithms + lots of data) for learning are now advocating for better algorithms, presumably with a focus on inference time compute (read: search). I agree that it's a little opaque, especially for people who haven't been paying attention to past and current research, but this message seems pretty clear to me. Noam Brown had a talk recently titled "Parables on the Power of Planning in AI" [0] which addresses this point more head on. I will also point out that the scaling hypothesis is closely related to "The Bitter Lesson" by Rich Sutton [1]. Most people focus on the "learning" aspect of scaling but "The Bitter Lesson" very clearly articulates learning and search as the methods most amenable to compute. From Sutton: """ ... Search and learning are the two most important classes of techniques for utilizing massive amounts of computation in AI research. ... """ [0] https://youtube.com/watch?v=eaAonE58sLU https://youtube.com/watch?v=eaAonE58sLU [1] https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
- abetusk 2y agoHere's a more pithy summary: "We've made a copy of the internet, run current state of the art methods on it and GPT-O1 is the best we can do. We need better (inference/search) algorithms to make progress"
- eldenring 2y agoHe mentions this in the video, but the talk is specifically tailored for the "Test of Time" award. This being his 3rd year in a row recieving the award, I think he's earned permission to speak prophetically.
- tylerchilds 2y agodid the last two prophecies come true?
- katamari-damacy 2y agoHe also said in an interview with Jensen soon after ChatGPT's launch that "before 2003, machines couldn't learn" ... LOL. I was stunned when I heard that nonsensical assertion. I guess it depends on his definition of "learn" ...
- bbor 2y ago"learn" is usually used in opposition to "taught", which refers to "expert systems"-type engineering; in other words, providing data and a success heuristic and asking it to devise its own optimal strategies vs. providing strategies hand-designed by humans. Obviously Perceptrons came out well before 2003, but I don't think it's necessarily out of line to say that they had limited efficacy before then, both for theoretical and compute reasons. But maybe I'm misunderstanding your criticism?
- remexre 2y agoILP goes back to the 80s, and was used to do drug discovery in the 90s. Bayes nets go back to the 80s as well.
- YeGoblynQueenne 2y ago"ILP" (as in Inductive Logic Programming not Integer Linear Programming) was first named in 1991 in a paper by Stephen Muggleton ("Inductive Logic Programming and Progol). The paper properly launched the field and generated a great deal of excitement at the time. There were precursors. At least Ehud Shapiro's doctoral thesis ("Automated Debugging") in the 1980's and Gordon Plotkin's doctoral thesis in the 1970's ("Automated Methods of Inductive Inference"). Sorry for not giving the exact years off the top of my head but I think it was 1983 and 1976, respectively. The point you are making is very right however because modern machine learning as a field started in the 1980's with the fall of expert systems, in fact it basically started as an effort to overcome one of the major limitations of expert systems, the so-called "knowledge acquisition bottleneck", which is to say, the difficulty of creating and maintaining huge databases of expert knowledge (in the form of production rules). In any case the seminal textbook in the field for the first 20 years, Tom Mitchell's Machine Learning came out in 1997 (https://www.cse.iitb.ac.in/~cs725/notes/slides/tom_mitchell/mlbook.html https://www.cse.iitb.ac.in/~cs725/notes/slides/tom_mitchell/...) and includes probabilistic, neural-net based and symbolic, logic-based approaches. So not only machines could "learn" way before 2003 but they could also learn in many different ways than what Ilya Sutskever means. We can go further back, to Donald Michie's 1961 MENACE (the first Reinforcement Learning system, implemented on a computer made of matchboxes with coloured beads used to encode state) and Arthur Samuel's 1959 checkers player (a paper on which gave the name to the field of machine learning). Lots of learning all over the place long, long before 2003.
- deleted 2y ago[deleted]
- wills_forward 2y agoIt was funny to hear the same guy warning LMMs were getting too powerful now talking about the limits of available original training data.
- sashank_1509 2y agoReminds me a little of a Feynman quote. Once physicists win a Nobel prize, their output falls because now they no longer can work on small problems. Everything they work on must be grand. Every speech they give must discover secrets of the universe. Seems to fit Ilya.
- eli_gottlieb 2y agoIt's a test-of-time talk. The point was to let him have a moment to brag and celebrate about the success of GANs, 10 years later.