15 ms·
Ilya Sutskever: We're moving from the age of scaling to the age of research
- el_jay 10mo agoSuggest tagline: “Eminent thought leader of world’s best-funded protoindustry hails great leap back to the design stage.”
- rglover 10mo agoHahahahahaha okay that was good.
- andy_ppp 10mo agoSo is the translation endless scaling has stopped being as effective?
- deleted 10mo ago[deleted]
- jsheard 10mo agoThe translation is that SSI says that SSIs strategy is the way forward so could investors please stop giving OpenAI money and give SSI the money instead. SSI has not shown anything yet, nor does SSI intend to show anything until they have created an actual Machine God, but SSI says they can pull it off so it's all good to go ahead and wire the GDP of Norway directly to Ilya.
- gessha 10mo agoIt’s a snake oil salesman’s world.
- aunty_helen 10mo agoIf we take AGI as a certainty, ie we think we can achieve AGI using silicon, then Ilya is one of the best bets you can take if you are looking to invest in this space. He has a history and he's motivated to continue working on this problem. If you think that AGI is not possible to achieve, then you probably wouldn't be giving anyone money in this space.
- deleted 10mo ago[deleted]
- bossyTeacher 10mo agoThis hinges on his company achieving AGI while he's still alive. He's 38 years old. He has about 4 decades to deliver AGI in his lifetime. When he dies, there is no guarantee whoever takes over will share his values. "If you think that AGI is not possible to achieve, then you probably wouldn't be giving anyone money in this space." If you think other people think AGI is possible, you sell them shovels and ready yourself for a shovel market dip in the near future. Strike while the iron is hot.
- shwaj 10mo agoAre you asking whether the whole podcast can be boiled down to that translation, or whether you can infer/translate that from the title? If the former, no. If the latter, sure, approximately.
- Animats 10mo agoIt's stopped being cost-effective. Another order of magnitude of data centers? Not happening. The business question is, what if AI works about as well as it does now for the next decade or so? No worse, maybe a little better in spots. What does the industry look like? NVidia and TSMC are telling us that price/performance isn't improving through at least 2030. Hardware is not going to save us in the near term. Major improvement has to come from better approaches. Sutskever: "I think stalling out will look like…it will all look very similar among all the different companies. It could be something like this. I’m not sure because I think even with stalling out, I think these companies could make a stupendous revenue. Maybe not profits because they will need to work hard to differentiate each other from themselves, but revenue definitely." Somebody didn't get the memo that the age of free money at zero interest rates is over. The "age of research" thing reminds me too much of mid-1980s AI at Stanford, when everybody was stuck, but they weren't willing to admit it. They were hoping, against hope, that someone would come up with a breakthrough that would make it work before the house of cards fell apart. Except this time everything costs many orders of magnitude more to research. It's not like Sutskever is proposing that everybody should go back to academia and quietly try to come up with a new idea to get things un-stuck. They want to spend SSI's market cap of $32 billion on some vague ideas involving "generalization". Timescale? "5 to 20 years". This is a strange way to do corporate R&D when you're kind of stuck. Lots of little and medium sized projects seem more promising, along the lines of Google X. The discussion here seems to lean in the direction of one big bet. You have to admire them for thinking big. And even if the whole thing goes bust, they probably get to keep the house and the really nice microphone holder.
- energy123 10mo agoThe ideas likely aren't vague at all given who is speaking. I'd bet they're extremely specific. Just not transparently shared with the public because it's intellectual property.
- giardini 10mo agoWhat kind of ideas would be intellectual property that was not shared? Isn't every part of LLMs, except the order of processes, publicly known ? Is there some magic algorithm previously unrevealed and held secret by a cabal of insiders?
- Quothling 10mo agoNot really, but there is a finite amount of data to train models on. I found it rather interesting to hear him talk about how Gemini has been better at getting results out of the data than their competition, and how this is the first insights into a new way of dealing with how they train models on the same data to get different results. I think the title is an interesting thing, because the scaling isn't about compute. At least as I understand it, what they're running out of is data, and one of the ways they deal with this, or may deal with this, is to have LLM's running concurrently and in competition. So you'll have thousands of models competing against eachother to solve challenges through different approaches. Which to me would suggest that the need for hardware scaling isn't about to stop.
- imiric 10mo agoThe translation to me is: this cow has run out of milk. Now we actually need to deliver value, or the party stops.
- giardini 10mo agoI'll be convinced LLMs are a reasonable approach to AI when an LLM can give reasonable answers after being trained with approximately the same books and classes in school that I was once I completed my college education.
- alex43578 10mo agoI'll be convinced cars are a reasonable approach to transportation when it can take me as far as a horse can on a bale of hay.
- Jeff_Brown 10mo agoThat is such a beautiful analogy that now I will read your other comments.
- snapcaster 10mo agoWhy do you think this standard you're applying is reasonable or meaningful?
- giardini 10mo agoFor the same reason anyone would: if an AI can reason to a human level after having been educated in a manner similar to a human then it is likely that we [the educators] have captured something akin to human intelligence.
- gizmodo59 10mo agoEven as criticism targets major model providers, his inability to answer clearly about revenue & dismissing it as a future concern reveals a great deal about today's market. It's remarkable how effortlessly he, Mira, and others secure billions, confident they can thrive in such an intensely competitive field. Without a moat defined by massive user bases, computing resources, or data, any breakthrough your researchers achieve quickly becomes fair game for replication. May be there will be new class of products, may be there is a big lock-in these companies can come up with. No one really knows!
- markus_zhang 10mo agoTBH if you truly believe you are in the frontier of AI you probably don’t need to care too much about those numbers. Yes corporations need those numbers, but those few humans are way more valuable than any numbers out there. Of course, only when others believe that they are in the frontier too.
- SilverElfin 10mo agoMira was a PM who somehow was at the right place at the right time. She isn’t actually an AI expert. Ilya however, is. I find him to be more credible and deserving in terms of research investment. That said, I agree that revenue is important and he will need a good partner (another company maybe) to turn ideas into revenue at some point. But maybe the big players like Google will just acquire them on no revenue to get access to the best research, which they can then turn into revenue.
- fragmede 10mo agoThat’s kind of a shitty way to put it. Mira wasn’t a PM at OpenAI. She was CTO and before that VP of Engineering. Prior to OpenAI she was an engineer at Tesla on the Model X and Leap Motion. You’re right that she’s not a published ML researcher like Ilya, but "right place, right time" undersells leading the team that shipped ChatGPT, DALL-E, and GPT-4.
- Nextgrid 10mo ago
- SilverElfin 10mo agoHow did Dwarkesh manage to build a brand that can attract famous people to his podcast? He didn’t have prior fame from something else in research or business, right? Curious if anyone knows his growth strategy to get here.
- piker 10mo agoSeems like he’s Lex without the Rogan association so hardcore liberal folks can listen without having to buy morality offsets. He’s good, and he’s filling a void in an established underserved genre is my take.
- camillomiller 10mo agoFridman is a morally broken grifter, who just built a persona and a brand on proven lies, claiming an association with MIT that was de facto non-existent. Not wanting to give the guy recognition is not a matter of being liberal or conservative, but just interested in truthfulness.
- wahnfrieden 10mo agoPatel takes anticommunism to such an extreme that he repeatedly brings up and speculates (despite being met with repudiation by even the staunchest anticommunist of guests) whether naziism is preferable, that Hitler should have the war against Soviets, that the US should have collaborated with Hitler to defeat communism, and that the enduring spread of naziism would have been a good tradeoff to make.
- chermi 10mo agoWhere does he say this?
- wahnfrieden 10mo agothe Sarah Paine interviews
- oytis 10mo agoAges just keep flying by
- scotty79 10mo agoTranslation: Free lunch of getting results just by throwing money at the problem is over. Now for the first time in years we actually need to think what we are doing and firgure out why things that work, do work. Somehow, despite being vastly overpaid I think AI researchers will turn out to be deeply inadequate for the task. As they have been during the last few AI winters.
- eats_indigo 10mo agodid he just say locomotion came from squirrels
- jonny_eh 10mo agotimestamp?
- FergusArgyll 10mo agoI think he was referencing something Richard Sutton said (iirc); along the lines of "If we can get to the intelligence of a squirrel, we're most of the way there"
- Animats 10mo agoI've been saying that for decades now. My point was that if you could get squirrel-level common sense, defined as not doing anything really bad in the next thirty seconds while making some progress on a task, you were almost there. Then you can back-seat drive the low-level system with something goal-oriented. I once said that to Rod Brooks, when he was giving a talk at Stanford, back when he had insect-level robots and was working on Cog, a talking head. I asked why the next step was to reach for human-level AI, not mouse-level AI. Insect to human seemed too big a jump. He said "Because I don't want to go down in history as the creator of the world's greatest robot mouse". He did go down in history as the creator of the robot vacuum cleaner, the Roomba.
- alyxya 10mo agoThe impactful innovations in AI these days aren't really from scaling models to be larger. It's more concrete to show higher benchmark scores, and this implies higher intelligence, but this higher intelligence doesn't necessarily translate to all users feeling like the model has significantly improved for their use case. Models sometimes still struggle with simple questions like counting letters in a word, and most people don't have a use case of a model needing phd level research ability. Research now matters more than scaling when research can fix limitations that scaling alone can't. I'd also argue that we're in the age of product where the integration of product and models play a major role in what they can do combined.
- TheBlight 10mo ago"Scaling" is going to eventually apply to the ability to run more and higher fidelity simulations such that AI can run experiments and gather data about the world as fast and as accurately as possible. Pre-training is mostly dead. The corresponding compute spend will be orders of magnitude higher.
- alyxya 10mo agoThat's true, I expect more inference time scaling and hybrid inference/training time scaling when there's continual learning rather than scaling model size or pretraining compute.
- TheBlight 10mo agoSimulation scaling will be the most insane though. Simulating "everything" at the quantum level is impossible and the vast majority of new learning won't require anything near that. But answers to the hardest questions will require as close to it as possible so it will be tried. Millions upon millions of times. It's hard to imagine.
- emporas 10mo ago>Pre-training is mostly dead. I don't think so. Serious attempts for producing data specifically for training have not being achieved yet. High quality data I mean, produced by anarcho-capitalists, not corporations like Scale AI using workers, governed by laws of a nation etc etc. Don't underestimate the determination of 1 million young people to produce within 24 hours perfect data, to train a model to vacuum clean their house, if they don't have to do it themselves ever again, and maybe earn some little money on the side by creating the data. The other part of the comment I agree.
- a_state_full 10mo ago[dead]
- deleted 10mo ago[deleted]
- Herring 10mo agohttps://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com... He’s wrong we still scaling, boys.
- epistasis 10mo agoThat blog post is eight months old. That feels like pretty old news in the age of AI. Has it held since then?
- conception 10mo agoIt looks like it’s been updated as it has codex 5.1 max on it
- deleted 10mo ago[deleted]
- rockinghigh 10mo agoYou should read the transcript. He's including 2025 in the age of scaling. > Maybe here’s another way to put it. Up until 2020, from 2012 to 2020, it was the age of research. Now, from 2020 to 2025, it was the age of scaling—maybe plus or minus, let’s add error bars to those years—because people say, “This is amazing. You’ve got to scale more. Keep scaling.” The one word: scaling. > But now the scale is so big. Is the belief really, “Oh, it’s so big, but if you had 100x more, everything would be so different?” It would be different, for sure. But is the belief that if you just 100x the scale, everything would be transformed? I don’t think that’s true. So it’s back to the age of research again, just with big computers.
- Herring 10mo agoNope, Epoch.ai thinks we have enough to scale till 2030 at least. https://epoch.ai/blog/can-ai-scaling-continue-through-2030 https://epoch.ai/blog/can-ai-scaling-continue-through-2030 ^ /_\ ***
- deleted 10mo ago
- jmkni 10mo agoThis reveals a new source of frustration, I can't watch this in work, and I don't want to read and AI generated summary so...?
- cheeseblubber 10mo agoThere is a transcript of the entire conversation if you scroll down a little
- delichon 10mo agoIf the scaling reaches the point at which the AI can do the research at all better than natural intelligence, then scaling and research amount to the same thing, for the validity of the bitter lesson. Ilya's commitment to this path is a statement that he doesn't think we're all that close to parity.
- pron 10mo agoI agree with your conclusion but not with your premise. To do the same research it's not enough to be as capable as a human intelligence; you'd need to be as capable as all of humanity combined. Maybe Albert Einstein was smarter than Alexander Fleming, but Einstein didn't discover penicillin. Even if some AI was smarter than any human being, and even if it devoted all of its time to trying to improve itself, that doesn't mean it would have better luck than 100 human researchers working on the problem. And maybe it would take 1000 people? Or 10,000?
- delichon 10mo agoI'm afraid that turning sand and sunlight into intelligence is so much more efficient than doing that with zygotes and food, that people will be quickly out scaled. As with chess, we will shift from collaborators to bystanders.
- pron 10mo agoWho's "we", though, and aren't virtually all of us already bystanders in that sense? I have virtually zero power to shape world events and even if I want to believe that what I do isn't entirely negligible, someone else could do it, possibly better. I live in one of the largest, most important metropolises in the world, and even as a group, everything the entire population of my city does is next to nothing compared to everything being done in the world. As the world has grown, my city's share of it has been falling. If a continent with 20 billion people on it suddenly appeared, the output of my entire country will be negligible; would it matter if they were robots? In the grand scheme of things, my impact on the world is not much greater than my cat's, and I think he's quite content overall. There are many people more accomplished than me (although I don't think they're all smarter); should I care if they were robots? I may be sad that I won't be able to experience what the robots experience, but there are already many people in the world whose experience is largely foreign to mine. And here's a completely way of looking at it, since I won't lieve forever. A successful species eventually becomes extinct - replaced by its own eventual offspring. Homo erectus are extinct, as they (eventually) evolved into homo sapiens. Are you the "we" of homo erectus or a different "we"? If all that remains from homo sapiens some time in the future is some species of silicon-based machines, machina sapiens, that "we" create, will those beings not also be "us"? After all, "we" will have been their progenitors in not-too-dissimilar a way to how the home erectus were ours (the difference being that we will know we have created a new distinct species). You're probably not a descendent of William Shakespeare's, so what makes him part of the same "we" that you belong to, even though your experience is in some ways similar to his and in some ways different. Will not a similar thing make the machines part of the same "we"?
- itissid 10mo agoAll coding agents are geared towards optimizing one metric, more or less, getting people to put out more tokens — or $$$. If these agents moved towards a policy where $$$ were charged for project completion + lower ongoing code maintenance cost, moving large projects forward, _somewhat_ similar to how IT consultants charge, this would be a much better world. Right now we have chaos monkey called AI and the poor human is doing all the cleanup. Not to mention an effing manager telling me you now "have" AI push 50 Features instead of 5 in this cycle.
- torben-friis 10mo ago>this would be a much better world. Would it? We’d close one of the few remaining social elevators, displace higher educated people by the millions and accumulate even more wealth at the top of the chain. If LLMs manage similar results to engineers and everyone gets free unlimited engineering, we’re in for the mother of all crashes. On the other hand, if LLMs don’t succeed we’re in for a bubble bust.
- itissid 10mo ago> Would it? As compared to now. Yes. The whole idea is that if you align AI to human goals of meeting project implementation + maintenance only then can it actually do something worthwhile. Instead now its just a bunch of of middle managers yelling you to do more and laying off people "because you have AI". If projects getting done a lot of actual wealth could be actually generated because lay people could implement things that go beyond the realm of toy projects.
- hn_acc1 10mo agoYou think that you will be ALLOWED to continue to use AI for free once it can create a LOT of wealth? Or will you have to pay royalties? The rich CEOs don't want MORE competition - they want LESS competition for being rich. I'm sure they'll find a way to add a "any vibe-coded business owes us 25% royalties" clause any day now, once the first big idea makes some $$. If that ever happens. They're NOT trying to liberate "lay people" to allow them to get rich using their tech, and they won't stand for it.
- wrs 10mo ago"The idea that we’d be investing 1% of GDP in AI, I feel like it would have felt like a bigger deal, whereas right now it just feels...[normal]." Wow. No. Like so many other crazy things that are happening right now, unless you're inside the requisite reality distortion field, I assure you it does not feel normal. It feels like being stuck on Calvin's toboggan, headed for the cliff.
- hn_acc1 10mo agoAgreed.
- johnxie 10mo agoI don’t think he meant scaling is done. It still helps, just not in the clean way it used to. You make the model bigger and the odd failures don’t really disappear. They drift, forget, lose the shape of what they’re doing. So “age of research” feels more like an admission that the next jump won’t come from size alone.
- energy123 10mo agoIt still does help in the clean way it used to. The problem is that the physical world is providing more constraints like lack of power and chips and data. Three years ago there was scaling headroom created by the gaming industry, the existing power grid, untapped data artefacts on the internet, and other precursor activities.
- kmmlng 10mo agoThe scaling laws are also power laws, meaning that most of the big gains happen early in the curve, and improvements become more expensive the further you go along.
- _giorgio_ 10mo agoScaling is not over, there's no wall. Oriol Vinyals VP of Gemini research https://x.com/OriolVinyalsML/status/1990854455802343680?t=oCF94w9t4bURhoAjv0av1g&s=19 https://x.com/OriolVinyalsML/status/1990854455802343680?t=oC...
- neonate 10mo agohttps://xcancel.com/OriolVinyalsML/status/1990854455802343680 https://xcancel.com/OriolVinyalsML/status/199085445580234368...?
- JohnnyMarcone 10mo agoHe didn't say it's over, just that continued scaling won't be transformational.
- _giorgio_ 10mo agoOriol Vinyals said that.
- lvl155 10mo agoYou have LLMs but you also need to model actual intelligence, not its derivative. Reasoning models are not it.
- xeckr 10mo agoHe is, of course, incentivised to say that.
- londons_explore 10mo ago> These models somehow just generalize dramatically worse than people. It's a very fundamental thing My guess is we'll discover that biological intelligence is 'learning' not just from your experience, but that of thousands of ancestors. There are a few weak pointers in that direction. Eg. A father who experiences a specific fear can pass that fear to grandchildren through sperm alone. [1]. I believe this is at least part of the reason humans appear to perform so well with so little training data compared to machines. [1]: https://www.nature.com/articles/nn.3594 https://www.nature.com/articles/nn.3594
- deleted 10mo ago[deleted]
- HarHarVeryFunny 10mo agoFrom both an architectural and learning algorithm perspective, there is zero reason to expect an LLM to perform remotely like a brain, nor for it to generalize beyond what was necessary for it to minimize training errors. There is nothing in the loss function of an LLM to incentivize it to generalize. However, for humans/animals the evolutionary/survival benefit of intelligence, learning from experience, is to correctly predict future action outcomes and the unfolding of external events, in a never-same-twice world. Generalization is key, as is sample efficiency. You may not get more than one or two chances to learn that life-saving lesson. So, what evolution has given us is a learning architecture and learning algorithms that generalize well from extremely few samples.
- jebarker 10mo ago> what evolution has given us is a learning architecture and learning algorithms that generalize well from extremely few samples. This sounds magical though. My bet is that either the samples aren’t as few as they appear because humans actually operate in a constrained world where they see the same patterns repeat very many times if you use the correct similarity measures. Or, the learning that the brain does during human lifetime is really just a fine-tuning on top of accumulated evolutionary learning encoded in the structure of the brain.
- l5870uoo9y 10mo ago> These models somehow just generalize dramatically worse than people. The whole mess surrounding Grok's ridiculous overestimation of Elon's abilities in comparison to other world stars, did not so much show Grok's sycophancy or bias towards Elon, as it showed that Grok fundamentally cannot compare (generalize) or has a deeper understanding of what the generated text is about. Calling for more research and less scaling is essentially saying; we don't know where to go from here. Seems reasonable.
- radicaldreamer 10mo agoI think the problem with that is that Grok has likely been prompted to do that in the system prompt or some prompts that get added for questions about Elon. That doesn't reflect on the actual reasoning or generalization abilities of the underlying model most likely.
- asolove 10mo agoYes it does. Today on X, people are having fun baiting Grok into saying that Elon Musk is the world’s best drinker of human piss. If you hired a paid PR sycophant human, even of moderate intelligence, it would know not to generalize from “say nice things about Elon” to “say he’s the best at drinking piss”.
- phs318u 10mo agoTrue. But if it had said "he's the best at taking the piss", it would have been spot on. https://en.wikipedia.org/wiki/Taking_the_piss https://en.wikipedia.org/wiki/Taking_the_piss
- l5870uoo9y 10mo agoYou can also give AI models Nobel-prize winning world literature and ask why this is bad and they will tear apart the text, without ever thinking "wait this is some of the best writing produced by man".
- 10mo ago
- tmp10423288442 10mo agoHe's talking his book. Doesn't mean he's wrong, but Dwarkesh is now big enough that you should assume every big name there is talking their book.
- delichon 10mo agoHere's a world class scientist here not because we had a hole in the schedule or he happened to be in town, but to discuss this subject that he thought and felt about so deeply that he had to write a book about it. That's a feature not a bug.
- giardini 10mo ago"Here's a world class scientist here not because we had a hole in the schedule or he happened to be in town, but to discuss this subject that he " had invested himself so fully personally and financially that, should it fail, he would be ruined. FTFY
- NaomiLehman 10mo agoruined how?
- gloftus 10mo agoEgo death.
- alexnewman 10mo agoA lot more of human intelligence is hard coded
- JimmyBuckets 10mo agoI respect Ilya hugely as a researcher in ML and quite admire his overall humility, but I have to say I cringed quite a bit at the start of this interview when he talks about emotions, their relative complexity, and origin. Emotion is so complex, even taking all the systems in the body that it interacts with. And many mammals have very intricate socio-emotional lives - take Orcas or Elephants. There is an arrogance I have seen that is typical of ML (having worked in the field) that makes its members too comfortable trodding into adjacent intellectual fields they should have more respect and reverence for. Anyone else notice this? It's something physicists are often accused of also.
- deleted 10mo ago[deleted]
- fumeux_fume 10mo agoYeah, that's bothered me as well. Andrej Karpathy does this all the time when he talks about the human brain and making analogies to LLMs. He makes speculative statements about how the human brain works as though it's established fact.
- mips_avatar 10mo agoAndrej does use biological examples, but he's a lot more cautious about biomimicry, and often uses biological examples to show why AI and bio are different. Like he doesn't believe that animals use classical RL because a baby horse can walk after 5 minutes which definitely wasn't achieved through classical RL. He doesn't pretend to know how a horse developed that ability, just that it's not classical RL. A lot of Ilya's takes in this interview felt like more of a stretch. The emotions and LLM argument felt like of like "let's add feathers to planes because birds fly and have feathers". I bet continual learning is going to have some kind of internal goal beyond RL eval functions, but these speculations about emotions just feel like college dorm discussions. The thing that made Ilya such an innovator (the elegant focus on next token prediction) was so simple, and I feel like his next big take is going to be something about neuron architecture (something he eluded to in the interview but flat out refused to talk about).
- 10mo ago
- river_otter 10mo agoOne thing from the podcast that jumped out to me was the statement that in pre training "you don't have to think closely about the data". Like I guess the success of pre training supports the point somewhat but it feels to me slightly opposed to Karpathy talking about what a large percentage of pretraining data is complete garbage. I guess I would hope that more work in cleaning the pre training data would result in stronger and more coherent base models.
- orbital-decay 10mo ago>You could actually wonder that one possible explanation for the human sample efficiency that needs to be considered is evolution. Evolution has given us a small amount of the most useful information possible. It's definitely not small. Evolution performed a humongous amount of learning, with modern homo sapiens, an insanely complex molecular machine, as a result. We are able to learn quickly by leveraging this "pretrained" evolutionary knowledge/architecture. Same reason as why ICL has great sample efficiency. Moreover, the community of humans created a mountain of knowledge as well, communicating, passing it over the generations, and iteratively compressing it. Everything that you can do beyond your very basic functions, from counting to quantum physics, is learned from the 100% synthetic data optimized for faster learning by that collective, massively parallel, process. It's pretty obvious that artificially created models don't have synthetic datasets of the quality even remotely comparable to what we're able to use.
- FloorEgg 10mo agoAren't you agreeing with his point? The process of evolution distilled down all that "humongous" amount to what is most useful. He's basically saying our current ML methods to compress data into intelligence can't compare to billions of years of evolution. Nature is better at compression than ML researchers, by a long shot.
- samrus 10mo agoSample efficiency isnt the ability to distill alot of data into good insights. Its the ability to get good insights from less data. Evolution didnt do that it had a lot of samples to get to where it did
- FloorEgg 10mo ago> Sample efficiency isnt the ability to distill alot of data into good insights Are you claiming that I said this? Because I didn't.... There's two things going on. One is compressing lots of data into generalizable intelligence. The other is using generalized intelligence to learn from a small amount of data. Billions of years and all the data that goes along with it -> compressed into efficient generalized intelligence -> able to learn quickly with little data
- roman_soldier 10mo agoScaling got us here and it wasn't obvious that it would produce the results we have now, so who's to say sentience won't emerge from scaling another few orders of magnitude? Of course there will always be research to squeeze more out of the compute, improving efficiency and perhaps make breakthroughs.
- hn_acc1 10mo agoAnother few orders of magnitude? Like 100-1000x more than we're already doing? Got a few extra suns we can tap for energy? And a nanobot army to build various power plants? There's no way to do 1000x of what we're already doing any time soon.
- roman_soldier 10mo ago10x to 100x and order of magnitude is a factor of 10.
- measurablefunc 10mo agoI didn't learn anything new from this. What exactly has he been researching this entire time?
- xoac 10mo agoBest time to sell his ai portfolio
- pxc 10mo agoIf "Era of Scaling" means "era of rapid and predictable performance improvements that easily attract investors", it sounds a lot like "AI summer". So... is "Era of Research" a euphemism for "AI winter"?
- zombiwoof 10mo ago[dead]
- techblueberry 10mo agoYes
- hiddencost 10mo agoThat presumes that performance improvements are necessary for commercialization. From what I've seen the models are smart enough, what we're lacking is the understanding and frameworks necessary to use them well. We've barely scratched the surface on commercialization. I'd argue there are two things coming: -> Era of Research -> Era of Engineering Previous AI winters happened because we didn't have a commercially viable product, not because we weren't making progress.
- ares623 10mo agoThe labs can't just stop improvements though. They made promises. And the capacity to run the current models are subsidized by those promises. If the promise is broken, then the capacity goes with it.
- wmf 10mo agoMaybe those promises can be better fulfilled with products based on current models.
- selectodude 10mo ago> the capacity goes with it. Sort of. The GPUs exist. Maybe LLM subs can’t pay for electricity plus $50,000 GPUs, but I bet after some people get wiped out, there’s a market there.
- mrcwinn 10mo agoHe also suggested the "revenue opportunities" would reveal themselves later, given enough investment. I have the same plan if anyone is interested.
- venturecruelty 10mo agoWhen are we going to call out these charlatans for the frauds that they are?
- bilsbie 10mo agoI don’t think either of those ages is correct. I’d like to see the age of efficiency and bringing decent models to personal devices.
- zerosizedweasle 10mo agoSure but that will also be a research age.
- deleted 10mo ago[deleted]
- nakamoto_damacy 10mo agoIlya "Undercover Genocide Supporter" Sutskever... ¯\_(ツ)_/¯
- nsoonhui 10mo agoOne thing I’m curious about is this: Ilya Sutskever wants to build Safe Superintelligence, but he keeps his company and research very secretive. Given that building Safe Superintelligence is extraordinarily difficult — and no single person’s ideas or talents could ever be enough — how does secrecy serve that goal?
- NateEag 10mo agoIf he (or his employees) are actually exploring genuinely new, promising approaches to AGI, keeping them secret helps avoid a breakneck arms race like the one LLM vendors are currently engaged in. Situations like that do not increase all participants' level of caution.
- 4b11b4 10mo agoDoesn't sound like you listened to the interview. He addresses this and says he may make releases that would be otherwise held back because he believes it's important for developments to be seen by the public.
- giardini 10mo agoNo reasonable person would do that! That is, if you had the key to AI, you wouldn't share it and you would do everything possible to prevent it's dissemination. Meanwhile you would use it to conquer the world! Bwahahahaaaah!
- photochemsyn 10mo agoOpen source the training corpus. Isn't this humanity's crown jewels? Our symbolic historical inheritance, all that those who came before us created? The net informational creation of the human species, our informational glyph, expressed as weights in a model vaster than anything yet envisionaged, a full vectorial representation of everything ever done by a historical ancestor... going right back to LUCA, the Last Universal Common Ancestor? Really the best way to win with AI is use it to replace the overpaid executives and the parasitic shareholders and investors. Then you put all those resources into cutting edge R & D. Like Maas Biosciences. All edge. (just copy and paste into any LLM then it will be explained to you).
- joelthelion 10mo ago> When do you expect that impact? I think the models seem smarter than their economic impact would imply. > Yeah. This is one of the very confusing things about the models right now. As someone who's been integrating "AI" and algorithms into people's workflows for twenty years, the answer is actually simple. It takes time to figure out how exactly to use these tools, and integrate them into existing tooling and workflows. Even if the models don't get any smarter, just give it a few more years and we'll see a strong impact. We're just starting to figure things out.
- otabdeveloper4 10mo ago> the models seem smarter than their economic impact would imply Key word is "seem".
- zeroonetwothree 10mo agoKind of like how some humans seem smart during the interview but then are incapable of actually doing anything properly.
- coder-3 10mo agoAs someone who is building an LLM-powered product on the side, using AI coding agents to help with development of said LLM-powered product and for my day job, and has a long-tail of miscellaneous uses for AI, I suspect you're right.
- tim333 10mo agoBeyond that the smartness is very patchy. They can do math problems beyond 99% of humans but lack the common sense understanding to take over most jobs.
- ACCount37 10mo agoMost jobs involve complex long term tasks - which isn't something that's natural to LLMs.
- 10mo ago
- wartywhoa23 10mo agoA steady progress implies transitioning between ages at least ⌊(year-2020)^2/10⌋ times a year, and entering at least one new era once in a decade.
- fuzzfactor 10mo agoThis AI stuff is really taking off fast. And hasn't Ilya been on the cutting edge for a while now? I mean, just a few hours earlier there was a dupe of this artice with almost no interest at all, and now look at it :) This was my feelings way back then when it comes to major electronics purchases: Sometimes you grow to utilize the enhanced capabilities to a greater extent than others, and time frame can be the major consideration. Also maybe it's just a faster processor you need for your own work, or OTOH a hundred new PC's for an office building, and that's just computing examples. Usually, the owner will not even explore all of the advantages of the new hardware as long as the purchase is barely justified by the original need. The faster-moving situations are the ones where fewest of the available new possibilities have a chance to be experimented with. IOW the hardware gets replaced before anybody actually learns how to get the most out of it in any way that was not foreseen before purchase. Talk about scaling, there is real massive momentum when it's literally tonnes of electronics. Like some people who can often buy a new car without ever utilizing all of the features of their previous car, and others who will take the time to learn about the new internals each time so they make the most of the vehicle while they do have it. Either way is very popular, and the hardware is engineered so both are satisfying. But only one is "research". So whether you're just getting a new home entertainment center that's your most powerful yet, or kilos of additional PC's that would theoretically allow you to do more of what you are already doing (if nothing else), it's easy for anybody to purchase more than they will be able to technically master or even fully deploy sometimes. Anybody know the feeling? The root problem can be that the purchasing gets too far ahead of the research needed to make the most of the purchase :\ And if the time & effort that can be put in is at a premium, there will be more waste than necessary and it will be many times more costly. Plus if borrowed money is involved, you could end up with debts that are not just technical. Scale a little too far, and you've got some research to catch up on :)
- myrmidon 10mo agoI really liked this podcasts; the host generally does a really good job, his series with Sarah Paine on geopolitics is also excellent (can find it on youtube).
- highfrequency 10mo agoGreat respect for Ilya, but I don’t see an explicit argument why scaling RL in tons of domains wouldn’t work.
- kubb 10mo agoNot sure why they care about his opinion and discard yours. They’re just as valid and well informed.
- never_inline 10mo agoI think that scaling RL for all common domains is already done to death by big labs.
- anthonypasq 10mo agodoesnt RL by definition not generalize? thats Ilya's entire criticism of the current paradigm
- podgorniy 10mo agoBack to drawing board! -- ~Don't mind all those trillions of unreturned investments. Taxpayers will bail out the too-bog-to-fail ones.~
- zombot 10mo agoShouldn't research have come first? Am I making any sense?
- shaism 10mo agoIlya mentioned in the video that 2012 and 2020 was the “Age of Research”, followed by the “Age of Scaling” from 2020 to 2025. Now, we are about to reenter the “Age of Research”.
- bentt 10mo agoIs this like if everyone suddenly got 1gb fiber connections in 1996? We put money into the thing we know (infra), but there's no youtube, netflix, dropbox, etc etc etc. Instead we're still loading static webpages with progressive jpegs and it's like... a waste?
- twodave 10mo agoI’d settle for the “age of being able to point the LLM at the entire codebase, describe a new feature, and see it implemented based on the patterns and idioms already present in that codebase”. My impression is the only thing between here and there is context size.
- Mars008 10mo agoIf I remember correctly after leaving OpenAI with a bang Ilya founded a company and attracted billions of $$ promising AGI soon. Now what?
- illwrks 10mo agoIn the “teenagers learn to drive in 10 hours” part… that’s active learning, but they have spent countless hours in their life in a car, on a bus or other forms of transport, even watching the shows and movies featuring driving, playing with toys and computer games etc. There is years of passive information absorbed already before that 10 hours of active learning begins.
- martin82 10mo agoThat's a very diplomatic way of saying "we burnt all this money but we have not the faintest clue about how to proceed"