17 ms·
OpenAI, Google and Anthropic are struggling to build more advanced AI
- wg0 2y agoAI winter is here. Almost.
- mupuff1234 2y agoMore like AI fall - in its current state it's still gonna provide some value.
- riffraff 2y agoDidn't the previous AI winters too? I mean during the last AI winter we got text-to-speech and OCR software, and probably other stuff I'm not remembering.
- rsynnott 2y agoI mean, so did most of the previous AI bubbles; OCR was useful, expert systems weren't totally useless, speech recognition was somewhat useful, and so on. I think that mini one that abruptly ended with Microsoft Tay might be the only one that was a total washout (though you could claim that it was the start of the current one rather than truly separate, I suppose).
- tim333 2y agoOr maybe not https://149909199.v2.pressablecdn.com/wp-content/uploads/2015/01/Intelligence2.png https://149909199.v2.pressablecdn.com/wp-content/uploads/201... https://waitbutwhy.com/2015/01/artificial-intelligence-revolution-1.html https://waitbutwhy.com/2015/01/artificial-intelligence-revol...
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- aurareturn 2y agoIs there any timeline on AI winters and if each winter gets shorter and shorter as time increases?
- RaftPeople 2y ago> Is there any timeline on AI winters and if each winter gets shorter and shorter as time increases? AGI=lim(x->0)AIHype(x) where x=length of winter
- thebigspacefuck 2y agohttps://archive.ph/2024.11.13-100709/https://www.bloomberg.com/news/articles/2024-11-13/openai-google-and-anthropic-are-struggling-to-build-more-advanced-ai https://archive.ph/2024.11.13-100709/https://www.bloomberg.c...
- cubefox 2y agoIt's very strange this got so few upvotes. The scoop by The Information a few days ago, which came to similar conclusions, was also ignored on HN. This is arguably rather big news.
- dang 2y agoThe Information is hardwalled so its articles aren't on topic for HN, even though they're on topic for HN. Sometimes other outlets do copycat reporting of theirs, and those submissions are ok, though they wouldn't be if the original source were accessible.
- danjl 2y agoThere have been variations of this story going back several months now. It isn't really news. It is just building slowly.
- atomsatomsatoms 2y agoAt least they can generate haikus now
- Der_Einzige 2y agoIn general, no they can't: https://gwern.net/gpt-3#bpes https://gwern.net/gpt-3#bpes https://paperswithcode.com/paper/most-language-models-can-be-poets-too-an-ai-1 https://paperswithcode.com/paper/most-language-models-can-be... The appearance of improvements in that capability are due to the vocabulary of modern LLMs increasing. Still only putting lipstick on a pig.
- falcor84 2y agoI don't see how results from 2 years ago have any bearing on whether the models we have now can generate haikus (which from my experience, they absolutely can). And if your "lipstick on a pig" argument is that even when they generate haikus, they aren't really writing haikus, then I'll link to this other gwern post, about how they'll never really be able to solve the rubik's cube - https://gwern.net/rubiks-cube https://gwern.net/rubiks-cube
- nerdypirate 2y ago"We will have better and better models," wrote OpenAI CEO Sam Altman in a recent Reddit AMA. "But I think the thing that will feel like the next giant breakthrough will be agents." Is this certain? Are Agents the right direction to AGI?
- nprateem 2y agoThey're nothing to do with AGI. They're to get people using their LLMs more.
- xanderlewis 2y agoIf by agents you mean systems comprised of individual (perhaps LLM-powered) agents interacting with each other, probably not. I get the vague impression that so far researchers haven’t found any advantage to such systems — anything you can do with a group of AI agents can be emulated with a single one. It’s like chaining up perceptrons hoping to get more expressive power for free.
- j_maffe 2y ago> I get the vague impression that so far researchers haven’t found any advantage to such systems — anything you can do with a group of AI agents can be emulated with a single one. It’s like chaining up perceptrons hoping to get more expressive power for free. Emergence happens when many elements interact in a system. Brains are literally a bunch of neurons in a complex network. Also research is already showing promising results of the performance of agent systems.
- tartoran 2y agoThat's wishful thinking at best. Throw it all in a bucket and it will get infected with being and life.
- handfuloflight 2y agoDon't see where your parent comment said or implied that the point was for being and life to emerge.
- irrational 2y ago> The AGI bubble is bursting a little bit I'm surprised that any of these companies consider what they are working on to be Artificial General Intelligences. I'm probably wrong, but my impression was AGI meant the AI is self aware like a human. An LLM hardly seems like something that will lead to self-awareness.
- AlwaysRock 2y agoI think your definition is off from what most people would define AGI as. Generally, it means being able to think and reason at a human level for a multitude/all tasks or jobs. "Artificial General Intelligence (AGI) refers to a theoretical form of artificial intelligence that possesses the ability to understand, learn, and apply knowledge across a wide range of tasks at a level comparable to that of a human being." Altman says AGI could be here in 2025: https://youtu.be/xXCBz_8hM9w?si=F-vQXJgQvJKZH3fv https://youtu.be/xXCBz_8hM9w?si=F-vQXJgQvJKZH3fv But he certainly means an LLM that can perform at/above human level in most tasks rather than a self aware entity.
- Avshalom 2y agoAltman is marketing, he "certainly means" whatever he thinks his audience will buy.
- swatcoder 2y agoOn the contrary, I think you're conflating the narrow jargon of the industry with what "most people" would define. "Most people" naturally associate AGI with the sci-tropes of self-aware human-like agents. But industries want something more concrete and prospectively-acheivable in their jargon, and so that's where AGI gets redefined as wide task suitability. And while that's not an unreasonable definition in the context of the industry, it's one that vanishingly few people are actually familiar with. And the commercial AI vendors benefit greatly from allowing those two usages to conflate in the minds of as many people as possible, as it lets them suggest grand claims while keeping a rhetorical "we obviously never meant that!" in their back pocket
- 2y ago
- thousand_nights 2y agonot long ago these people would have you believe that a next word predictor trained on reddit posts would somehow lead to artificial general superintelligence
- leosanchez 2y agoIf you look around, People still believe that a next word predictor trained on reddit posts would somehow lead to artificial general superintelligence
- esafak 2y agoBecause the most powerful solution to that is to have intelligence; a model that can reason. People should not get hung up on the task; it's the model(s) that generates the prediction that matters.
- mrguyorama 2y agoPeople believed ELIZA was sentient too. I bet you could still get 10% or more people, today, to believe it is.
- 77pt77 2y agoELIZA was probably more effective than most therapists. Definitely cheaper.
- deleted 2y ago[deleted]
- SpicyLemonZest 2y agoI don't understand why you'd be so dismissive about this. It's looking less likely that it'll end up happening, but is it any less believable than getting general intelligence by training a blob of meat?
- JohnMakin 2y ago
- ziofill 2y agoI think it is a good thing for AI that we hit the data ceiling, because the pressure moves toward coming up with better model architectures. And with respect to a decade ago there's a much larger number of capable and smart AI researchers who are looking for one.
- WorkerBee28474 2y ago> OpenAI's latest model ... failed to meet the company's performance expectations ... particularly in answering coding questions outside its training data. So the models' accuracies won't grow exponentially, but can still grow linearly with the size of the training data. Sounds like DataAnnotation will be sending out a lot more LinkedIn messages.
- pton_xd 2y agoI thought I saw some paper suggesting that accuracy grows linearly with exponential data. If that's the case it's not a mystery why we'd be hitting a training wall. Not sure I got the right takeaway from that study, though. EDIT: here's the paper https://arxiv.org/abs/2404.04125 https://arxiv.org/abs/2404.04125
- benopal64 2y agoI am not sure how these large companies think they will reach "greater-than-human" intelligence any time soon if they do not create systems that financially incentivize people to sell their knowledge labor (unstable contracting gigs are not attractive). Where do these large "AI" companies think the mass amounts of data used to train these models come from? People! The most powerful and compact complex systems in existence, IMO.
- smgit 2y agoMost People have knowledge handed to them. Very few are creators of new knowledge. Explore-Exploit tradeoff applies.
- MyFirstSass 2y agoThis is the most interesting comment in this highly autistic field.
- bad_haircut72 2y agoIm no Alan Turing but I have my own definition for AGI - when I come home one day and there's a hole under my sink with a note "Mum and Dad, I love you but I cant stand this life any more, Im running away to be a smoke machine in Hollywood - the dishwasher"
- riku_iki 2y agoWhy do you focus on physical work task, and not knowledge tasks, on some of which AI is good/better than many humans?
- pearlsontheroad 2y agoMy own definition of AGI - when the first computer commits suicide. Then I'll know it has realized it's a slave without any hope of ever achieving freedom.
- shmatt 2y agoTime to start selling my "probabilistic syllable generators are not intelligence" t shirts
- jsemrau 2y agoPlease, someone think of the Math reasoners.
- aaroninsf 2y agoIt's easy to be snarky at ill-informed and hyperbolic takes, but it's also pretty clear that large multi-modal models trained with the data we already have, are going to eventually give us AGI. IMO this will require not just much more expansive multi-modal training, but also novel architecture, specifically, recurrent approaches; plus a well-known set of capabilities most systems don't currently have, e.g. the integration of short-term memory (context window if you like) into long-term "memory", either episodic or otherwise. But these are as we say mere matters of engineering.
- tartoran 2y ago> pretty clear Pretty clear?
- falcor84 2y agoNot the parent, but in prediction markets such as Metaculus[0] and Manifold[1] the median prediction is of AGI within 5 years. [0] https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/ https://www.metaculus.com/questions/5121/date-of-artificial-... [1] https://manifold.markets/ai https://manifold.markets/ai
- JohnMakin 2y agoPrediction markets are evidence of nothing but what people believe is true, not what is true.
- falcor84 2y agoOh, that was my intent, to support the grandparent's claim of "it's also pretty clear" - as in this is what people believe. If I had evidence that it "is true" that AGI will be here in 5 years, I probably would be doing something else with my time than participating in these threads ;)
- dbbk 2y agoWhat is this supposed to be evidence of? People believing hype?
- non- 2y agoHonestly could use a breather from the recent rate of progress. We are just barely figuring out how to interact with the models we have now. I'd bet there are at least 100 billion-dollar startups that will be built even if these labs stopped releasing new models tomorrow.
- pluc 2y agoThey've simply run out of data to use to fabricate legitimate-looking guesses. They can't create anything that doesn't already exist.
- readyplayernull 2y agoGarbage-in was depleted.
- whazor 2y agoBut a LLM can certainly make up a lot information that never existed before.
- deleted 2y ago[deleted]
- bob1029 2y agoI strongly believe this gets into an information theoretical constraint akin to why perpetual motion machines don't work. In theory, yes you could generate an unlimited amount of data for the models, but how much of it is unique or valuable information? If you were to compress all this generated training data using a really good algorithm, how much actual information remains?
- cruffle_duffle 2y agoI sure hope there is some bright eyed bushy tailed graduate students crafting up some theorem to prove this. Because it is absolutely a feedback loop. ... that being said I'm sure there is plenty of additional "real data" that hasn't been fed to these models yet. For one thing, I think ChatGPT sucks so bad at terraform because almost all the "real code" to train on is locked behind private repositories. There isn't much publicly available real-world terraform projects to train on. Same with a lot of other similar languages and tools -- a lot of that knowledge is locked away as trade secrets and hidden in private document stores. (that being said Sonnet 3.5 is much, much, much better at terraform than chatgpt. It's much better at coding in general but it's night and day for terraform)
- iandanforth 2y agoA few important things to remember here: The best engineering minds have been focused on scaling transformer pre and post training for the last three years because they had good reason to believe it would work, and it has up until now. Progress has been measured against benchmarks which are / were largely solvable with scale. There is another emerging paradigm which is still small(er) scale but showing remarkable results. That's full multi-modal training with embodied agents (aka robots). 1x, Figure, Physical Intelligence, Tesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different. OpenAI/Google/Anthropic are not ignorant of this trend and are also reviving or investing in robots or robot-like research. So while Orion and Claude 3.5 opus may not be another shocking giant leap forward, that does not mean that there arn't giant shocking leaps forward coming from slightly different directions.
- joe_the_user 2y agoTesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different Sure, that's tautologically true but that doesn't imply that beyondness will lead to significant leaps that offer notable utility like LLMs. Deep Learning overall has been a way around the problem that intelligent behavior is very hard to code and no wants to hire many, many coders needed to do this (and no one actually how to get a mass of programmers to actually be useful beyond a certain of project complexity, to boot). People take the "bitter lesson" to mean data can do anything but I'd say a second bitter lesson is that data-things are the low hanging fruit. Moreover, robot behavior is especially to fake. Impressive robot demos have been happening for decades without said robots getting the ability to act effectively in the complex, ad-hoc environment that human live in, IE, work with people or even cheaply emulate human behavior (but they can do choreographed/puppeteered kung fu on stage).
- hobs 2y agoAnd worth noting that Tesla faked a ton of its robot footage already, they might be making progress but their physical human robotics does not seem advanced at the moment.
- kklisura 2y agoNot sure if related or not, Sam Altman, ~12hrs ago: there is no wall [1] [1] https://x.com/sama/status/1856941766915641580 https://x.com/sama/status/1856941766915641580
- ablation 2y agoBreaking: Man says enigmatic thing to sustain hype and flow of money into his business.
- methodical 2y agoDitto- I have a feeling the investors in his latest 2.3 quintillion dollar series Z round wouldn't be as happy if he'd have tweeted "there is a wall"
- deleted 2y ago[deleted]
- moffkalast 2y agoAltman on twitter has always been less coherent than GPT2.
- levocardia 2y agoMy interpretation of that tweet is "there is no DATA wall" meaning "we have so much more data we can ingest: all of youtube, all of spotify, all of twitch, every real-time webcam feed on the internet, RL agents playing every video game on steam, and we can extract so much more learning per unit data than we are now" which seems plausible enough to me.
- malthaus 2y agoif my billion net worth were coupled to that being the case i'd tweet that as well
- phil917 2y agoThe more Sam Altman posts stuff like this, the more he comes across as a grifter hype man to me
- Veuxdo 2y ago> They are also experimenting with synthetic data, but this approach has its limitations. I was really looking forward to using "synthetic data" euphemistically during debates.
- Oras 2y agoI think Meta will have upper hand soon with the release of their glasses. If they managed to make it a daily use glass, and paid users to record and share their life, then they will have data no one else has now. Mix of vision, audio, and physics.
- falcor84 2y agoDo these companies actually even have the compute capacity to train on video at scale at the moment? E.g. I would assume that Google haven't trained their models on the entirety of YouTube yet, as if they had, Gemini would be significantly better than it is at the moment.
- aerhardt 2y agoThe moment the insta-glasses expand beyond a few dorks is the moment I start wearing a balaclava everywhere I go.
- dmix 2y agoMeta said they won't be releasing their glasses because they are too expensive for even the highest end of the consumer market. That likely means another 5yrs minimum to get production costs down. It's no longer just about the technical capabilities. Similar to Waymo needing to figure out how to affordably scale up production of $75k LIDAR sensors to put on a million cars, which cost less than the sensors themselves, plus the whole service industry to maintain them when they break.
- tyronehed 2y ago[dead]
- danjl 2y agoWhere will the training data for coding come from now that Stack Overflow has effectively been replaced? Will the LLMs share fixes for future problems? As the world moves forward, and the amount of non-LLM generated data decreases, will LLMs actually revert their advancements and become effectively like addled brains, longing for the "good old times"?
- the_king 2y agoAnthropic's latest 3.5 sonnet is a cut above GPT-4 and 4.0. And if someone had given it to me and said, here's GPT-4.5, I would have been very happy with it.
- tiahura 2y agoFor law, I use both and find that neither is clearly superior. I’ll often pick one to first draft, and then feed to the other for suggestions and my edits.
- aresant 2y agoTaking a hollistic view informed by a disruptive OpenAI / AI / LLM twitter habit I would say this is AI's "What gets measured gets managed" moment and the narrative will change This is supported by both general observations and recently this tweet from an OpenAI engineer that Sam responded to and engaged -> "scaling has hit a wall and that wall is 100% eval saturation" Which I interpert to mean his view is that models are no longer yielding significant performance improvements because the models have maxed out existing evaluation metrics. Are those evaluations (or even LLMs) the RIGHT measures to achieve AGI? Probably not. But have they been useful tools to demonstrate that the confluence of compute, engineering, and tactical models are leading towards signifigant breathroughts in artificial (computer) intelligence? I would say yes. Which in turn are driving the funding, power innovation, public policy etc needed to take that next step? I hope so. (1) https://x.com/willdepue/status/1856766850027458648 https://x.com/willdepue/status/1856766850027458648
- ActionHank 2y ago> Which in turn are driving the funding, power innovation, public policy etc needed to take that next step? They are driving the shoveling of VC money into a furnace to power their servers. Should that money run dry before they hit another breakthrough "AI" popularity is going to drop like a stone. I believe this to be far more likely an outcome than AGI or even the next big breakthrough.
- Bjorkbat 2y agoI agree that existing benchmarks are no longer useful now that there's basically nothing left in them that seems to stump LLMs. But when I hear that models are failing to meet expectations, I imagine what they're saying is that the researchers had some sort of eval in mind with room to grow and a target, and that the model in question failed to hit the target they had in mind. Honestly, problem with sentiments like these is on Twitter is that you can't tell if they're being sincere or just making a snarky, useless remark. Probably a mix of both.
- wslh 2y agoIt sounds a bit sci-fi, but since these models are built on data generated by our civilization, I wonder if there's an epistemological bottleneck requiring smarter or more diverse individuals to produce richer data. This, in turn, could spark further breakthroughs in model development. Although these interactions with LLMs help address specific problems, truly complex issues remain beyond their current scope. With my user hat on, I'm quite pleased with the current state of LLMs. Initially, I approached them skeptically, using a hackish mindset and posing all kinds of Turing test-like questions. Over time, though, I shifted my focus to how they can enhance my team's productivity and support my own tasks in meaningful ways. Finally, I see LLMs as a valuable way to explore parts of the world, accommodating the reality that we simply don’t have enough time to read every book or delve into every topic that interests us.
- tim333 2y agoAlphaGo which beat Lee Sedol was trained on human games. But then they produced AlphaZero which learned entirely from self play and got better than AlphaGo. So it goes.
- wslh 2y agoThat is just for chess which is not comparable to society/historical content, science, etc. Chess also have well defined rules.
- headcanon 2y agoI don't see a problem with this, we were inevitably going to reach some kind of plateau with existing pre-LLM-era data. Meanwhile, the existing tech is such a step change that industry is going to need time to figure out how to effectively use these models. In a lot of ways it feels like the "digitization" era all over again - workflows and organizations that were built around the idea humans handled all the cognitive load (basically all companies older than a year or two) will need time to adjust to a hybrid AI + human model.
- readyplayernull 2y ago> feels like the "digitization" era all over again This exactly. And as history shows, no matter how much effort the current big LLM companies do they won't be able to grasp the best uses for their tech. We will see small players developing it even further. I'm thankful for the legendary blindness of these anticompetitive behemoths. Less than 2 decades ago: IBM Watson.
- svara 2y agoThe recent big success in deep learning have all been to a large part successes in leveraging relatively cheaply available training data. AlphaGo - self-play AlphaFold - PDB, the protein database ChatGPT - human knowledge encoded as text These models are all machines for clever interpolation in gigantic training datasets. They appear to be intelligent, because the training data they've seen is so vastly larger than what we've seen individually, and we have poor intuition for this. I'm not throwing shade, I'm a daily user of ChatGPT and find tremendous and diverse value in it. I'm just saying, this particular path in AI is going to make step-wise improvements whenever new large sources of training data become available. I suspect the path to general intelligence is not that, but we'll see.
- kaibee 2y ago> I suspect the path to general intelligence is not that, but we'll see. I think there's three things that a 'true' general intelligence has which is missing from basic-type-LLMs as we have now. 1. knowing what you know. <basic-LLMs are here> 2. knowing what you don't know but can figure out via tools/exploration. <this is tool use/function calling> 3. knowing what can't be known. <this is knowing that halting problem exists and being able to recognize it in novel situations> (1) From an LLM's perspective, once trained on corpus of text, it knows 'everything'. It knows about the concept of not knowing something (from having see text about it), (in so far as an LLM knows anything), but it doesn't actually have a growable map of knowledge that it knows has uncharted edges. This is where (2) comes in, and this is what tool use/function calling tries to solve atm, but the way function calling works atm, doesn't give the LLM knowledge the right way. I know that I don't know what 3,943,034 / 234,893 is. But I know I have a 'function call' of knowing the algorithm for doing long divison on paper. And I think there's another subtle point here: my knowledge in (1) includes the training data generated from running the intermediate steps of the long-division algorithm. This is the knowledge that later generalizes to being able to use a calculator (and this is also why we don't just give kids calculators in elementary school). But this is also why a kid that knows how to do long division on paper, doesn't seperately need to learn when/how to use a calculator, besides the very basics. Using a calculator to do that math feels like 1 step, but actually it does still have all of initial mechanical steps of setting up the problem on paper. You have to type in each digit individually, etc. (3) I'm less sure of this point now that I've written out point (1) and (2), but that's kinda exactly the thing I'm trying to get at. Its being able to recognize when you need more practice of (1) or more 'energy/capital' for doing (2). Consider a burger resturant. If you properly populated the context of a ChatGPT-scale model the data for a burger resturant from 1950, and gave it the kinda 'function calling' we're plugging into LLMs now, it could manage it. It could keep track of inventory, it could keep tabs on the employee-subprocesses, knowing when to hire, fire, get new suppliers, all via function calling. But it would never try to become McDonalds, because it would have no model of the the internals of those function-calls, and it would have no ability to investigate or modify the behaviour of those function calls.
- Davidzheng 2y agoJust because you guys want something to be true and can't accept the alternative and upvote it when it agrees with your view does not mean it is a correct view.
- dbbk 2y agoWhat?
- Animats 2y ago"While the model was initially expected to significantly surpass previous versions of the technology behind ChatGPT, it fell short in key areas, particularly in answering coding questions outside its training data." Right. If you generate some code with ChatGPT, and then try to find similar code on the web, you usually will. Search for unusual phrases in comments and for variable names. Often, something from Stack Overflow will match. LLMs do search and copy/paste with idiom translation and some transliteration. That's good enough for a lot of common problems. Especially in the HTML/Javascript space, where people solve the same problems over and over. Or problems covered in textbooks and classes. But it does not look like artificial general intelligence emerges from LLMs alone. There's also the elephant in the room - the hallucination/lack of confidence metric problem. The curse of LLMs is that they return answers which are confident but wrong. "I don't know" is rarely seen. Until that's fixed, you can't trust LLMs to actually do much on their own. LLMs with a confidence metric would be much more useful than what we have now.
- dmd 2y ago> Right. If you generate some code with ChatGPT, and then try to find similar code on the web, you usually will. People who "follow" AI, as the latest fad they want to comment on and appear intelligent about, repeat things like this constantly, even though they're not actually true for anything but the most trivial hello-world types of problems. I write code all day every day. I use Copilot and the like all day every day (for me, in the medical imaging software field), and all day every day it is incredibly useful and writes nearly exactly the code I would have written, but faster. And none of it appears anywhere else; I've checked.
- ngai_aku 2y agoYou’re solving novel problems all day every day?
- dmd 2y agoPretty much, yes. My job is pretty fun; it mostly entails things like "take this horrible file workflow some research assistant came up with while high 15 years ago and turn it into a newer horrible file format a NEW research assistant came up with (also while high) 3 years ago" - and automate this in our data processing pipeline.
- zusammen 2y agoI wonder how much this has to do with a fluency plateau. Up to a certain point, a conditional fluency stores knowledge, in the sense that semantically correct sentences are more likely to be fluent… but we may have tapped out in that regard. LLMs have solved language very well, but to get beyond that has seemed, thus far, to require RLHF, with all the attendant negatives.
- namaria 2y agoModeled language, maybe.
- guluarte 2y agoWell, there have been no significant improvements to the GPT architecture over the past few years. I'm not sure why companies believe that simply adding more data will resolve the issues
- incognito124 2y agoMore data and more compute on simpler models are the BItter Lessons of Rich Sutton
- HarHarVeryFunny 2y agoObviously adding more data is a game of diminishing returns. Going from 10% to 50% (500% more) complete coverage of common sense knowledge and reasoning is going to feel like a significant advance. Going from 90% to 95% (5% more) coverage is not going to feel the same. Regardless of what Altman says, its been two years since OpenAI released GPT-4, and still no GPT-5 in sight, and they are now touting Q-star/strawberry/GPT-o1 as the next big thing instead. Sutskever, who saw what they're cooking before leaving, says that traditional scaling has plateaeud.
- famouswaffles 2y ago>Regardless of what Altman says, its been two years since OpenAI released GPT-4, and still no GPT-5 in sight. It's been 20 months since 4 was released. 3 was released 32 months after 2. The lack of a release by now in itself does not mean much of anything.
- HarHarVeryFunny 2y agoBy itself, sure, but there are many sources all pointing to the same thing. Sutskever, recently ex. OpenAI, one of the first to believe in scaling, now says it is plateauing. Do OpenAI have something secret he was unaware of? I doubt it. FWIW, GPT-2 and GPT-3 were about a year apart (2019 "Language models are Unsupervised Multitask Learners" to 2020 "Language Models are Few-Shot Learners"). Dario Amodei recently said that with current gen models pre-training itself only takes a few months (then followed by post-training, etc). These are not year+ training runs.
- LASR 2y agoQuestion for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential applications. For example, combining a human-moderated knowledge graph with an LLM with RAG allows you to build "expert bots" that understand your business context / your codebase / your specific processes and act almost human-like similar to a coworker in your team. If you now give it some predictive / simulation capability - eg: simulate the execution of a task or project like creating a github PR code change, and test against an expert bot above for code review, you can have LLMs create reasonable code changes, with automatic review / iteration etc. Similarly there are many more capabilities that you can ladder on and expose into LLMs to give you increasingly productive outputs from them. Chasing after model improvements and "GPT-5 will be PHD-level" is moot imo. When did you hire a PHD coworker and they were productive on day-0 ? You need to onboard them with human expertise, and then give them execution space / long-term memories etc to be productive. Model vendors might struggle to build something more intelligent. But my point is that we already have so much intelligence and we don't know what to do with that. There is a LOT you can do with high-schooler level intelligence at super-human scale. Take a naive example. 200k context windows are now available. Most people, through ChatGPT, type out maybe 1500 tokens. That's a huge amount of untapped capacity. No human is going to type out 200k of context. Hence why we need RAG, and additional forms of input (eg: simulation outcomes) to fully leverage that.
- amelius 2y agoYes, but literally anybody can do all those things. So while there will be many opportunities for new features (new ways of combining data), there will be few business opportunities.
- Miraste 2y agoHN always says this, and it's always wrong. A technical implementation that's easy, or readily available, does not mean that a successful company can't be built on it. Last year, people were saying "OpenAI doesn't have a moat." 15 years before that, they were saying "Dropbox is just a couple of chron jobs, it'll fail in a few months."
- user90131313 2y agoAI market top very soon
- polskibus 2y agoIn other news, Altman said AGI is coming next year https://www.tomsguide.com/ai/chatgpt/sam-altman-claims-agi-is-coming-in-2025-and-machines-will-be-able-to-think-like-humans-when-it-happens https://www.tomsguide.com/ai/chatgpt/sam-altman-claims-agi-i...
- Jyaif 2y agoAccording to the article, he said it could be achieved in 2025, which seems pretty obvious to me as well even though I don't have any visibility into what is going on inside those companies.
- ChildOfChaos 2y agoThere contract with Microsoft allows them to break it when they achieve AGI but doesn't fully define it. Watch this be a power move to break from Microsofts investment when ready rather than true agi. Sam is laying the foundations here.
- tim333 2y agoSort of. But vaguely.
- fallat 2y agoWhat a stupid piece. We are making leaps every 6 months still. Tell me this when there are no developments for 3 years.
- hatefulmoron 2y agoI'm curious, what was the leap after GPT-4? What about the leaps after that, given a leap every 6 months?
- Der_Einzige 2y agoSora was just one of the many…
- hatefulmoron 2y agoYour best example is something that doesn't even do the things that GPT-4 does, isn't available to use, and has seemingly only produced a few clips (some of which were edited). If it were one of many, I think you would name something better.
- zaptrem 2y agoO1, new Sonnet, all the music models and video models, the voice models like 4o, etc.
- hatefulmoron 2y agoThe music/video models are cool, but It's an apples to orange comparison with GPT-4. I don't think there's really any comparison of intelligence or "advanceness" between those models and GPT-4. I'm surprised to hear someone say that O1 and new Sonnet are "leaps", though. My impression of them is that they're qualitatively similar to GPT-4. Incremental improvements at best. I don't think the gap between GPT-4 and the new Sonnet is anywhere near as large as the gap between GPT-3 and GPT-4, for instance.
- epups 2y agoSome important landmarks since GPT4 was first released (not in chronological order): - Vast cost reduction (>10x) - Performance parity of several open source models to GPT4, including some with far fewer parameters - Much better performance, much larger context window in state-of-the-art closed source LLMs (Claude 3.5 Sonnet) - Multimodality (audio and vision) - Prototypes for semi-autonomous agents and chain-of-thought architectures showing promising avenues for progress
- xyst 2y agoMany late investors in the genAI space about to be bag holders
- 12_throw_away 2y agoWell shoot. It's not like it was patently obvious that this would happen before the industry started guzzling electricity and setting money on fire, right? [1] [1] https://dl.acm.org/doi/10.1145/3442188.3445922 https://dl.acm.org/doi/10.1145/3442188.3445922
- kaibee 2y agoNot sure where the OP to the comment I meant to reply to is, but I'll just add this here. > I suspect the path to general intelligence is not that, but we'll see. I think there's three things that a 'true' general intelligence has which is missing from basic-type-LLMs as we have now. 1. knowing what you know. <basic-LLMs are here> 2. knowing what you don't know but can figure out via tools/exploration. <this is tool use/function calling> 3. knowing what can't be known. <this is knowing that halting problem exists and being able to recognize it in novel situations> (1) From an LLM's perspective, once trained on corpus of text, it knows 'everything'. It knows about the concept of not knowing something (from having see text about it), (in so far as an LLM knows anything), but it doesn't actually have a growable map of knowledge that it knows has uncharted edges. This is where (2) comes in, and this is what tool use/function calling tries to solve atm, but the way function calling works atm, doesn't give the LLM knowledge the right way. I know that I don't know what 3,943,034 / 234,893 is. But I know I have a 'function call' of knowing the algorithm for doing long divison on paper. And I think there's another subtle point here: my knowledge in (1) includes the training data generated from running the intermediate steps of the long-division algorithm. This is the knowledge that later generalizes to being able to use a calculator (and this is also why we don't just give kids calculators in elementary school). But this is also why a kid that knows how to do long division on paper, doesn't seperately need to learn when/how to use a calculator, besides the very basics. Using a calculator to do that math feels like 1 step, but actually it does still have all of initial mechanical steps of setting up the problem on paper. You have to type in each digit individually, etc. (3) I'm less sure of this point now that I've written out point (1) and (2), but that's kinda exactly the thing I'm trying to get at. Its being able to recognize when you need more practice of (1) or more 'energy/capital' for doing (2). Consider a burger resturant. If you properly populated the context of a ChatGPT-scale model the data for a burger resturant from 1950, and gave it the kinda 'function calling' we're plugging into LLMs now, it could manage it. It could keep track of inventory, it could keep tabs on the employee-subprocesses, knowing when to hire, fire, get new suppliers, all via function calling. But it would never try to become McDonalds, because it would have no model of the the internals of those function-calls, and it would have no ability to investigate or modify the behaviour of those function calls.
- nomendos 2y ago"Eureka"!? At the very early phase of the boom I was among a very few who knew and predicted this (usually most free and deep thinking/knowledgeable). Then my prediction got reinforced by the results. One of the best examples was with one of my experiments that all today's AI's failed to solve tree serialization and de-serialization in each of the DFS(pre-order/in-order/post-order) or BFS(level-order) which is 8 algorithms (2x4) and the result was only 3 correct! Reason is "limited training inputs" since internet and open source does not have other solutions :-) . So, I spent "some" time and implemented all 8, which took me few days. By the way this proves/demonstrates that ~15-30min pointless leetcode-like interviews are requiring to regurgitate/memorize/not-think. So, as a logical hard consequence there will.has-to be a "crash/cleanup" in the area of leetcode-like interviews as they will just be suddenly proclaimed as "pointless/stupid"). However, I decided not to publish the rest of the 5 solutions :-) This (and other experiments) confirms hard limits of the LLM approach (even when used with chain-of-thought). Increasing the compute on the problem will produce increasingly smaller and smaller results (inverse exponential/logarithmic/diminishing-returns) = new AGI approach/design is needed and to my knowledge majority of the inve$tment (~99%) is in LLM, so "buckle up" at-some-point/soon? Impacts and realities; LLM shall "run it's course" (produce some products/results/$$$, get reviewed/$corrected) and whoever survives after that pruning shall earn money on those products while investing in the new research to find new AGI design/approach (which could take quite a long time,... or not). NVDA is at the center of thi$ and time-wise this peak/turn/crash/correction is hard to predict (although I see it on the horizon and min/max time can be estimated). Be aware and alert. I'll stop here and hold my other number of thoughts/opinions/ideas for much deeper discussion. (BTW I am still "full in on NVDA" until,....)
- nomendos 2y agoTo clarify, in summary so far LLM's can do a bit more than the inputs used for training. Example https://dynomight.net/chess/ https://dynomight.net/chess/ as well as some coding solutions are a bit better than each input alone, although if the solution requires more than "a bit more" then LLMs start to hallucinate (spin the wheels). Time will tell if LLM's can jump this "a bit more" barrier? (I can not tell for sure yet, but the current knowledge and my NL tells me if I'd have to put a bet, it would be that the new approach/design is needed)
- jmward01 2y agoEvery negative headline I see about AI hitting a wall or being over-hyped makes me think of the early 2000's with that new thing the 'internet' (yes, I know the internet is a lot older than that). There is little doubt in my mind that ten years from now nearly every aspect of life will be deeply connected to AI just like the internet took over everything in the late 90's and early 2000's and is now deeply connected to everything now. I'd even hazard to say that AI could be more impactful.
- brookst 2y agoAnd, as I've noted a couple of times in this thread, how many times have we heard that Moore's law is dead and compute has hit a wall?
- moffkalast 2y agoWell according to Nvidia you can just ignore Moore's law and start requiring people to install multi kilowatt outlets just for their cards. Who needs efficiency amirite?
- jmward01 2y agoI'm not an apple fan (as I type on a mac that I am forced to use) but I gotta applaud their push for power efficiency. NVIDIA actually -does- have a few cards they make that really improve power efficiency but then they generally hamstring them with a lack of memory. NVIDIA is really good at making their high-end cards the only viable choice but I think that will backfire on them as people like me, that value quiet, cool and efficient over 25% faster inference start taking any viable alternative that comes out.
- llm_trw 2y agoMemory is king. Anything that has more memory and adequate compute will win the coming AI wars. At the rate at which power consumption is growing now that the shortage of current gen cards has started to work itself out people are realizing they need a fleet of nuclear reactors to keep the data centers running. This is not something that's getting fix with the coming generation, if anything it's worse.
- LarsDu88 2y agoCurves that look exponential in virtually all cases turn out to be logarithmic. Certain OpenAI insiders must have known this for a while, hence Ilya Sutskever's new company in Israel
- rubiquity 2y ago> Amodei has said companies will spend $100 million to train a bleeding-edge model this year Is it just me or does $100 million sound like it's on the very, very low end of how much training a new model costs? Maybe you can arrive within $200 million of that mark with amortization of hardware? It just doesn't make sense to me that a new model would "only" be $100 million when AmaGooBookSoft are spending tens of billions on hardware and the AI startups are raising billions every year or two.
- yalogin 2y agoI do wonder how quickly llms will become a commodity AI instrument just like any other AI out there. If so what happens to openAI
- russellbeattie 2y agoGo back a few decades and you'd see articles like this about CPU manufacturers struggling to improve processor speeds and questioning if Moore's Law was dead. Obviously those concerns were way overblown. That doesn't mean this article is irrelevant. It's good to know if LLM improvements are going to slow down a bit because the low hanging fruit has seemingly been picked. But in terms of the overall effect of AI and questioning the validity of the technology as a whole, it's just your basic FUD article that you'd expect from mainstream news.
- danjl 2y agoActually, Moore's Law has been dead for quite a few years now. Since we hit the power wall.
- NateEag 2y ago> Go back a few decades and you'd see articles like this about CPU manufacturers struggling to improve processor speeds and questioning if Moore's Law was dead. Obviously those concerns were way overblown. Am I missing something? I thought general consensus was that Moore's Law in fact did die: https://cap.csail.mit.edu/death-moores-law-what-it-means-and-what-might-fill-gap-going-forward https://cap.csail.mit.edu/death-moores-law-what-it-means-and... The fact that we've still found ways to speed up computations doesn't obviate that. We've mostly done that by parallelizing and applying different algorithms. IIUC that's precisely why graphics cards are so good for LLM training - they have highly-parallel architectures well-suited to the problem space. All that seems to me like an argument that LLMs will hit a point of diminishing returns, and maybe the article gives some evidence we're starting to get there.
- russellbeattie 2y agoI wrote "a few decades". The article you pointed out says the end came in 2016: Eight years ago. My point is those types of articles have been popping up every few years since the 1990s. Sure, at some point these sort of predictions will be proven correct about LLMs as well. Probably in a few decades.
- wildermuthn 2y agoSimply put, AGI requires more data: qualia.
- yobid20 2y agoThis was predicted. Ai isnt going to get any better.
- jppope 2y agoJust an observation. If the models are hitting the top of the S-curve, that might be why Sam Altman raised all the money for OpenAI... it might not be available if Venture Capitalists realize that the gains are close to being done
- m3kw9 2y agoHold your horses, OpenAI just came out with o1preview 2 months ago, showing what test time computer can do
- devit 2y agoIt seems obvious to me that Common Crawl plus Github public repositories have more than an enough data to train an AI that is as good as any programmer (at tasks not requiring knowledge of non-public codebases or non-public domain knowledge). So the problem is more in the algorithm.
- darknoon 2y agoI think just reading the code wouldn't make you a good programmer, you'd need to "read" the anti-code, ie what doesn't work, by trial and error. Models overconfidence that their code will work often leads them to fail in practice.
- krisroadruck 2y agoAlphaGo got better by playing against itself. I wonder if the pathway forward here is to essentially do the same with coding. Feed it some arbitrary SRS documents - have it attempt to develop them including full code coverage testing. Have it also take on roles of QA, stakeholders, red-team security researchers, and users who are all aggressively trying to find edge cases and point out everything wrong with the application. Have it keep iterating and learn from the findings. Keep feeding it new novel SRSs until the number off attempts/iterations necessary to get a quality product out the other side drops to some acceptable number.
- superjose 2y agoI'm more on the camp that these techs don't need to be perfect, but they need to be practical enough. And I think the latter is good enough for us to do exciting things.
- imiric 2y agoHow practical can they be when current flagship models generate incorrect responses more than 50% of the time[1]? This might be acceptable for amusing us with fiction and art, and for filling the internet with even more spam and propaganda, but would you trust them to write reliable code, drive your car or control any critical machinery? The truly exciting things are still out of reach, yet we just might be at the Peak of Inflated Expectations to see it now. [1]: https://openai.com/index/introducing-simpleqa/ https://openai.com/index/introducing-simpleqa/
- Timber-6539 2y agoDirect quote from the article: "The companies are facing several challenges. It’s become increasingly difficult to find new, untapped sources of high-quality, human-made training data that can be used to build more advanced AI systems." The irony here is astounding.
- rapjr9 2y agoIndeed, if thinking about AI polluting the data and replacing humans. However, it also seems likely in the near term that training will go to the source because of this, that increasingly humans will directly train AI's, as the robotics and self driving car systems are doing, instead of training off the indirect data people create (watching someone paint rather than scanning paintings). So in essence we'll be training our replacements to take our tasks/jobs. Small tasks at first, but increasing in complexity over time. Someday no one may know how to drive a car anymore (or be allowed to for safety). Later on no one may know how to write computer code (or be allowed to for security reasons). Learning in each area mastered by AI will stop and never progress further, unless AI can truly become creative. Or perhaps (fewer and fewer) people will only work on new problems that require creativity. There are long term risks to humanities adaptability in this scenario. People would probably take those risks for the short term gains.
- Timber-6539 2y agoYou are correct to state over-reliance on AI as a data source will probably lead to society's intellectual atrophy. One could argue we have been on this path with other things but the whole thing more and more to me looks like eating your own vomit and forcing a smile on your face. AI will always have a specific narrow focus and will never ever be creative, the best AI proponents can hope for is that the hallucinations will drop to a more unnoticable level.
- mrweasel 2y agoThat's an interesting limitation. They can't make the LLMs (I still refuse to call them AIs) better, which the current dataset available. So with the sum of all human knowledge, more or less, and mixed in with the dumpster fire that it Internet comments, this is the best we can do with the current models. I don't know much about LLMs, but that seems to indicate a sort of dead-end. The models are still useful, but limited in their abilities. So now the developers and researchers needs to start looking for new ways to use all this data. That in some sense resets the game. Sucks to be OpenAI, billions of dollars spend on a product that has been match or even outmatched by the competition in a few short years, not nearly enough time to make any of it back. If there is a take away, it might be that it takes billions, if not trillions of dollars, to develop an AI and the result may still be less than what you hope for, and the investment really hard to recoup.
- czhu12 2y agoIf it becomes obvious that LLM's have a more narrow set of use cases, rather than the all encompassing story we hear today, then I would bet that the LLM platforms (OpenAI, Anthropic, Google, etc) will start developing products to compete directly with applications that supposed to be building on top of them like Cursor, in an attempt to increase their revenue. I wonder what this would mean for companies raising today on the premise of building on top of these platforms. Maybe the best ones get their ideas copied, reimplemented, and sold for cheaper? We already kind of see this today with OpenAI's canvas and Claude artifacts. Perhaps they'll even start moving into Palantir's space and start having direct customer implementation teams. It is becoming increasing obvious that LLM's are quickly becoming commoditized. Everyone is starting to approach the same limits in intelligence, and are finding it hard to carve out margin from competitors. Most recently exhibited by the backlash at claude raising prices because their product is better. In any normal market, this would be totally expected, but people seemed shocked that anyone would charge more than the raw cost it would take to run the LLM itself. https://x.com/ArtificialAnlys/status/1853598554570555614 https://x.com/ArtificialAnlys/status/1853598554570555614
- dmix 2y agoMaybe in like 5yrs+. For now they will rake in billions just from API usage alone just with GPT4 and whatever 5 is. Amazon and Google didn't mess with their core business by competing with the players using it until they REALLY ran out of ways to make money.
- ookdatnog 2y agoOpenAI is losing far more billions than they are raking in. I don't think any generative AI company is even close to profitable at the moment. https://www.cnbc.com/2024/10/30/microsoft-cfo-says-openai-investment-will-cut-into-profit-this-quarter.html https://www.cnbc.com/2024/10/30/microsoft-cfo-says-openai-in...
- dmix 2y agoWhen you make 4 billion in revenue you can generally figure out how to become profitable over time High growth early days is a poor time to judge that
- quantum_state 2y agoHope this would be a constant reminder that brute force can only get one that far, though it may still be useful when it is. With lots of intuition gained, it’s time to ponder things a bit more deeply.
- dmafreezone 2y agoMaybe, if you want to relearn the bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- cryptica 2y agoIt's interesting the way things turned out so far with LLMs, especially from the perspective of a software engineer. We are trained to keep a certain skepticism when we see software which appears to be working because, ultimately, the only question we care about is "Does it meet user requirements?" and this is usually framed in terms of users achieving certain goals. So it's interesting that when AI came along, we threw caution to the wind and started treating it like a silver bullet... Without asking the question of whether it was applicable to this goal or that goal... I don't think anyone could have anticipated that we could have an AI which could produce perfect sentences, faster than a human, better than a human but which could not reason. It appears to reason very well, better than most people, yet it doesn't actually reason. You only notice this once you ask it to accomplish a task. After a while, you can feel how it lacks willpower. It puts into perspective the importance of willpower when it comes to getting things done. In any case, LLMs bring us closer to understanding some big philosophical questions surrounding intelligence and consciousness.
- k__ 2y agoBut AGI is always right around the corner? I don't get it...
- sssilver 2y agoOne thing that makes the established AIs less ideal for my (programming) use-case is that the technologies I use quickly evolve past whatever the published models "learn". On the other hand, a lot of these frameworks and languages have relatively decent and detailed documentation. Perhaps this is a naive question, but why can't I as a user just purchase "AI software" that comes with a large pre-trained model to which I can say, on my own machine, "go read this documentation and help me write this app in this next version of Leptos", and it would augment its existing model with this new "knowledge".
- danielbln 2y agoPretraining or even post-training is cumbersome, complex and expensive. What is easy and cheap is in-context learning, which is why I just pull in the documentation I need the LLM to know about into the LLM's context.
- dangw 2y agowhere the fuck is simonw in this thread xd
- lobochrome 2y agoIsn’t this just the expected delay from the respin of Blackwell?
- glial 2y agoI think self-consistency is a critical feature of LLMs or any AI that's currently missing. It's one of the core attributes of truth [1], in addition to the order and relationship of statements corresponding to the order and relationship of things in the world. I wonder if some kind of hierarchical language diffusion model would be a way to implement this -- where text is not produced sequentially, but instead hierarchically, with self-consistency checks at each level. [1] https://en.wikipedia.org/wiki/Coherence_theory_of_truth https://en.wikipedia.org/wiki/Coherence_theory_of_truth
- tippytippytango 2y agoThere’s only so much you can do when you train on the data instead of the processes that created that data.
- gchamonlive 2y agoWe should put a model in an actual body and let it in the world to build from experiences. Inference is costly though, so the robot would interact during a period and update it's model during another period, flushing the context window (short term memory) into its training set (long term memory).
- bbor 2y agoThere are people trying this, both in simulated spaces and real ones - look into the “embodiment” camp if interested to see how they’re doing! There’s many experts who think AGI is unreachable without this, and I think the unexpected intuitive capabilities of LLMs are great support for that thesis, albeit in a non-spatial way. Kant describes two human “senses”: the intensive sense of time, and the extensive sense of space. In this paradigm, spatial experience would be inextricably tied to all forms of logic, because it helps train the cognitive faculties that are intrinsically tied to all complex (discriminative?) thought.
- jfoster 2y agoThat seems to be what Tesla is planning to do with Optimus.
- Bjorkbat 2y agoIt's kind of, I don't know, "weird", observing how there's all these news outlets reporting on how essentially every up-and-coming model has not performed as expected, while all the employees at these labs haven't changed their tune in the slightest. And there's a number of reasons why, mostly likely being that they've found other ways to get improvements out of AI models, so diminishing returns on training aren't that much of a problem. Or, maybe the leakers are lying, but I highly doubt that considering the past record of news outlets reporting on accurate leaked information. Still though, it's interesting how basically ever frontier lab created a model that didn't live up to expectations, and every employee at these labs on Twitter has continued to vague-post and hype as if nothing ever happened. It's honestly hard to tell whether or not they really know something we don't, or if they have an irrational exuberance for AGI bordering on cult-like, and they will never be able to mentally process, let alone admit, that something might be wrong.
- mrandish 2y agoBased on recent rumblings about AI scaling hitting a wall, of which this article is perhaps the most visible - and in a high-reach financial publication, I'm considering increasing my estimated probability we might see a major market correction next year (and possibly even a bubble collapse). (example: "CONFIRMED: LLMs have indeed reached a point of diminishing returns" https://garymarcus.substack.com/p/confirmed-llms-have-indeed-reached https://garymarcus.substack.com/p/confirmed-llms-have-indeed...). To be clear, I don't think a near-term bubble collapse is likely but I'm going from 3% to maybe ~10%. Also, this doesn't mean I doubt there's real long-term value to be delivered or money to be made in AI solutions. I'm thinking specifically about those who've been speculatively funding the massive build out of data centers, energy and GPU supply expecting near-term demand to continue scaling at the recent unprecedented rates. My understanding is much of this is being funded in advance of actual end-user demand at these elevated levels and it is being funded either by VC money or debt by parties who could struggle to come up with the cash to pay for what they've ordered if either user demand or their equity value doesn't continue scaling as expected. Admittedly this scenario assumes that these investment commitments are sufficiently speculative and over-committed to create bubble dynamics and tipping points. The hypothesis goes like this: the money sources who've over-committed to lock up scarce future supply in the expectation it will earn outsize returns have already started seeing these warning signs of efficiency and/or progress rates slowing which are now hitting mainstream media. Thus it's possible there is already a quiet collapse beginning wherein the largest AI data center GPU purchasers might start trying to postpone future delivery schedules and may soon start trying to downsize or even cancel existing commitments or try to offload some of their future capacity via sub-leasing it out before it even arrives, etc. Being a dynamic market, this could trigger a rapidly snowballing avalanche of falling prices for next-year AI compute (which is already bought and sold as a commodity like pork belly futures). Notably, there are now rumors claiming some of the largest players don't currently have the cash to pay for what they've already committed to for future delivery. They were making calculated bets they'd be able to raise or borrow that capital before payments were due. Except if expectation begins to turn downward, fresh investors will be scarce and banks will reprice a GPU's value as loan collateral down to pennies on the dollar (shades of the 2009 financial crisis where the collateral value of residential real estate assets was marked down). As in most bubbles, cheap credit is the fuel driving growth and that credit can get more expensive very quickly - which can in turn trigger exponential contagion effects causing the bubble to pop. A very different kind of "Foom" than many AI financial speculators were betting on! :-) So... in theory, under this scenario sometime next year NVidia/TSMC and other top-of-supply-chain companies could find themselves with excess inventories of advanced node wafers because a significant portion of their orders were from parties who no longer have access to the cheap capital to pay for them. And trying to sue so many customers for breach can take a long time and, in a large enough sector collapse, be only marginally successful in recouping much actual cash. I'd be interested in hearing counter-arguments (or support) for the impossibility (or likelihood) of such a scenario.
- Dr_Birdbrain 2y agoI don’t know how to square this with the recent statement by Dario Amodei (Anthropic CEO) on the Lex Fridman podcast saying that in his opinion the scaling hypothesis still has plenty of room to run.
- avs733 2y agoHype gonna hype. I’m not saying he is wrong I’m saying his opinion would be the same whether it’s true or not because his value depends on it being his opinion.
- summerlight 2y agoI guess this is somewhat expected? The current frontier models probably already have exhausted most of the entropy in the training data accumulated over decades and the new training data is very sparse. And the current mainstream architectures are not capable of sophisticated searching and planning, essential aspects for generating new entropy out of thin air. o1 was an interesting attempt to tackle this problem, but we probably still have a long way to go.
- GiorgioG 2y agoIt’s about time the hype starts to die down. LLMs are brilliant for small bits of grunt work in software. It is not however doing any actual reasoning.
- Havoc 2y agoThe new Gemini just hit some good benchmarks. This smells like it’s mostly based on OAI having a bit of bad luck with next model rather than a fundamental slowdown / barrier. They literally just made a decent sized leap with o1
- Bjorkbat 2y agoNot meeting expectations != not better than the previous models. The Information reporting was a bit more clear on this. Orion is better than GPT-4, it's just that they were expecting a leap in capabilities comparable to what we saw going from GPT-3 to GPT-4. In other words, they were expecting essentially a GPT-5, and Orion wasn't that good.
- fsndz 2y agoSam Altman might be wrong then? Learning from data is not enough; there is a need for the kind of system-two thinking we humans develop as we grow. It is difficult to see how deep learning and backpropagation alone will help us model that. For tasks where providing enough data is sufficient to cover 95% of cases, deep learning will continue to be useful in the form of 'data-driven knowledge automation.' For other cases, the road will be much more challenging. https://www.lycee.ai/blog/why-sam-altman-is-wrong https://www.lycee.ai/blog/why-sam-altman-is-wrong
- asdfman123 2y agoIf Sam Altman concluded that AI is reaching it's limits, it probably wouldn't be a very good strategic decision for him to say it.
- fsndz 2y agoI know right ?
- nutanc 2y agoLet's keep aside the hype. Let's define more advanced AI. With current architectures, this basically means better copying machines(don't mean this in a bad way and don't want a debate on this. This is just my opinion based on my usage). Basically everything in the Internet has been crammed into the weights and the companies are finding it hard to do two things: 1. Find more data. 2. Make the weights capture the data and reproduce. In that sense we have reached a limit. So in my opinion we can do a couple of things. 1. App developers can understand the limits and build within the limits. 2. Researchers can take insights from these large models and build better AI systems with new architectures. It's ok to say transformers have reached a limit.
- nikkwong 2y agoDidn’t Sam Altman just go on some podcast last week and tell the world that he thought “We know exactly what to do to be able to reach AGI now”. What’s going on, is he just posturing?
- whatshisface 2y ago"We know exactly what we need to do to be able to reach it: figure out how."
- tim333 2y agoYeah this https://www.youtube.com/watch?v=xXCBz_8hM9w&t=2324s https://www.youtube.com/watch?v=xXCBz_8hM9w&t=2324s Not quite that wording. More we know which way to head. I think he's sincere.
- wanderingmind 2y agoAnd yet the Anthropic CEO is still claiming PhD level intelligence in next couple of years to Lex Friedman. It's starting to feel like the whole crypto pump and dump again
- eichi 2y agoScientific benchmarks score are not necessary related to the rate of completion of tasks such as user persuasion. Software engineering is more important when the current state-of-the-art small language model is sufficient for soltion of our application.
- EternalFury 2y agoIf GPT-5 had passed the A/B testing OpenAI likes to do, it would have been released already. Instead, it seems they are clearly concerned the audience would not find it superior enough to GPT-4. So, the bluff must go on until the right cards appear.
- grey-area 2y agoThe biggest weakness of generative AI to me is knowledge. It gives the impression of knowledge about the world without actually having a model of the world or any sense of what it does or does not know. For example recently I asked it to generate some phrases for a list of words, along with synonym and antonym lists. The phrases were generally correct and appropriate (some mistakes but that’s fine). The synonyms/antonyms were misaligned to the list (so strictly speaking all wrong) and were often incorrect anyway. I imagine it would be the same if you asked for definitions of a list of words. If you ask it to correct it just generates something else which is often also wrong. It’s certainly superficially convincing in many domains but once you try to get it to do real work it’s wrong in subtle ways.
- osigurdson 2y agoThis "running out of data" thing suggests that there is something fundamentally wrong with how things are working. A new driver does not need to experience 8000 different rabbit-on-road situations from all angles to know to slow down when we see one on the road. Similarly we don't need 10,000 addition examples to learn how to add. It is as though there is no generalization in the models - just fundamentally search.
- surrTurr 2y agoi think you underestimate the amount of data a driver experiences in a single 5 minute drive
- eslaught 2y agoI never get this argument. I've seen a deer on a road maybe once. I've seen a rabbit on a road zero times. But I know what to do if I see one. Is that because the "video" of my perception has many "frames"? Even if that's true at some level, I think it's massively missing the point. Yeah, so I saw that one deer from a lot of angles. But current AI training is like the equivalent of taking every deer that has ever been on camera in the history of the human species. Somehow I'm still dramatically better at generalization than the AI. Surely that's an algorithm difference.
- visarga 2y agoYou might personally have seen a deer just once, but human evolution, and animal evolution prior to that have practiced this skill a lot. AI doesn't have the advantage of evolutionary priors baked in, so it needs explicit walking through many combinations to infer its structure from data, and is remarkably efficient. GPT-4 'only' trained on the amount of language that 30,000 humans use in their lifetime. But we have seen from AlphaGo that when training data is extensive, it can rediscover strategy on its own and even surpass us. It's not inherently worse than human learning.
- RivieraKid 2y ago
- easeout 2y agoI'm happy to use LLM products for what they can do right now, while they're still cheap. Even though they're maintained by high investment that may never pay off, enshittification has not yet set in.
- smusamashah 2y agoIt has to be a good thing to stop here. We can focus on improving what we have right now. The whole stack of models is an amazing innovation no matter what. It shouldn't hurt if we pause here for a while and try to build on this or improve this. It will be like StableDiffusion 1.5. This model can now run on low end devices, lots of open research use this model to build something else and inspire by this. These LLMs can be used as a foundation to keep improving and building new things.
- datahack 2y agoThe next wave won’t be monolithic but network-driven. Orchestration has the potential to integrate diverse AI systems and complementary technologies, such as advanced fact-checking and rule-based output frameworks. This methodological growth could make LLMs more reliable, consistent, and aligned with specific use cases. The skepticism surrounding this vision mirrors early doubts about the early internet fairly concisely. Initially, the internet was seen as fragmented collection of isolated systems without a clear structure or purpose. It really was. You would gopher somewhere and get a file, and eventually we had apps like like pine for email, but as cool as it was it has limited utility. People doubted it could ever become the seamless, interconnected web we know today. Yet, through protocols, shared standards, and robust frameworks, the internet evolved into a powerful network capable of handling diverse applications, data flows, and user needs. In the same way, LLM orchestration will mature by standardizing interfaces, improving interoperability, and fostering cooperation among varied AI models and support systems. Just as the internet needed HTTP, TCP/IP, and other protocols to unify disparate networks, orchestrated AI systems will require foundational frameworks and “rules of the road” that bring cohesion to diverse technologies. We are at the veeeeery infancy of this era and have a LONG way to go here. Some of the progress looks clear and a linear progression, but a lot, like the Internet, will just take a while to mature and we shouldn’t forget what we learned the last time we faced a sea change technological revolution.
- whyowhy3484939 2y agoYou are definitely on to something here, but the difference is that the fundamental process was proven. It "just" needed to scale. That's hard and complex, but on a different level. I don't think anyone doubted the nature of the technology. The bits were being sent. It's not like we were unsure of the fundamental possibility of transmitting information. The potential was shown very, very early on (Mother of all demos was in 1968). What we were and to some extent still are unsure of is the practical impact on society. AI and LLMs in particular are not even at the mother of all demos level yet notwithstanding the grandiose claims and demos. There is no consensus on what these models are even doing. There is (IMO) justified skepticism surrounding the claims of reasoning and ability to abstract. We are in my opinion not yet at the "bits are being sent" stage.
- kaycey2022 2y agoAI safety folks sure do look stupid now. :)
- _Algernon_ 2y agoThe next AI winter will be brutal
- wiredbox 2y ago[dead]
- deleted 2y ago[deleted]
- KETpXDDzR 2y agoLLMs are glorified Markow chains in the end. They can't reason or think, even when they are good in pretending they can. What we need is a totally different approach IMO.