24 ms·
The Bitter Lesson Is Misunderstood
- singulargalaxy 1y agoPerhaps the author should have acknowledged the catalyst for this blog post? https://x.com/andrewgwils/status/1953814226188841417 https://x.com/andrewgwils/status/1953814226188841417
- FloorEgg 1y agoThe problem I am facing in my domain is that all of the data is human generated and riddled with human errors. I am not talking about typos in phone numbers, but rather fundamental errors in critical thinking, reasoning, semantic and pragmatic oversights, etc. all in long-form unstructured text. It's very much an LLM-domain problem, but converging on the existing data is like trying to converge on noise. The opportunity in the market is the gap between what people have been doing and what they are trying to do, and I have developed very specialized approaches to narrow this gap in my niche, and so far customers are loving it. I seriously doubt that the gap could ever be closed by throwing more data and compute at it. I imagine though that the outputs of my approach could be used to train a base model to close the gap at a lower unit cost, but I am skeptical that it would be economically worth while anytime soon.
- stego-tech 1y agoThis is my current drum I bang on when an uninformed stakeholder tries shoving LLMs blindly down everyone’s throats: it’s the data, stupid. Current data aggregates outside of industries wholly dependent on it (so anyone not in web advertising, GIS, or intelligence) are garbage, riddled with errors and in awful structures that are opaque to LLMs. For your AI strategy to have any chance of success, your data has to be pristine and fresh, otherwise you’re lighting money on fire. Throwing more compute and data at the problem won’t magically manifest AGI. To reach those lofty heights, we must first address the gaping wounds holding us back.
- FloorEgg 1y agoYes, for me both customers and colleagues continually suggested "hey let's just take all these samples of past work and dump it in the magical black box and then replicate what they have been doing". Instead I developed a UX that made it as easy as possible for people to explain what they want to be done, and a system that then goes and does that. Then we compare the system's output to their historical data and there is always variance, and when the customer inspects the variance they realize that their data was wrong and the system's output is far more accurate and precise than their process (and ~3 orders of magnitude cheaper). This is around when they ask how they can buy it. This is the difference between making what people actually want and what they say they want: it's untangling the why from the how.
- marlott 1y agoInteresting! Could you give an example with a bit more specific detail here? I take it there's some kind of work output, like a report, in a semi-structured format, and the goal is to automate creation of these. And you would provide a UX that lets them explain what they want the system to create?
- FloorEgg 1y agoYes, essentially. There are multiple long-form text inputs, one set is provided by User A, and another set by User B. User A inputs act as a prompt for User B, and then User A analyzes User B's input according to the original User A inputs, producing an output. My system takes User A and B inputs and produces the output with more accuracy and precision than User As do, but a wide margin. Instead of trying to train a model on all the history of these inputs and outputs, the solution was a combination of goal->job->task breakdown (like a fixed agentic process), and lots of context and prompt engineering. I then test against customer legacy samples, and inspect any variances by hand. At first the variances were usually system errors, which informed improvements to context and prompt engineering, and after working through about a thousand of these (test -> inspect variance -> if system mistake improve system -> repeat) iterations, and benefiting from a couple base-model upgrades, the variances are now about 99.9% user error (bad historical data or user inputs) and 0.1% system error. Overall it took about 9 months to build, and this one niche is worth ~$30m a year revenue easy, and everywhere I look there are market niches like this... it's ridiculous. (and a basic chat interface like ChatGPT doesn't work for these types of problems, no matter how smart it gets, for a variety of reasons) So to summarize: Instead of training a model on the historical inputs and outputs, the solution was to use the best base model LLMs, a pre-determined agentic flow, thoughtful system prompt and context engineering, and an iterative testing process with a human in the loop (me) to refine the overall system by carefully comparing the variances between system outputs and historical customer input/output samples.
- mediaman 1y agoThis is one reason why verifiable rewards works really well, if it's possible for a given domain. Figuring out how to extract signal and verify it for an RL loop will be very popular for a lot of niche fields.
- incompatible 1y agoWhen studying human-created data, you always need to be aware of these factors, including bias from doctrines, such as religion, older information becoming superseded, outright lies and misinformation, fiction, etc. You can't just swallow it all uncritically.
- simianwords 1y agoYou just need data to be directionally correct. It doesn’t have to be absolutely correct. We still got pretty far by scraping internet data which we all know is not fully trust worthy.
- cs702 1y agoI don't think Sutton's essay is misunderstood, but I agree with the OP's conclusion: We're reaching scaling limits with transformers. The number of parameters in our largest transformers, N, is now in the order of trillions, which is the most we can apply given the total number of tokens of training data available worldwide, D, also in the order of trillions, resulting in a compute budget C = 6N × D, which is in the order of D². OpenAI and Google were the first to show these transformer "scaling laws." We cannot add more compute to a given compute budget C without increasing data D to maintain the relationship. As the OP puts it, if we want to increase the number of GPUs by 2x, we must also increase the number of parameters and training tokens by 1.41x, but... we've already run out of training tokens. We must either (1) discover new architectures with different scaling laws, and/or (2) compute new synthetic data that can contribute to learning (akin to dreams).
- FloorEgg 1y agoWhat about or (3) models that interact with the real world? To be clear I also agree with your (1) and (2).
- jvanderbot 1y agoPlay in the real world generates a data point every few minutes. Seems a bit slow?
- FloorEgg 1y agoWhat are you basing that statement on? What exactly are you considering a "data point"? Are you assuming one model = one agent instance? I am pretty sure that there is more information (molecular structure) and functional information (I(Ex )) just in the room I am sitting in than all the unique, useful, digitized information on earth.
- pizzly 1y agoHumans experience (play in the real world) is multi modal though vision, sound, touch, pressure, muscle feedback, gravitational, etc. Its extremely rich in data. Its also not a data point its continuous stream of information. Also I would bet that humans synthesize data at the same time. Everytime we run multiple scenarios in our mind before choosing the one we execute without even thinking about it is synthesizing data. Also humans dream which is another form of data synthesizing. Allowing AI to interact with the real world is definitely a way to go.
- TheDudeMan 1y agoI interpret The Bitter Lesson as suggesting that you should be selecting methods that do not need all that data (in many domains, we don't know those methods yet).
- NooneAtAll3 1y agowhile I don't disagree with the facts, I don't understand the... tone? when Dennard scaling (single core performance) started to fail in 90s-00s, I don't think there was a sentiment "how stupid was it to believe such a scaling at all"? sure, people were compliant (and we still meme about running Crysis), but in the end the discussion resulted in "no more free lunch" - progress in one direction has hit a bottleneck, so it's time to choose some other direction to improve on (and multi-threading has now become mostly the norm) I don't really see much of a difference?
- geetee 1y agoI don't understand why we need more data for training. Assuming we've already digitized every book, magazine, research paper, newspaper, and other forms of media, why do we need this "second internet?" Legal issues aside, don't we already have the totality of human knowledge available to us for training?
- incompatible 1y agoA lot of newspapers seem to be stuck behind paywalls, even when in the public domain.
- deleted 1y ago[deleted]
- dr_dshiv 1y agoLet’s keep in mind that we don’t have most of the renaissance through the early modern period (1400-1800) because it was published in neolatin with older typefaces— and only about 10% is even digitized. We probably don’t have most of the Arabic corpus either — and barely any Sanskrit. Classical Chinese is probably also lacking — only about 1% of it is translated to English.
- jacobolus 1y agoThe volume of text in English and digitized from the past few years dwarfs the volume of Latin text from all time. Unless you are wondering about a very niche historical topic there’s more written in English than Latin about basically everything.
- dr_dshiv 1y agoWell, if you are looking for diversity of perspective— temporal diversity may be valuable. Marsilio Ficino was hired by the Medici to translate Plato and other classical Greek works into Latin. He directly taught DaVinci, Raphael, Michelangelo, Toscanelli, etc. I mean to say that his ideas and perspectives helped spark the renaissance. Insofar as we hope for an AI renaissance and not an AI apocalypse, it might benefit us to have the actual renaissance in the training data.
- back2dafucha 1y agoAbout 28 years ago a wise person said to me: "Data will kill you" Even mainframe programmers knew it.
- bwhiting2356 1y agoAudio and video data can be collected from the real world. It won't be immediate and won't be cheap.
- frankenstine 1y ago> The path forward: data alchemists (high-variance, 300% lottery ticket) or model architects (20-30% steady gains) No, the paths forward are: better design, training, feeding in more video, audio, and general data from the outside world. The web is just a small part of our experience. What about apps, webcam streams, radio from all over the world in its many forms, OTA TV, interacting with streaming content via remote, playing every video game, playing board games with humans, feeds and data from robots LLMs control, watching everyone via their phones and computers, car cameras, security footage and CCTV, live weather and atmospheric data, cable television, stereoscopic data, ViewMaster reels, realtime electrical input from various types of brains while interacting with their attached creatures, touch and smell, understanding birth, growth, disease, death, and all facets of life as an observer, observing those as a subject, expanding to other worlds, solar systems, galaxies, etc., affecting time and space, search and communication with a universal creator, and finally understanding birth and death of the universe.
- Quarrelsome 1y agoI really enjoyed reading this article as I found its content extremely insightful, but I fear I must whine for far too long about something entirely minor. As someone that didn't go to expensive maths club, the way people who did, talk about maths is disgraceful imho. Consider the equasion in this article: (C ~ 6 N⋅D) I can look up the symbol for "roughly equals", that was super cool and is a great part of curiousity. But this _implied_ multiplication between the 6 and the N combined with using a fucking diamond symbol (that I already despise given how long it took me to figure the first time I encountered it) is just gross. I figured it was likely that but then I was like: "but why not just 6ND? Maybe there's a reason why N⋅D but 6 N? Does that mean there's a difference between those operations"? Thankfully I can use gippity these days to get by, but before gippity I had to look up an entire list of maths symbols to find the diamond symbol to work out what it meant. Its why I love code because there's considerably less implicit behaviour once you slap down the formula into code and you can play with the input/output. I don't think mathsy people realise how exclusionary their communication is, but its so frustrating when I end up fumbling around in slow-mo when the maths kicks in, because "oh the /2 when discussing logarithms in comp sci is _obvious_, so we just don't put it in the equasion" just kills me. Idiot me, staring at the equasion thinking it actually makes sense without knowing the special maths knowledge of implication means that it actually doesn't solve as it reads on the page. Unless of course you went to expensive maths club where they tell you all this. What drives me nuts is that every time I spend ages finally grokking something, I realise how obvious it is and how non-trivial it is to explain it simply. Comp sci isn't much better to be honest, where we use CQRS instead of "read here, write there". Which results in thousands of newbies trying to parse the unfathomable complexity of "Command Query Responsibility Segregation" and spending as much time staring at its opaqueness as I did the opening sentence of the wikipedia article on logarithms. Idk what my point is, I just don't understand what's wrong with 6⋅N⋅D or 6*N*D. Do mathmeticians feel ugly if they write something down like that or smth?
- Chinjut 1y agoWhat diamond symbol?
- Quarrelsome 1y ago
- Mistletoe 1y ago>And herein lies the problem — we’ve basically ingested the entire Internet, and there is no second Internet. One of the best things I've read in a while about AI.
- throwaway314155 1y agoThe scaling laws for transformers _deliberately_ factor in the amount of data as well as the amount of compute needed in order to scale. The premise of this article, that data is more important than compute has been obvious to people who are paying attention. Sorry but the unnecessary sensationalism in this article was mildly annoying to me. Like the author discovered some novel new insight. A bit like that doctor who published a "no el" paper about how to find the area under a curve.
- gavmor 1y ago> The premise of this article... has been obvious to people who are paying attention. Well, forgive me but I feel that the article is a much-needed injection of context into my thinking around the Bitter Lesson. I like the imperative to preface compute requests with data roadmaps. I'm not an AI guy. Not an ML engineer. I've been studiously avoiding the low-level stuff, actually, because I didn't want to half-ass it when off-the-shelf solutions were still providing tremendous novelty and value for my customers. So, for most of my career, "compute" has been practically irrelevant! RAM and disk constraints presented more frequent obstacles than processor cycles'. I would have easily told you that data presents more of a bottleneck to value than CPU. But that's just the era of computing I came up in. The last few years have been different. Suddenly compute is at a premium, again. So it's easy to think, "if only I had more," and "line goes up!" and forget about s-curves and logarithmic scaling. Is the article unnecessarily sensationalist? I don't know, maybe you've been overestimating how much the rest of us are "paying attention."[0] 0. https://xkcd.com/2501/ https://xkcd.com/2501/
- kushalc 1y agoHey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards and IMO that's exactly it: if you can find a way to produce verifiable rewards about a target world, you can essentially produce unlimited amounts of data and (likely) scale past the current bottleneck. Then the question becomes, working backwards from the set of interesting 4-8h METR tasks, what worlds can we make verifiable rewards for and how do we scalably make them? [1] Which is to say, it's not about more data in general, it's about the specific kind of data (or architecture) we need to break a specific bottleneck. For instance, real-world data is indeed verifiable and will be amazing for robotics, etc. but that frontier is further behind: there are some cool labs building foundational robotics models, but they're maybe ~5 years behind LMs today. [1] There's another path with better design, e.g. CLIP that improves both architecture and data, but let's leave that aside for now.
- Quarrelsome 1y ago> if you can find a way to produce verifiable rewards about a target world I feel like there's an interesting symmetry here between the pre and post LLM world, where I've always found that organisations over-optimise for things they can measure (e.g. balance sheets) and under-optimise for things they can't (e.g. developer productivity), which explains why its so hard to keep a software product up to date in an average org, as the natural pressure is to run it into the ground until a competitor suddenly displaces it. So in a post LLM world, we have this gaping hole around things we either lack the data for, or as you say: lack the ability to produce verifiable rewards for. I wonder if similar patterns might play out as a consequence and what unmodelled, unrecorded, real-world things will be entirely ignored (perhaps to great detriment) because we simply lack a decent measure/verifiable-reward for it.
- FloorEgg 1y ago10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), and yet... where are all the awesome robots? Why can't I buy a robot that loads my dishwasher for me? Last year this really started to bug me, and after digging into it with some friends I think we collectively realized something that may be a hint at the answer. As far as we know, it took roughly 100M-1B years to evolve human level "embodiment" (evolve from single celled organisms to human), but it only took around ~100k-1M for humanity to evolve language, knowledge transfer and abstract reasoning. So it makes me wonder, is embodiment (advanced robotics) 1000x harder than LLMs from an information processing perspective?
- benlivengood 1y agoI don't think anyone has yet trained on all videos on the Internet. Plenty of petabytes left there to pretrain on, and likely just as useful once the text/audio/image pretraining is done.
- simianwords 1y agoIt might have been trained on a select high quality of videos, say more than 10k views and only trained on its transcripts.
- scrivna 1y agoSeems like reading a transcript of the commentary from a football game, it’s obviously missing a lot of information.
- g42gregory 1y agoThe D here is not exactly defined (or maybe I just missed that). Does synthetic data count? What about making several more passes through already available data?
- paulsutter 1y agoPhysical simulation is the most important underutilized data source. It’s very large, but also finite. And once you’ve learned the complexity of reality you won’t need more data you’ll be done
- madrox 1y agoIn any field where there is a creative element, progress comes in fits and starts that are difficult to predict in advance. No one can accurately predict when we'll get the cure for cancer, for example, in spite of people working on it. But that isn't how investors operate. They want to know what they will get in exchange for giving a company a billion dollars. If you're running an AI business, you need to set expectations. How do you do that? Go do the thing you know you can do on a schedule, like standing up a new GPU data center. I don't think the bitter lesson is misunderstood in quite the way the author describes. I think most are well aware we're approaching the data wall within a couple years. However, if you're not in academia you're not trying to solve that problem; you're trying to get your bag before it happens. That may sound a little flip, but this is yet another incarnation of the hungry beast: https://stvp.stanford.edu/clips/the-hungry-beast-and-the-ugly-baby/ https://stvp.stanford.edu/clips/the-hungry-beast-and-the-ugl...
- simianwords 1y agoWhy do you assume investors don’t know about this? They know some investments follow the power law - very few of them work out but they bring most value. The very existence of openAI and Anthropic are proof of it happening. Imagine you were an investor and you know what you know now (creativity can’t be predicted). How would you then invest in companies? Your answer might converge on existing VC strategies.
- madrox 1y agoI don't assume that at all. Investors absolutely know, but investment is predicated on returns. You can't do that if you can't give a timeline for when value will be generated unless your investment is so small it's practically a donation. Obviously you can invest in moonshots, but you don't want to bet your whole portfolio. Why do you think OpenAI had the governing structure it did before it made its breakthroughs but suddenly both them and Anthropic can do insane raises?
- theahura 1y agoHas HRM really dramatically changed the landscape? My read of the paper thus far is that it is an impressive result, but there have been a few of those in the past that have fizzled out, so I'm still in wait-and-see mode
- casey2 1y agoIt's a boot-strapping problem. LLMs have shown that we can reproduce data that's already in the form we want, and use that data to solve novel problems. There is no shortage of data, it's just data that's in a form you want is hard to come by. You want to create a model that generates steps for a robot with a particular shape? First you have to create a robot with that shape that can walk, then create a million of them and record them walking all over the place. Now you have something that's probably going to be too slow to run. Not fesible in the real world, the closest we have today is something like driverless car, (which is already a solved problem they are called trains) This is why I think China will ultimately win the AI race, they will be able to put tens of millions of people to a specific task until there is enough data generated to replace humans on that task in 99.99% of cases, and they have the manufacturing capability to make the millions of IO devices needed for this. Yes, humanoid robots are a good idea, but only if you can train them with walking data from real people, I think it will probably translate well enough to most humanoid robots, but ideally you are designing the physical robot from the ground up to model human movement as close as possible. You have to accept that if we go the LM route for AI that the optimal hardware behaves like human wetware. The neuromorphic computing people get it, robotics people should too.
- datadrivenangel 1y agoIf a problem is worth throwing 10 million people at it, it's worth putting the problem into a deterministically solvable form. Legal AI would be easy if we made our legal code more robust
- credit_guy 1y ago> There is no second internet I don't know about that. LLMs have been trained mostly on text. If you add photos, audio and videos, and later even 3D games, or 3D videos, you get massively more data than the old plain text. Maybe by many orders of magnitude. And this is certainly that can improve cognition in general. Getting to AGI without audio and video, and 3D perception seems like a non-starter. And even if we think AGI is not the goal, further improvements from these new training datasets are certainly conceivable.
- 1970-01-01 1y agoYes. It's a complete oversight and wrong. This paper missed: darknets, the deep web, Usenet, BBS, Internet2, and all other paywalled archives .
- Symmetry 1y agoAlso, even if we lacked the data to proceed with Chinchilla-optimal scaling that wouldn't be the same as being unable to proceed with scaling, it would just require larger models and more flops than we would prefer.
- qcnguy 1y agoThat's been done already for years. OpenAI were training on bulk AI transcribed YouTube vids already in the GPT-4 era. Modern models are all multi-modal and cotrained on audio and image tokens together with text. The AI companies are not only out of such data but their access to it is shrinking as the people who control the hosting sites wall them off (like YouTube).
- sfpotter 1y agoOh man, I love crazy stuff like this on HN. For a community which espouses rationality and careful thought, somehow an article with "C ~ D^2" has floated to the top. No notes.
- nightsd01 1y agoI am not an expert in AI by any means but I think I know enough about it to comment on one thing: there was an interesting paper not too long ago that showed if you train a randomly-initialized model from scratch on questions, like a bank of physics questions & answers, models will end up with much higher quality if you teach it the simple physics questions first, and then move up to more complex physics questions. This shows that in some ways, these large language models really do learn like we do. I think the next steps will be more along this vain of thinking. Treating all training data the same is a mistake. Some data is significantly more valuable to developing an intelligent model than most other training data, even when you pass quality filters. I think we need to revisit how we 'train' these models in the first place, and come up with a more intelligent/interactive system of doing so
- flux3125 1y agoCurriculum learning: https://en.wikipedia.org/wiki/Curriculum_learning https://en.wikipedia.org/wiki/Curriculum_learning
- nikki93 1y agoA relevant paper: https://arxiv.org/abs/2306.11644 https://arxiv.org/abs/2306.11644 -- the Phi models (and many others too) are based on this idea.
- FloorEgg 1y agoWow. I really like this take. I've seen how time and time again nature follows the Pareto principle. It makes sense that training data would follow this principle as well. Further that the order of training matters is novel to me and seems so obvious in hindsight. Maybe both of these points are common knowledge/practice among current leading LLM builders. I don't build LLMs, I build on and with them, so I don't know.
- a2128 1y agoFrom my personal experience training models this is only true when the parameter count is a limiting factor. When the model is past a certain size, it doesn't really lead to much improvement to use curriculum learning. I believe most research also applies it only to small models (e.g. Phi)
- aerospades 1y agoI disagree with the author's thesis about data scarcity. There's an infinite amount of data available in the real world. The real world is how all generally intelligent humans have been trained. Currently, LLMs have just been trained on the derived shadows (as in Plato's allegory of the cave). The grounding to base reality seems like an important missing piece. The other data type missing is the feedback: more than passively training/consuming text (and images/video), being able to push on the chair and have it push back. Once the AI can more directly and recursively train on the real world, my guess is we'll see Sutton's bitter lesson proven out once again.
- lawrencechen 1y agoIn the bitter lesson essay [0], the word "data" is not mentioned a single time. The author fundamentally misunderstands the bitter lesson. [0] https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
- heresie-dabord 1y ago"We have to learn the bitter lesson that building in how we think we think does not work in the long run. The bitter lesson is based on the historical observations that 1) AI researchers have often tried to build knowledge into their agents, 2) this always helps in the short term, and is personally satisfying to the researcher, but 3) in the long run it plateaus and even inhibits further progress, and 4) breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning. The eventual success is tinged with bitterness, and often incompletely digested, because it is success over a favored, human-centric approach."
- imtringued 1y agoWhat if the meta bitter lesson is that data scaling is just a more extreme form of the human-centric approach of building knowledge into agents? After all, we're telling the model what to say, think and how to behave. A true general method wouldn't rely on humans at all! Human data would be worthless beyond bootstrapping!
- heresie-dabord 1y agoAnother meta bitter lesson: we don't understand ourselves well enough to define and build something that thinks as we do.
- mehulashah 1y agoI’m surprised by the argument. It’s not wrong. You need more data, but that presumes that the task is to pre-train on data. Additional compute is also useful for unearthing tacit capabilities in the models. This requires inference time scaling and post training usually on specific downstream tasks using RL. Sure that generates data, but it’s not the same as the Internet, and can be scaled.
- mdemare 1y agoJust using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on < 1M tokens of math texts, never mind the 70T tokens that GPT5 appears to be trained on. A difference of eight orders of magnitude. In other words, data scarcity is not a fundamental problem, just a problem for the current paradigm.
- flooo 1y agoNow consider that the genius cannot physically interact with the world or the people therein, and uses her eyes only for reading text.
- nosianu 1y agoYes - we train only on a subset of human communication, the one using written symbols (even voice has much much more depth to it), but human brains train on the actual physical world. Human students who only learned some new words but have not (yet) even began to really comprehend a subject will just throw around random words and sentences that sound great but have no basis in reality too. For the same sentence, for example, "We need to open a new factory in country XY", the internal model lighting up inside the brain of someone who has actually participated when this was done previously will be much deeper and larger than that of someone who only heard about it in their course work. That same depth is zero for an LLM, which only knows the relations between words and has no representation of the world. Words alone cannot even begin to represent what the model created from the real-world sensors' data, which on top of the direct input is also based on many times compounded and already-internalized prior models (nobody establishes that new factory as a newly born baby with a fresh neural net, actually, even the newly born has inherited instincts that are all based on accumulated real world experiences, including the complex very structure of the brain). Somewhat similarly, situations reported in comments like this one (client or manager vastly underestimating the effort required to do something): https://news.ycombinator.com/item?id=45123810 https://news.ycombinator.com/item?id=45123810 The internal model for a task of those far removed from actually doing it is very small compared to the internal models of those doing the work, so trying to gauge required effort falls short spectacularly if they also don't have the awareness.
- benob 1y agoStop thinking about text being the data. There are so many other sources, even some that you can generate. https://arxiv.org/pdf/2506.20057 https://arxiv.org/pdf/2506.20057
- JumpCrisscross 1y ago> Stop thinking about text being the data Path #2 in TFA.
- TheDong 1y agoIs the data input into ChatGPT not a large enough source of new data to matter? People are constantly inputting novel data, telling ChatGPT about mistakes it made and suggesting approaches to try, and so on. For local tools, like claude code, it feels like there's an even bigger goldmine of data in that you can have a user ask claude code to do something, and when it fails they do it themselves... and then if only anthropic could slurp up the human-produced correct solution, that would be high quality training data. I know paid claude-code doesn't slurp up local code, and my impression is paid ChatGPT also doesn't use input for training... but perhaps that's the next thing to compromise on in the quest for more data.
- johnecheck 1y agoNO! CLEARLY THE ENTIRE CORPUS OF HUMAN LITERATURE AND THE INTERNET DOESN'T CONTAIN ENOUGH INFORMATION TO EDUCATE AN EXPERT!!!! I JUST NEED ANOTHER BILLION DOLLARS PLS PLS PLS I PROMISE THE SCALING LAWS ARE ACTUALLY LAWS THIS TIME
- EZ-Cheeze 1y agoThe AI companies won't run out of data to train on. Almost every user interaction is a significant source of data. Chains of interactions are even more significant, especially the longer and more sophisticated they are. Yesterday I was given A/B tests from both GPT5-Thinking and Gemini 2.5 Pro, something neither of then had done before. OpenAI also just acquired Statsig for $1.1 billion. Statsig does A/B testing and other analytics. The data scrapped from the Internet and scanned books served its purpose: it bootstrapped something that we all love talking to and discussing ANYTHING with. That's the new source of data and intelligence.
- WesolyKubeczek 1y ago> it bootstrapped something that we all love talking to and discussing ANYTHING with. We all? Speak for yourself, dude
- EZ-Cheeze 1y agoHyperbole, I meant a huge number of people, not ALL people. I'm into advaita vedanta, priority cosmopsychism, open individualism
- nahuel0x 1y agoAlso to consider, how the massive datasets powering LLMs were generated? For the case of text, it was generations of humans and humans lives, experiences and interactions with the real world that coagulated into masses of text and the language itself.. not to mention the evolutionary process that made that possible. There is an history of biological computation and interaction behind what it seems to be static data.
- d--b 1y agoHow does the brain do it? A baby's brain isn't wired to the entire internet. A 2-year-old has access to at most 2 years of HD video data, plus some other belly-ache and poo-smell stimuli. And a baby's brain has no replay capacity. That's not a lot to work with. Yet, a 2-year-old clearly thinks, is conscious, can understand and create sentences, and wants to annihilate everything just as much as Grok. Sure you can scale data all you want. But there should be enough to work with without scaling like crazy. Having AI know all CSS tricks out there is one thing that requires a lot of data, AGI is different.
- deleted 1y ago[deleted]
- eirikbakke 1y agoHumans require a _lot_ less training data to become, for instance, fluent in English. If a given AI algorithm needs to be trained on the entire Internet to accomplish the same, then it seems safe to assume that the data has not really been "mined out". Generating more training data from the same original data should not be fundamentally problematic in that sense.
- felipeerias 1y agoIt only seems that way because much of the data that humans use is not in a format that computers would understand. A toddler learning to talk is engaging their full body.
- ausbah 1y agohumans also have billions of years of evolution and trillions of organisms to develop a receptacle biased towards learning language
- eirikbakke 1y agoBillions of years of evolution, but still limited to the data that is replicated in human genome/DNA, which is about 3 gigabytes (+epigenome).
- ausbah 1y agoisn’t the problem of not enough data just a problem of not having grounding in the world? world models and everything feel like they’re just dancing around the problem with thin veneers of human-in-the-loop and the verifiable domains we already have
- deleted 1y ago[deleted]
- nobodywillobsrv 1y agoThey mention symmetries and invariances and what not but I wonder if it would be better to clearly emphasize that, in some problems, when you remove certain kinds of symmetries you are are combinatorially worse off. And this is ridiculously bad in some settings. Learning symmetries automatically used to be something I would see people working on but haven't kept up lately.