22 ms·
>What takes the long amount of time and the way to think about it is that it’s a march of nines. Every single nine is a constant amount of work. Every single ni
by Imnimo 1y ago
>What takes the long amount of time and the way to think about it is that it’s a march of nines. Every single nine is a constant amount of work. Every single nine is the same amount of work. When you get a demo and something works 90% of the time, that’s just the first nine. Then you need the second nine, a third nine, a fourth nine, a fifth nine. While I was at Tesla for five years or so, we went through maybe three nines or two nines. I don’t know what it is, but multiple nines of iteration. There are still more nines to go.
I think this is an important way of understanding AI progress. Capability improvements often look exponential on a particular fixed benchmark, but the difficulty of the next step up is also often exponential, and so you get net linear improvement with a wider perspective.
- sdenton4 1y agoHa, I often speak of doing the first 90% of the work, and then moving on to the following 90% of the work...
- inerte 1y agoI use "The project is 90% ready, now we only have to do the other half"
- typpilol 1y ago92% is half actually - RuneScape Players
- JimDabell 1y ago> The first 90 percent of the code accounts for the first 90 percent of the development time. The remaining 10 percent of the code accounts for the other 90 percent of the development time. — Tom Cargill, Bell Labs (September 1985) https://dl.acm.org/doi/pdf/10.1145/4284.315122 https://dl.acm.org/doi/pdf/10.1145/4284.315122
- czk 1y agolike leveling to 99 in old school runescape
- fbrchps 1y agoThe first 92% and the last 92%, exactly.
- zeroonetwothree 1y agoOr Diablo 2
- genewitch 1y agoi don't remember the end-game of the original Diablo; however, in diablo III and IV everyone i've tried to play the game gets bored in the run up to max level. I always tell them "i skip that part as much as possible, because that's not the game. That's just the story!" Once you hit max level in III and IV, the game actually "begins." and to explain the Diablo 2 Reference, the amount of time/effort it takes to go from level 98 to level 99 (the max level), is the same amount of time it takes to go from level 1 to level 98. I've heard "2 weeks" as a rough estimate of "unhealthy playtime", at least solo.
- wilfredk 1y agoPerfect analogy.
- somanyphotons 1y agoThis is an amazing quote that really applies to all software development
- zeroonetwothree 1y agoWell, maybe not all. I’ve definitely built CRUD UIs that were linear in effort. But certainly anything technically challenging or novel.
- Veserv 1y agoDrawn from Karpathy killing a bunch of people by knowingly delivering defective autonomous driving software instead of applying basic engineering ethics and refusing to deploy the dangerous product he was in charge of.
- kordlessagain 1y agoIn Texas, it's a $5K a day fine to call yourself an engineer, even in a job title on LinkedIn, and actually not be a certified engineer.
- bdangubic 1y agohas anyone ever been convicted/fined based on this statute?
- zeroonetwothree 1y agoWhen I worked at Facebook they had a slogan that captured this idea pretty well: “this journey is 1% finished”.
- gowld 1y agoCopied from Amazon's "Day 1".
- fair_enough 1y agoReminds me of a time-honored aphorism in running: A marathon consists of two halves: the first 20 miles, and then the last 10k (6.2mi) when you're more sore and tired than you've ever been in your life.
- tylerflick 1y agoI think I hated life most after 20 miles. Especially in training.
- jakeydus 1y agoThis is 100% unrelated to the original article but I feel like there's an underreported additional first half. As a bigger runner who still loves to run, the first two or three miles before I have enough endorphins to get into the zen state that makes me love running is the first half, then it's 17 miles of this amazing meditative mindset. Then the last 10k sucks.
- awesome_dude 1y agoJust, ftr, endorphins cannot pass the blood brain barrier http://hopkinsmedicine.org/health/wellness-and-prevention/the-truth-behind-runners-high-and-other-mental-benefits-of-running http://hopkinsmedicine.org/health/wellness-and-prevention/th...
- lovecg 1y agoSo a runner’s high is more like a literal high then? Interesting
- awesome_dude 1y agoHa! Endorphins are "endogenous opioid peptides produced by the pituitary and hypothalamus glands that function as the body's natural painkillers and mood regulators". "They are part of the endogenous opioid system", so either way was talking about literal highs. The endocannabinoid system (I hope that I have the spelling correct) is a relatively recent discovery (1980s on), and is quite fascinating on how integral to the human body it is
- ekjhgkejhgk 1y agoThe interview which I've watched recently with Rich Sutton left me with the impression that AGI is not just a matter of adding more 9s. The interviewer had an idea that he took for granted: that to understand language you have to have a model of the world. LLMs seem to udnerstand language therefore they've trained a model of the world. Sutton rejected the premise immediately. He might be right in being skeptical here.
- sysguest 1y agoyeah that "model of the world" would mean: babies are already born with "the model of the world" but a lot of experiments on babies/young kids tell otherwise
- rwj 1y agoLots of experiments show that babies develop import capabilities at roughly the same times. That speaks to inherited abilities.
- ben_w 1y ago> babies are already born with "the model of the world" > but a lot of experiments on babies/young kids tell otherwise I believe they are born with such a model? It's just that model is one where mummy still has fur for the baby to cling on to? And where aged something like 5 to 8 it's somehow useful for us to build small enclosures to hide in, leading to a display of pillow forts in the modern world?
- 1y ago
- jlas 1y agoNotably the scaling law paper shows result graphs on log-scale
- omidsa1 1y agoI also quite like the way he puts it. However, from a certain point onward, the AI itself will contribute to the development—adding nines—and that’s the key difference between this analogy of nines in other systems (including earlier domain‑specific ML ones) and the path to AGI. That's why we can expect fast acceleration to take off within two years.
- AnimalMuppet 1y agoIsn't that one of the measures of when it becomes an AGI? So that doesn't help you with however many nines we are away from getting an AGI. Even if you don't like that definition, you still have the question of how many nines we are away from having an AI that can contribute to its own development. I don't think you know the answer to that. And therefore I think your "fast acceleration within two years" is unsupported, just wishful thinking. If you've got actual evidence, I would like to hear it.
- scragz 1y agoAGI is when it is general. a narrow AI trained only on coding and training AIs would contribute to the acceleration without being AGI itself.
- ben_w 1y agoAI has been helping with the development of AI ever since at least the first optimising compiler or formal logic circuit verification program. Machine learning has been helping with the development of machine learning ever since hyper-parameter optimisers became a thing. Transformers have been helping with the development of transformer models… I don't know exactly, but it was before ChatGPT came out. None of the initials in AGI are booleans. But I do agree that: > "fast acceleration within two years" is unsupported, just wishful thinking Nobody has any strong evidence of how close "it" is, or even a really good shared model of what "it" even is.
- Yoric 1y agoIt's a possibility, but far from certainty. If you look at it differently, assembly language may have been one nine, compilers may have been the next nine, successive generations of language until ${your favorite language} one more nine, and yet, they didn't get us noticeably closer to AGI.
- breve 1y ago[flagged]
- wcoenen 1y agoThis is not exactly new information[1]. You may have a point that it was not presented to customers this way though. [1] https://x.com/elonmusk/status/1382458022367162370 https://x.com/elonmusk/status/1382458022367162370
- breve 1y agoThe lie is older than that: https://web.archive.org/web/20161020091022/https://tesla.com/blog/all-tesla-cars-being-produced-now-have-full-self-driving-hardware https://web.archive.org/web/20161020091022/https://tesla.com... https://motherfrunker.ca/fsd/ https://motherfrunker.ca/fsd/ The fact that lie is old only makes it worse that Musk, Karpathy, and Tesla generally have still not taken responsibility for the lie. They are still not willing to refund the money they took for something they did not deliver.
- jakeydus 1y agoYou know what they say, a Silicon Valley 9 is a 10 anywhere else. Or something like that.
- Yoric 1y agoI assume you're describing the fact that Silicon Valley culture keeps pushing out products before they're fully baked?
- jakeydus 1y agoIt was a half-baked joke along the lines of "A New York 6 is a Utah 10" or something like that, one of those that crops up in a sitcom every once in a while. Like a silicon valley product, I should have let it develop further before pushing it out for public consumption.
- tekbruh9000 1y agoInfinitely big little numbers Academia has rediscovered itself Signal attenuation, a byproduct of entropy, due to generational churn means there's little guarantee. Occam's Razor; Karpathy knows the future or he is self selecting biology trying to avoid manual labor? His statements have more in common with Nostradamus. It's the toxic positivity form of "the end is nigh". It's "Heaven exists you just have to do this work to get there." Physics always wins and statistics is not physics. Gamblers fallacy; improvement of statistical odds does not improve probability. Probability remains the same this is all promises of some people who have no idea or interest in doing anything else with their lives; so stay the course.
- startupsfail 1y ago>> Heaven exists you just have to do this work to get there. Or perhaps Karpathy has a higher level understanding and can see a bigger picture? You've said something about heaven. Are you able to understand this statement, for example: "Heaven is a memeplex, it exists." ?
- tekbruh9000 1y ago"Higher" than an EE with an MSc in elastic structures, ~30 years industry experience, now working with PhDs across the spectrum on energy models to embed in chips? Energy models in part, inferred from categorization of LLM contents and compression of those contents into geometric functions like I described? "Higher level" implies acceptance of geometric structure. You place tokens like a Chomsky diagrams at each step up and down, where you should see parameters to transform geometry of the structure. My team works "above" the contrived state management of software workers to more efficiently sync memory matrix to display matrix. LLMs are a form of compression [1]. My team is working on compressing them further into sets of points that make up each glyph and functions to recreate them. Electromagnetic geometry transforms hardcoded[2] into hardware so reduce energy use of all the outdated string mangling of software dev as most know it. What's higher level, relative to our machines, than design and implementation of the machine? DnD dungeon master versus WOTC game designer. Notice outside how there are no words and philosophy? Just color gradient and geometry? Notice inside the human body no philosophy or words? Language is not intelligence it's an emergent phenomena of geometry created by fundamental forces of physics organizing matter at various speeds relative to light. You've read too much into an ultimately arbitrary statement meant to invoked a subtext, a subtle emotion context. You think of language as Legos, when it is music to feel. [1] https://arxiv.org/abs/2309.10668 https://arxiv.org/abs/2309.10668 [2] https://iopscience.iop.org/article/10.1088/1742-6596/2987/1/012001/pdf https://iopscience.iop.org/article/10.1088/1742-6596/2987/1/...
- godelski 1y agoIt's a good way to think about lots of things. It's Pareto efficiency. The 80/20 rule 20% of your effort gets you 80% of the way. But most of your time is spent getting that last 20%. People often don't realize that this is fractal like in nature, as it draws from the power distribution. So of that 20% you still have left, the same holds true. 20% of your time (20% * 80% = 16% -> 36%) to get 80% (80% * 20% => 96%) again and again. The 80/20 numbers aren't actually realistic (or constant) but it's a decent guide. It's also something tech has been struggling with lately. Move fast and break things is a great way to get most of the way there. But you also left a wake of destruction and tabled a million little things along the way. Someone needs to go back and clean things up. Someone needs to revisit those tabled things. While each thing might be little, we solve big problems by breaking them down into little ones. So each big problem is a sum of many little ones, meaning they shouldn't be quickly dismissed. And like the 9's analogy, 99.9% of the time is still 9hrs of downtime a year. It is still 1e6 cases out of 1e9. A million cases is not a small problem. Scale is great and has made our field amazing, but it is a double edged sword. I think it's also something people struggle with. It's very easy to become above average, or even well above average at something. Just trying will often get you above average. It can make you feel like you know way more but the trap is that while in some domains above average is not far from mastery in other domains above average is closer to no skill than it is to mastery. Like how having $100m puts your wealth closer to a homeless person than a billionaire. At $100m you feel way closer to the billionaire because you're much further up than the person with nothing but the curve is exponential.
- 010101010101 1y agohttps://youtu.be/bpiu8UtQ-6E?si=ogmfFPbmLICoMvr3 https://youtu.be/bpiu8UtQ-6E?si=ogmfFPbmLICoMvr3 "I'm closer to LeBron than you are to me."
- red75prime 1y agoThe question is how many nines are humans.
- notTooFarGone 1y agoHumans adapt and become more nines the more they learn about something. Humans also are liable in a lawful sense. This is a huge factor in any AI use case.
- red75prime 1y agoSo, it's not really nines, but the lack of continuous learning and legal issues.
- mcmoor 1y agoContinuous learning seems like one of real criteria of AGI
- yoyohello13 1y agoI think a ton of people see a line going up and they think exponential. When in Reality, the vast majority of the time it’s actually logistic.
- tibbar 1y agoGiven the physical limits of the universe and our planet in particular, yeah, this is pretty much always true. The interesting question is: what is that limit, and: how many orders of magnitude are we away from leveling off?
- Ianjit 11mo agoFor reasoning the data suggests models are in the logistic domain? Grok 3 - Grok 3 reasoning: 15% increase in training compute for a 25% uplift in inteligence Grok 3 reasoning - Grok 4: 80% increase in training compute for a 15% uplift in inteligence. Inteligence: Source Atrifical Analysis Training Compute: Source https://youtu.be/MtYsUdfZPMA?t=162 https://youtu.be/MtYsUdfZPMA?t=162
- misnome 1y agoI mean the cost line does look somewhat exponential…
- imadierich 1y ago[dead]
- ojr 1y agoif it works 90% of the time that means it fails 10% of the time, to get to 1% failure rate is a 10x improvement and from 1% failure rate to a 0.1% failure rate is also a 10x improvement First time being hearing it be called "march of nines", did Tesla make the term, I thought it was an Amazon thing
- danielvaughn 1y agoI have a very surface level understanding of AI, and yet this always seemed obvious to me. It's almost a fundamental law of the universe that complexity of any kind has a long tail. So you can get AI to faithfully replicate 90% of a particular domain skill. That's phenomenal, and by itself can yield value for companies. But the journey from 90%-100% is going to be a very difficult march.
- Forgeties79 1y agoThe last mile problem is inescapable!
- tim333 1y agoThe nines comment was in the context of self driving cars which I can see because you are never perfect driving and accidents can be fatal. Some AI is like chess though, where they steadily advance in ELO ranking.
- DanHulton 1y agoThe thing about this, though - cars have been built before. We understand what's necessary to get those 9s. I'm sure there were some new problems that had to be solved along the way, but fundamentally, "build good car" is known to be achievable, so the process of "adding 9s" there makes sense. But this method of AI is still pretty new, and we don't know it's upper limits. It may be that there are no more 9s to add, or that any more 9s cost prohibitively more. We might be effectively stuck at 91.25626726...% forever. Not to be a doomer, but I DO think that anyone who is significantly invested in AI really has to have a plan in case that ends up being true. We can't just keep on saying "they'll get there some day" and acting as if it's true. (I mean you can, just not without consequences.)
- danielmarkbruce 1y agoWhile you are right about the broader (and sort of ill defined) chase toward 'AGI' - another way to look at it is the self driving car - they got there eventually.And, if you work on applications using LLMs you can pretty easily see that Karpathy's sentiment is likely correct. You see it because you do it. Even simple applications are shaped like this, albeit each 9 takes less time than self driving cars for a simple app.. it still feels about right.
- vasco 1y ago> another way to look at it is the self driving car - they got there eventually Current self driving cars only work in American roads. Maybe Canada too, not sure how their roads are. Come to Europe/anywhere else and every other road would be intractable. Much tighter lanes, many turns you have a little mirror to see who's coming on the other side, single car at a time lanes that you need to "understand" who goes first, mountain roads where you sometimes need to reverse for 100m when another car is coming so it's wide enough that they can pass before you can keep going forward, etc. Many things like this that would require another 2 or 3 "nines" as the guy put it than acceptable quality in American huge roads. https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQ4NWItLjc-FtP4012fVczTngYeFLMi8EOnW1b2ru8sA72LyER-wqhJKQlv&s=10 https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQ4NWIt...
- TeMPOraL 1y agoFWIW, Karpathy literally says, multiple times, that he thinks we never left the exponential - that all human progress over last 4+ centuries averages out to that smooth ~2% growth rate exponential curve, that electricity and computing and AI are just ways we keep it going, and we'll continue on that curve for the time being. It's the major point of contention between him and the host (who thinks growth rate will increase).
- rcxdude 1y agoIn my experience with AI it's steeper than that: the jump from 90% to 99% is much harder than the jump from 0 to 90%
- gaindustries 1y ago[dead]
- atleastoptimal 1y agosomething that replaces humans doesn’t need to be 99.9999% reliable, it just has to be better than the humans it replaces.
- rrrrrrrrrrrryan 1y agoBut to be accepted by people, it has to be better than humans in the specific ways that humans are good at things. And less bad than humans in the ways that they're bad at things. When automated solutions fail in strange alien ways, it understandably freaks people out. Nobody wants to worry about if a car will suddenly serve into oncoming traffic because of a sensor malfunction. Comparing incidents-per-miles-driven might make sense from a utilitarian perspective, just isn't good enough for humans to accept replacement tech psychologically, so we do have to chase those 9s until they can handle all the edge cases at least as well as humans.
- atleastoptimal 1y agoWaymo has been growing rapidly. It still makes mistakes, but leas often than humans, and its riders are willing to accept the trade off given the benefits.
- joe_the_user 1y agoThe thing is, the example of the "march of nines" is self-driving cars. These deal with roads and roads are interface between the chaos of the overall world and a system that has quite well-defined rules. I can imagine other task on a human/rules-based "frontier" would have a similar quality. But I think there are others that are going to be inaccessible entirely "until AGI" (or something). Humanoid robots moving freely in human society would an example I think.
- HarHarVeryFunny 1y agoI think the point Andrej was making here is that in some areas, such as self driving, the cost of failure is extremely high (maybe death), so 99.9% reliable doesn't cut it, and therefore doesn't mean you are almost done, or have done 99.9% of the work. It's "The last 10% is 90% of the work" recursively applied. He was also pointing out that the same high cost of failure consideration applies to many software systems (depending on what they are doing/controlling). We may already be at the level where AI coding agents are adequate for some less critical applications, but yet far away from them being a general developer replacement. I see software development as something that uses closer to 100% of your brain than 10% - we may well not see AI coding agents approach human reliability levels until we have human level AGI. The AI snake oil salesmen/CEOs like to throw out competitive coding or math olympiad benchmarks as if they are somehow indicative of the readiness of AI for other tasks, but reliability matters. Nobody dies or loses millions of dollars if you get a math problem wrong.
- kordlessagain 1y agoSo when you say first 9, you mean like Anthropic's uptime on models, right?