116 ms·
OpenAI O3 breakthrough high score on ARC-AGI-PUB
- gxt 2y agoI don't care about some scores going up. Newer models need to stop regressing on tasks they were already good at. 4o sucks at LLVM and related tasks were as legacy GPT 4 is relatively ok at it.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- razodactyl 2y agoGreat. Now we have to think of a new way to move the goalposts.
- tines 2y agoI mean, what else do you call learning?
- deleted 2y ago[deleted]
- Pesthuf 2y agoWell right now, running this model is really expensive, but we should prepare a new cope for when equivalent models no longer are, ahead of time.
- cchance 2y agoYa getting costs down will be the big one, i imagine quantization, distillation and lots and lots of improvements on the compute side both hardware and software wise.
- a_wild_dandan 2y agoLet's just define AI as "whatever computers still can't do." That'll show those dumb statistical parrots!
- dboreham 2y ago[flagged]
- foobarqux 2y agoThis is just as silly as claiming that people "moved the goalposts" when a computer beat Kasparov at chess to claim that it wasn't AGI: it wasn't a good test and some people only realize this after the computer beat Kasparov but couldn't do much else. In this case the ARC maintainers specifically have stated that this is a necessary but not sufficient test of AGI (I personally think it is neither).
- famouswaffles 2y agoIt's not silly. The computer that could beat Kasparov couldn't do anything else so of course it wasn't Artificial General Intelligence. o3 can do much much more. There is nothing narrow about SOTA LLMs. They are already General. It doesn't matter what ARC Maintainers have said. There is no common definition of General that LLMs fail to meet. It's not a binary thing. By the time a single machine covers every little test humanity can devise, what comes out of that is not 'AGI' as the words themselves mean but a General Super Intelligence.
- foobarqux 2y agoIt is silly, the logic is the same: "Only a (world-altering) 'AGI' could do [test]" -> test is passed -> no (world-altering) 'AGI' -> conclude that [test] is not a sufficient test for (world-altering) 'AGI' -> chase new benchmark. If you want to play games about how to define AGI go ahead. People have been claiming for years that we've already reached AGI and with every improvement they have to bizarrely claim anew that now we've really achieved AGI. But after a few months people realize it still doesn't do what you would expect of an AGI and so you chase some new benchmark ("just one more eval"). The fact is that there really hasn't been the type of world-altering impact that people generally associate with AGI and no reason to expect one.
- famouswaffles 2y ago>It is silly, the logic is the same: "Only a (world-altering) 'AGI' could do [test]" -> test is passed -> no (world-altering) 'AGI' -> conclude that [test] is not a sufficient test for (world-altering) 'AGI' -> chase new benchmark. Basically nobody today thinks beating a single benchmark and nothing else will make you a General Intelligence. As you've already pointed out out, even the maintainers of ARC-AGI do not think this. >If you want to play games about how to define AGI go ahead. I'm not playing any games. ENIAC cannot do 99% of the things people use computers to do today and yet barely anybody will tell you it wasn't the first general purpose computer. On the contrary, it is people who seem to think "General" is a moniker for everything under the sun (and then some) that are playing games with definitions. >People have been claiming for years that we've already reached AGI and with every improvement they have to bizarrely claim anew that now we've really achieved AGI. Who are these people ? Do you have any examples at all. Genuine question >But after a few months people realize it still doesn't do what you would expect of an AGI and so you chase some new benchmark ("just one more eval"). What do you expect from 'AGI'? Everybody seems to have different expectations, much of it rooted in science fiction and not even reality, so this is a moot point. What exactly is World Altering to you ? Genuinely, do you even have anything other than a "I'll know it when i see it ?" If you introduce technology most people adopt, is that world altering or are you waiting for Skynet ?
- deleted 2y ago[deleted]
- famouswaffles 2y agoThis is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.
- attentionmech 2y agoI legit see that if there is not even a new breakthrough just one week, people start shouting plateau plateau.. Our rate of progress is extraordinary and any downplay of it seems like stupid
- deleted 2y ago[deleted]
- throwup238 2y ago> So much for a plateau lol. It’s been really interesting to watch all the internet pundits’ takes on the plateau… as if the two years since the release of GPT3.5 is somehow enough data for an armchair ponce to predict the performance characteristics of an entirely novel technology that no one understands.
- jgalt212 2y agoYou could make an equivalently dismissive comment about the hypesters.
- throwup238 2y agoYeah but anyone with half a brain knows to ignore them. Vapid cynicism is a lot more seductive to the average nerd.
- bandwidth-bob 2y agoThe pundits response to the (alleged) plateau was proportional to the certainty with which CEOs of frontier labs discussed pre-training scaling. The o3 result is from scaling test time compute, which represents a meaningful change in how you would build out compute for scaling (single supercluster --> presence in regions close to users). Thus it is important to discuss.
- deleted 2y ago[deleted]
- maxdoop 2y agoHow much longer can I get paid $150k to write code ?
- tsunamifury 2y agoOften what happens is the golf-course phenomenon. As golfing gets less popular, low and mid tier golf courses go out of business as they simply aren't needed. But at the same time demand for high end golf courses actually skyrockets because people who want to golf either can give it up or go higher end. This I think will happen with programmers. Rote programming will slowly die out, while demand for super high end will go dramatically up in price.
- CapcomGo 2y agoWhere does this golf-course phenomenon come from? It doesn't really match the real world or how golfing works.
- tsunamifury 2y agohow so, witnessed it quite directly in California. Majority have closed and remaining have gone up in price and are up scale. This has been covered in various new programs like 60 minutes. You can look up death of golfing. Also unsure what you mean by...'how golfing works'. This is the economics of it, not the game
- EVa5I7bHFq9mnYK 2y agoMaybe its CA thing? Plenty of $50 golf courses here in Phoenix.
- CapcomGo 2y agoGolfing has had a huge surge in popularity since 2020. Prices are going up but courses aren't closing.
- colesantiago 2y ago
- braden-lk 2y agoIf people constantly have to ask if your test is a measure of AGI, maybe it should be renamed to something else.
- deleted 2y ago[deleted]
- OfficialTurkey 2y agoFrom the post > Passing ARC-AGI does not equate achieving AGI, and, as a matter of fact, I don't think o3 is AGI yet. o3 still fails on some very easy tasks, indicating fundamental differences with human intelligence.
- cchance 2y agoIts funny when they say this, as if all humans can solve basic ass question/answer combos, people seem to forget theirs a percentage of the population that honestly believe the world is flat along with other hallucinations at the human level
- jppittma 2y agoI don't believe AGI at that level has any commercial value.
- Jensson 2y agoHumans works in groups, so you are wrong a group of human is extremely reliable on tons of tasks. These AI models also work in groups, or they don't improve from working in a group since the company uses whatever does the best on the benchmark, so it is only fair to compare AI vs group of people, AI compared to an individual will always be an unfair comparison since an AI is never alone.
- modeless 2y agoCongratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that these models can't perform as well as humans on everyday tasks despite blowout scores on the hardest tests we give to humans. Yet nobody could quantify exactly the ways the models were deficient. ARC is the best effort in that direction so far. We don't need more "hard" benchmarks. What we need right now are "easy" benchmarks that these models nevertheless fail. I hope Francois has something good cooked up for ARC 2!
- dtquad 2y agoAre there any single-step non-reasoner models that do well on this benchmark? I wonder how well the latest Claude 3.5 Sonnet does on this benchmark and if it's near o1.
- throwaway71271 2y ago| Name | Semi-private eval | Public eval | |--------------------------------------|-------------------|-------------| | Jeremy Berman | 53.6% | 58.5% | | Akyürek et al. | 47.5% | 62.8% | | Ryan Greenblatt | 43% | 42% | | OpenAI o1-preview (pass@1) | 18% | 21% | | Anthropic Claude 3.5 Sonnet (pass@1) | 14% | 21% | | OpenAI GPT-4o (pass@1) | 5% | 9% | | Google Gemini 1.5 (pass@1) | 4.5% | 8% | https://arxiv.org/pdf/2412.04604 https://arxiv.org/pdf/2412.04604
- kandesbunzler 2y agowhy is this missing the o1 release / o1 pro models? Would love to know how much better they are
- wilg 2y agofun! the benchmarks are so interesting because real world use is so variable. sometimes 4o will nail a pretty difficult problem, other times o1 pro mode will fail 10 times on what i would think is a pretty easy programming problem and i waste more time trying to do it with ai
- behnamoh 2y agoSo now not only are the models closed, but so are their evals?! This is a "semi-private" eval. WTH is that supposed to mean? I'm sure the model is great but I refuse to take their word for it.
- deleted 2y ago[deleted]
- ZeroCool2u 2y agoThe private evaluation set is private from the public/OpenAI so companies can't train on those problems and cheat their way to a high score by overfitting.
- jsheard 2y agoIf the models run on OpenAIs servers then surely they could still see the questions being put into it if they wanted to cheat? That could only be prevented by making the evaluation a one-time deal that can't be repeated, or by having OpenAI distribute their models for evaluators to run themselves, which I doubt they're inclined to do.
- foobarqux 2y agoYes that's why it is "semi"-private: From the ARC website "This set is "semi-private" because we can assume that over time, this data will be added to LLM training data and need to be periodically updated." I presume evaluation on the test set is gated (you have to ask ARC to run it).
- cchance 2y agothe evals are the question/answers, ARC-AGI doesn't share the questions and answers for a portion so that models can't be trained on them, the public ones... the public knows the questions so theres a chance they could have been at least partially been trained on the question (if not the actual answer). Thats how i understand it
- neom 2y agoWhy would they give a cost estimate per task on their low compute mode but not their high mode? "low compute" mode: Uses 6 samples per task, Uses 33M tokens for the semi-private eval set, Costs $17-20 per task, Achieves 75.7% accuracy on semi-private eval The "high compute" mode: Uses 1024 samples per task (172x more compute), Cost data was withheld at OpenAI's request, Achieves 87.5% accuracy on semi-private eval Can we just extrapolate $3kish per task on high compute? (wondering if they're withheld because this isn't the case?)
- WiSaGaN 2y agoThe withheld part is really a red flag for me. Why do you want to withhold a compute number?
- zebomon 2y agoMy initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is with any sufficiently complex program, if you could discern the result beforehand, you wouldn't have had to execute the program in the first place. I'm not trying to be a downer on the 12th day of Christmas. Perhaps because my first instinct is childlike excitement, I'm trying to temper it with a little reason.
- amarcheschi 2y agoI just googled arc agi questions, and it looks like it is similar to an iq test with raven matrix. Similar as in you have some examples of images before and after, then an image before and you have to guess the after. Could anyone confirm if this is the only kind of questions in the benchmark? If yes, how come there is such a direct connection to "oh this performs better than humans" when llm can be quite better than us in understanding and forecasting patterns? I'm just curious, not trying to stir up controversies
- zebomon 2y agoIt's a test on which (apparently until now) the vast majority of humans have far outperformed all machine systems.
- patrickhogan1 2y agoBut it’s not a test that directly shows general intelligence. I am excited no less! This is huge improvement. How does this do on SWE Bench?
- famouswaffles 2y ago>How does this do on SWE Bench? 71.7%
- attentionmech 2y agoIsn't this at the level now where it can sort of self improve. My guess is that they will just use it to improve the model and the cost they are showing per evaluation will go down drastically. So, next step in reasoning is open world reasoning now?
- dyauspitr 2y agoI don’t believe so. If it’s at the point where you could just plug it into a bunch of camera feeds around the world and it could only filter out a useful training set for itself out of that data then we truly would have AGI. I don’t think it’s there yet.
- attentionmech 2y agomay be for some sub-domains like math and code it can do that since the verification process can be done / relatively tractable
- yawnxyz 2y agoO3 High (tuned) model scored an 88% at what looks like $6,000/task haha I think soon we'll be pricing any kind of tasks by their compute costs. So basically, human = $50/task, AI = $6,000/task, use human. If AI beats human, use AI? Ofc that's considering both get 100% scores on the task
- deleted 2y ago[deleted]
- cchance 2y agoIsn't that generally what ... all jobs are? Automation Cost vs Longterm Human cost... its why amazon did the weird "our stores are AI driven" but in reality was cheaper to higher a bunch of guys in a sweat shop to look at the cameras and write things down lol. The thing is given what we've seen from distillation and tech, even if its 6,000/task... that will come down drastically over time through optimization and just... faster more efficient processing hardware and software.
- cryptoegorophy 2y agoI remember hearing Tesla trying to automate all of production but some things just couldn’t , like the wiring which humans still had to do.
- dyauspitr 2y agoCompute can get optimized and cheap quickly.
- karmasimida 2y agoIs it? The moore’s law is dead dead, I don’t think this is a given.
- jsheard 2y agoThat's the elephant in the room with the reasoning/COT approach, it shifts what was previously a scaling of training costs into scaling of training and inference costs. The promise of doing expensive training once and then running the model cheaply forever falls apart once you're burning tens, hundreds or thousands of dollars worth of compute every time you run a query.
- philip1209 2y ago[flagged]
- deleted 2y ago[deleted]
- Delmolokolo 2y ago[dead]
- spaceman_2020 2y agoJust as an aside, I've personally found o1 to be completely useless for coding. Sonnet 3.5 remains the king of the hill by quite some margin
- cchance 2y agoThe new gemini's are pretty good too
- lysecret 2y agoActually prefer new geminis too. 2.0 experimental especially.
- spaceman_2020 2y agoThe new ai studio from Google is fantastic
- famouswaffles 2y agoTo be fair, until the last checkpoint released 2 days ago, o1 didn't really beat sonnet (and if so, barely) in most non-competitive coding benchmarks
- vessenes 2y agoTo fill this out, I find o1-pro (and -preview when it was live) to be pretty good at filling in blindspots/spotting holistic bugs. I use Claude for day to day, and when Claude is spinning, o1 often can point out why. It's too slow for AI coding, and I agree that at default its responses aren't always satisfying. That said, I think its code style is arguably better, more concise and has better patterns -- Claude needs a fair amount of prompting and oversight to not put out semi-shitty code in terms of structure and architecture. In my mind: going from Slowest to Fastest, and Best Holistically to Worst, the list is: 1. o1-pro 2. Claude 3.5 3. Gemini 2 Flash Flash is so fast, that it's tempting to use more, but it really needs to be kept to specific work on strong codebases without complex interactions.
- spaceman_2020 2y ago
- smy20011 2y agoIt seems O3 following trend of Chess engine that you can cut your search depth depends on state. It's good for games with clear signal of success (Win/Lose for Chess, tests for programming). One of the blocker for AGI is we don't have clear evaluation for most of our tasks and we cannot verify them fast enough.
- flakiness 2y agoThe cost axis is interesting. O3 Low is $10+ per task and 03 High is over $1000 (it's logarithmic graph so it's like $50 and $5000 respectively?)
- obblekk 2y agoHuman performance is 85% [1]. o3 high gets 87.5%. This means we have an algorithm to get to human level performance on this task. If you think this task is an eval of general reasoning ability, we have an algorithm for that now. There's a lot of work ahead to generalize o3 performance to all domains. I think this explains why many researchers feel AGI is within reach, now that we have an algorithm that works. Congrats to both Francois Chollet for developing this compelling eval, and to the researchers who saturated it! [1] https://x.com/SmokeAwayyy/status/1870171624403808366 https://x.com/SmokeAwayyy/status/1870171624403808366, https://arxiv.org/html/2409.01374v1 https://arxiv.org/html/2409.01374v1
- phillipcarter 2y agoAs excited as I am by this, I still feel like this is still just a small approximation of a small chunk of human reasoning ability at large. o3 (and whatever comes next) feels to me like it will head down the path of being a reasoning coprocessor for various tasks. But, still, this is incredibly impressive.
- qt31415926 2y agoWhich parts of reasoning do you think is missing? I do feel like it covers a lot of 'reasoning' ground despite its on the surface simplicity
- mistermann 2y agoOptimal phenomenological reasoning is going to be a tough nut to crack. Luckily we don't know the problem exists, so in a cultural/phenomenological sense it is already cracked.
- phillipcarter 2y agoI think it's hard to enumerate the unknown, but I'd personally love to see how models like this perform on things like word problems where you introduce red herrings. Right now, LLMs at large tend to struggle mightily to understand when some of the given information is not only irrelevant, but may explicitly serve to distract from the real problem.
- Imnimo 2y agoWhenever a benchmark that was thought to be extremely difficult is (nearly) solved, it's a mix of two causes. One is that progress on AI capabilities was faster than we expected, and the other is that there was an approach that made the task easier than we expected. I feel like the there's a lot of the former here, but the compute cost per task (thousands of dollars to solve one little color grid puzzle??) suggests to me that there's some amount of the latter. Chollet also mentions ARC-AGI-2 might be more resistant to this approach. Of course, o3 looks strong on other benchmarks as well, and sometimes "spend a huge amount of compute for one problem" is a great feature to have available if it gets you the answer you needed. So even if there's some amount of "ARC-AGI wasn't quite as robust as we thought", o3 is clearly a very powerful model.
- exe34 2y ago> the other is that there was an approach that made the task easier than we expected. from reading Dennett's philosophy, I'm convinced that that's how human intelligence works - for each task that "only a human could do that", there's a trick that makes it easier than it seems. We are bags of tricks.
- Jensson 2y ago> We are bags of tricks. We are trick generators, that is what it means to be a general intelligence. Adding another trick in the bag doesn't make you a general intelligence, being able to discover and add new tricks yourself makes you a general intelligence.
- falcor84 2y agoNot the parent, but remembering my reading of Dennett, he was referring to the tricks that we got through evolution, rather than ones we invented ourselves. As particular examples, we have neural functional areas for capabilities like facial recognition and spatial reasoning which seems to rely on dedicated "wetware" somewhat distinct from other parts of the brain.
- 2y ago
- whoistraitor 2y agoThe general message here seems to be that inference-time brute-forcing works as long as you have a good search and evaluation strategy. We’ve seemingly hit a ceiling on the base LLM forward-pass capability so any further wins are going to be in how we juggle multiple inferences to solve the problem space. It feels like a scripting problem now. Which is cool! A fun space for hacker-engineers. Also: > My mental model for LLMs is that they work as a repository of vector programs. When prompted, they will fetch the program that your prompt maps to and "execute" it on the input at hand. LLMs are a way to store and operationalize millions of useful mini-programs via passive exposure to human-generated content. I found this such an intriguing way of thinking about it.
- whimsicalism 2y ago> We’ve seemingly hit a ceiling on the base LLM forward-pass capability so any further wins are going to be in how we juggle multiple inferences to solve the problem space Not so sure - but we might need to figure out the inference/search/evaluation strategy in order to provide the data we need to distill to the single forward-pass data fitting.
- cchance 2y agoIs it just me or does looking at the ARC-AGI example questions at the bottom... make your brain hurt?
- drdaeman 2y agoLooks pretty obvious to me, although, of course, it took me a few moments to understand what's expected as a solution. c6e1b8da is moving rectangular figures by a given vector, 0d87d2a6 is drawing horizontal and/or vertical lines (connecting dots at the edges) and filling figures they touch, b457fec5 is filling gray figures with a given repeating color pattern. This is pretty straightforward stuff that doesn't require much spatial thinking or keeping multiple things/aspects in memory - visual puzzles from various "IQ" tests are way harder. This said, now I'm curious how SoTA LLMs would do on something like WAIS-IV.
- randyrand 2y agoI'll sound like a total douche bag - but I thought they were incredibly obvious - which I think is the point of them. What took me longer was figuring out how the question was arranged, i.e. left input, right output, 3 examples each
- airstrike 2y agoUhh...some of us are apparently living under a rock, as this is the first time I hear about o3 and I'm on HN far too much every day
- burningion 2y agoI think it was just announced today! You're fine!
- deleted 2y ago[deleted]
- cryptoegorophy 2y agoBesides higher scores - is there any improvements for a general use? Like asking to help setup home assistant etc etc?
- rvz 2y agoGreat results. However, let's all just admit it. It has well replaced journalists, artists and on its way to replace nearly both junior and senior engineers. The ultimate intention of "AGI" is that it is going to replace tens of millions of jobs. That is it and you know it. It will only accelerate and we need to stop pretending and coping. Instead lets discuss solutions for those lost jobs. So what is the replacement for these lost jobs? (It is not UBI or "better jobs" without defining them.)
- neom 2y agoDo you follow Jack Clark? I noticed he's been on the road a lot talking to governments and policy makers, and not just in the "AI is coming" way he used to talk.
- whynotminot 2y agoWhen none of us have jobs or income, there will be no ability for us to buy products. And then no reason for companies to buy ads to sell products to people who don’t have money. Without ad money (or the potential of future ad money), the people pushing the bounds of AGI into work replacement will lose the very income streams powering this research and their valuations. Ford didn’t support a 40 hour work week out of the kindness of his heart. He wanted his workers to have time off for buying things (like his cars). I wonder if our AGI industrialist overlords will do something similar for revenue sharing or UBI.
- whimsicalism 2y agoThis picture doesn't make sense. If most don't have any money to buy products, just invent some other money and start paying one of the other people who doesn't have any money to start making the products for you. In reality, if there really is mass unemployment, AI driven automation will make consumables so cheap that anyone will be able to buy it.
- whynotminot 2y ago> This picture doesn't make sense. If most don't have any money to buy products, just invent some other money and start paying one of the other people who doesn't have any money to start making the products for you. Uh, this picture doesn’t make sense. Why would anyone value this randomly invented money?
- mensetmanusman 2y agoI’m super curious as to whether this technology completely destroys the middle class, or if everyone becomes better off because productivity is going to skyrocket.
- mhogers 2y agoIs anyone here aware of the latest research that tries to predict the outcome? Please share - super curious as well
- te_chris 2y agoThere’s this https://arxiv.org/pdf/2312.05481v9 https://arxiv.org/pdf/2312.05481v9
- pdfernhout 2y agoSome thoughts I put together on all this circa 2010: https://pdfernhout.net/beyond-a-jobless-recovery-knol.html https://pdfernhout.net/beyond-a-jobless-recovery-knol.html "This article explores the issue of a "Jobless Recovery" mainly from a heterodox economic perspective. It emphasizes the implications of ideas by Marshall Brain and others that improvements in robotics, automation, design, and voluntary social networks are fundamentally changing the structure of the economic landscape. It outlines towards the end four major alternatives to mainstream economic practice (a basic income, a gift economy, stronger local subsistence economies, and resource-based planning). These alternatives could be used in combination to address what, even as far back as 1964, has been described as a breaking "income-through-jobs link". This link between jobs and income is breaking because of the declining value of most paid human labor relative to capital investments in automation and better design. Or, as is now the case, the value of paid human labor like at some newspapers or universities is also declining relative to the output of voluntary social networks such as for digital content production (like represented by this document). It is suggested that we will need to fundamentally reevaluate our economic theories and practices to adjust to these new realities emerging from exponential trends in technology and society."
- tivert 2y ago> I’m super curious as to whether this technology completely destroys the middle class, or if everyone becomes better off because productivity is going to skyrocket. Even if productivity skyrockets, why would anyone assume the dividends would be shared with the "destroy[ed] middle class"? All indications will be this will end up like the China Shock: "I lost my middle class job, and all I got was the opportunity to buy flimsy pieces of crap from a dollar store." America lacks the ideological foundations for any other result, and the coming economic changes will likely make building those foundations even more difficult if not impossible.
- croemer 2y agoThe programming task they gave o3-mini high (creating Python server that allows chatting with OpenAI API and run some code in terminal) didn't seem very hard? Strange choice of example for something that's claimed to be a big step forwards. YT timestamped link: https://www.youtube.com/watch?v=SKBG1sqdyIU&t=768s https://www.youtube.com/watch?v=SKBG1sqdyIU&t=768s (thanks for the fixed link @photonboom) Updated: I gave the task to Claude 3.5 Sonnet and it worked first shot: https://claude.site/artifacts/36cecd49-0e0b-4a8c-befa-faa5aaa102e6 https://claude.site/artifacts/36cecd49-0e0b-4a8c-befa-faa5aa...
- bearjaws 2y agoIt's good that it works since if you ask GPT-4o to use the openai sdk it will often produce invalid and out of date code.
- luke-stanley 2y agoBut they did use a prompt that included a full example of how to call their latest model and API!
- m3kw9 2y agoI would say they didn’t need to demo anything, because if you are gonna use the output code live on a demo it may make compile errors and then look stupid trying to fix it live
- croemer 2y agoIf it was a safe bet problem, then they should have said that. To me it looks like they faked excitement for something not exciting which lowers credibility of the whole presentation.
- sunaookami 2y agoThey actually did that the last time when they showed the apps integration. First try in Xcode didn't work.
- Bjorkbat 2y agoI was impressed until I read the caveat about the high-compute version using 172x more compute. Assuming for a moment that the cost per task has a linear relationship with compute, then it costs a little more than $1 million to get that score on the public eval. The results are cool, but man, this sounds like such a busted approach.
- futureshock 2y agoSo what? I’m serious. Our current level of progress would have been sci-fi fantasy with the computers we had in 2000. The cost may be astronomical today, but we have proven a method to achieve human performance on tests of reasoning over novel problems. WOW. Who cares what it costs. In 25 years it will run on your phone.
- Bjorkbat 2y agoIt's not so much the cost as much the fact that they got a slightly better result by throwing 172x more compute per/task. The fact that it may have cost somewhere north of $1 million simply helps to give a better idea of how absurd the approach is. It feels a lot less like the breakthrough when the solution looks so much like simply brute-forcing. But you might be right, who cares? Does it really matter how crude the solution is if we can achieve true AGI and bring the cost down by increasing the efficiency of compute?
- futureshock 2y ago“Simply brute-forcing” That’s the thing that’s interesting to me though and I had the same first reaction. It’s a very different problem than brute-forcing chess. It has one chance to come to the correct answer. Running through thousands or millions of options means nothing if the model can’t determine which is correct. And each of these visual problems involve combinations of different interacting concepts. To solve them requires understanding, not mimicry. So no matter how inefficient and “stupid” these models are, they can be said to understand these novel problems. That’s a direct counter to everyone who ever called these a stochastic parrot and said they were a dead-end to AGI that was only searching an in distribution training set. The compute costs are currently disappointing, but so was the cost of sequencing the first whole human genome. That went from 3 billion to a few hundred bucks from your local doctor.
- tripletao 2y agoTheir discussion contains an interesting aside: > Moreover, ARC-AGI-1 is now saturating – besides o3's new score, the fact is that a large ensemble of low-compute Kaggle solutions can now score 81% on the private eval. So while these tasks get greatest interest as a benchmark for LLMs and other large general models, it doesn't yet seem obvious those outperform human-designed domain-specific approaches. I wonder to what extent the large improvement comes from OpenAI training deliberately targeting this class of problem. That result would still be significant (since there's no way to overfit to the private tasks), but would be different from an "accidental" emergent improvement.
- onemetwo 2y agoIn (1) the author use a technique to improve the performance of an LLM, he trained sonnet 3.5 to obtain 53,6% in the arc-agi-pub benchmark moreover he said that more computer power would give better results. So the results of o3 could be produced in this way using the same method with more computer power, so if this is the case the result of o3 is not very interesting. (1) https://params.com/@jeremy-berman/arc-agi https://params.com/@jeremy-berman/arc-agi
- TypicalHog 2y agoThis is actually mindblowing!
- blixt 2y agoThese results are fantastic. Claude 3.5 and o1 are already good enough to provide value, so I can't wait to see how o3 performs comparatively in real-world scenarios. But I gotta say, we must be saturating just about any zero-shot reasoning benchmark imaginable at this point. And we will still argue about whether this is AGI, in my opinion because these LLMs are forgetful and it's very difficult for an application developer to fix that. Models will need better ways to remember and learn from doing a task over and over. For example, let's look at code agents: the best we can do, even with o3, is to cram as much of the code base as we can fit into a context window. And if it doesn't fit we branch out to multiple models to prune the context window until it does fit. And here's the kicker – the second time you ask for it to do something this all starts over from zero again. With this amount of reasoning power, I'm hoping session-based learning becomes the next frontier for LLM capabilities. (There are already things like tool use, linear attention, RAG, etc that can help here but currently they come with downsides and I would consider them insufficient.)
- vessenes 2y ago[flagged]
- jamiek88 2y ago> most of the people training these next-gen AIs are neurodiverse Citation needed. This is a huge claim based only on stereotype.
- vessenes 2y agoSo true. Perhaps I'm just thinking it's my people and need to update my priors.
- getpost 2y ago> most of the people training these next-gen AIs are neurodiverse and we are training the AI in our own image Do you have any evidence to support that? It would be fascinating if the field is primarly advancing due to a unique constellation of traits contributed by individuals who, in the past, may not have collaborated so effectively.
- vessenes 2y agoPURELY Anecdotal. But I'll say that as of 2024 1 in 36 US children are diagnosed on the spectrum according to the CDC(!), which would mean if you met 10 AI researchers and 4 were neurodivergent you'd reasonably expect that it's a higher-than-population average representation. I'm polling from the Effective Altruist AI folks in my mind, and the number is definitely, definitely higher than 4/10.
- EVa5I7bHFq9mnYK 2y agoAre there non-Effective Altruist AI folks?
- vessenes 2y agoI love how this might mean "non-Effective", non-"Effective Altruist" or non-"Effective Altruist AI" folks. Yes
- nopinsight 2y agoLet me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they deal with unfamiliar domains, although that is not the case with these models.) ARC has been challenging precisely because solving its problems often requires: 1) using multiple different *kinds* of core knowledge [1], such as symmetry, counting, color, AND 2) using the right level(s) of abstraction Achieving human-level performance in the ARC benchmark, as well as top human performance in GPQA, Codeforces, AIME, and Frontier Math suggests the model can potentially solve any problem at the human level if it possesses essential knowledge about it. Yes, this includes out-of-distribution problems that most humans can solve. It might not yet be able to generate highly novel theories, frameworks, or artifacts to the degree that Einstein, Grothendieck, or van Gogh could. But not many humans can either. [1] https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf https://www.harvardlds.org/wp-content/uploads/2017/01/Spelke... ADDED: Thanks to the link to Chollet's posts by lswainemoore below. I've analyzed some easy problems that o3 failed at. They involve spatial intelligence, including connection and movement. This skill is very hard to learn from textual and still image data. I believe this sort of core knowledge is learnable through movement and interaction data in a simulated world and it will not present a very difficult barrier to cross. (OpenAI purchased a company behind a Minecraft clone a while ago. I've wondered if this is the purpose.)
- xvector 2y agoAgree. AGI is here. I feel such a sense of pride in our species.
- deleted 2y ago[deleted]
- timabdulla 2y agoWhat's your explanation for why it can only get ~70% on SWE-bench Verified? I believe about 90% of the tasks were estimated by humans to take less than one hour to solve, so we aren't talking about very complex problems, and to boot, the contamination factor is huge: o3 (or any big model) will have in-depth knowledge of the internals of these projects, and often even know about the individual issues themselves (e.g. you can say what was Github issue #4145 in project foo, and there's a decent chance it can tell you exactly what the issue was about!)
- CliveBloomers 2y agoAnother meaningless benchmark, another month—it’s like clockwork at this point. No one’s going to remember this in a month; it’s just noise. The real test? It’s not in these flashy metrics or minor improvements. The only thing that actually matters is how fast it can wipe out the layers of middle management and all those pointless, bureaucratic jobs that add zero value. That’s the true litmus test. Everything else? It’s just fine-tuning weights, playing around the edges. Until it starts cutting through the fat and reshaping how organizations really operate, all of this is just more of the same.
- handfuloflight 2y agoAgreed, but isn't it management who decides that this would be implemented? Are they going to propogate their own removal?
- zamadatix 2y agoMiddle manager types are probably interested in their salary performance more than anything. "Real" management (more of their assets come from their ownership of the company than a salary) will override them if it's truthfully the best performing operating model for the company.
- oytis 2y agoSo far AI market seems to be focused on replacing meaningful jobs, meaningless ones look safe (which kind of makes sense if you think about it).
- akra 2y agoIts a common view from the "do'ers" (the people who made most of the value in the past; the hard workers, etc) that this will make management redundant. Sadly with a basic understanding of economics you can see this is probably wrong. The "do'ers" have given more power to the management class at their own expense with this solution - if I can get the AI to "do" all I need are the people who "decide what to do". Market power belongs with scarcity - all else being equal AI makes the barrier to development smaller meaning less scarcity on that side. In general technology developments have increased inequality especially since the 90's onwards. Generally with AI think the top of society stand to gain a lot more than the middle/bottom of it for a whole host of reasons. If you think anything different your framework you use to make your conclusion is probably wrong at least in IMO. I don't like saying this but there is a reason why the "AI bros", VC's, big tech CEO's, etc are all very very excited about this and many employees (some commenting here) are filled with dread/fear. The sales people, the managers, the MBA's, etc stand to gain a lot from this. Fear also serves as the best marketing tool; it makes people talk and spread OpenAI's news more so than everything else. Its a reason why targeting coding jobs/any jobs is so effective. I want to be wrong of course.
- 6gvONxR4sf7o 2y agoI'm glad these stats show a better estimate of human ability than just the average mturker. The graph here has the average mturker performance as well as a STEM grad measurement. Stuff like that is why we're always feeling weird that these things supposedly outperform humans while still sucking. I'm glad to see 'human performance' benchmarked with more variety (attention, time, education, etc).
- RivieraKid 2y agoIt sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.
- xvector 2y agoHumanity is about to enter an even steeper hockey stick growth curve. Progressing along the Kardashev scale feels all but inevitable. We will live to see Longevity Escape Velocity. I'm fucking pumped and feel thrilled and excited and proud of our species. Sure, there will be growing pains, friction, etc. Who cares? There always is with world-changing tech. Always.
- drcode 2y agolongevity for the AIs
- tokioyoyo 2y agoMy job should be secure for a while, but why would an average person give a damn about humanity when they might lose their jobs and comfort levels? If I had kids, I would absolutely hate this uncertainty as well. “Oh well, I guess I can’t give the opportunities to my kid that I wanted, but at least humanity is growing rapidly!”
- xvector 2y ago> when they might lose their jobs and comfort levels? Everyone has always worried about this for every major technology throughout history IMO AGI will dramatically increase comfort levels, lower your chance of dying, death, disease, etc.
- tokioyoyo 2y agoAgain, sure, but it doesn’t matter to an average person. That’s too much focus on the hypothetical future. People care about the current times. In the short term it will suck for a good chunk of people, and whether the sacrifice is worth it will depend on who you are. People aren’t really on uproar yet, because implementations haven’t affected the job market of the masses. Afterwards? Tume will show.
- bluecoconut 2y agoEfficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective experience) between 5 second and 5 minutes to solve the task. (So i'd argue a human is at 0.03USD - 1.67USD per puzzle at 20USD/hr, and they include in their document an average mechancal turker at $2 USD task in their document) Going the other direction: I am interpreting this result as human level reasoning now costs (approximately) 41k/hr to 2.5M/hr with current compute. Super exciting that OpenAI pushed the compute out this far so we could see he O-series scaling continue and intersect humans on ARC, now we get to work towards making this economical!
- riku_iki 2y ago> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
- bluecoconut 2y agoThat's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k
- ionwake 2y agosorry to be a noob, but can someone tell me doe sths mena o3 will be unaffordable for a typical user? Will only companies with thousands to spend per query be able to use this? Sorry for being thick Im just confused how they can turn this into an addordable service?
- JohnnyMarcone 2y agoThere are likely many efficiency gains that will be made before it's released, and after. Also they showed o3 mini to be better than o1 for less cost in multiple benchmarks, so there're already improvements there at a lower cost than what available.
- aithrowawaycomm 2y agoI would like to see this repeated with my highly innovative HARC-HAGI, which is ARC-AGI but it uses hexagons instead of squares. I suspect humans would only make slightly more brain farts on HARC-HAGI than ARC-AGI, but O3 would fail very badly since it almost certainly has been specifically trained on squares. I am not really trying to downplay O3. But this would be a simple test as to whether O3 is truly "a system capable of adapting to tasks it has never encountered before" versus novel ARC-AGI tasks it hasn't encountered before.
- falcor84 2y agoHere's my take - even if the o3 as currently implemented is utterly useless on your HARC-HAGI, it is obvious that o3 coupled with its existing training pipeline trained briefly on the hexagons would excel on it, such that passing your benchmark doesn't require any new technology. Taking this a level of abstraction higher, I expect that in the next couple of years we'll see systems like o3 given a runtime budget that they can use for training/fine-tuning smaller models in an ad-hoc manner.
- botro 2y agoThe LLM community has come up with tests they call 'Misguided Attention'[1] where they prompt the LLM with a slightly altered version of common riddles / tests etc. This often causes the LLM to fail. For example I used the prompt "As an astronaut in China, would I be able to see the great wall?" and since the training data for all LLMs is full of text dispelling the common myth that the great wall is visible from space, LLMs do not notice the slight variation that the astronaut is IN China. This has been a sobering reminder to me as discussion of AGI heats up. [1] https://github.com/cpldcpu/MisguidedAttention https://github.com/cpldcpu/MisguidedAttention
- kizer 2y agoIt could be that it “assumed” you meant “from China”; in the higher level patterns it learns the imperfection of human writing and the approximate threshold at which mistakes are ignored vs addressed by training on conversations containing these types of mistakes; e.g Reddit. This is just a thought. Try saying: As an astronaut in Chinese territory; or as an astronaut on Chinese soil. Another test would be to prompt it to interpret everything literally as written.
- deleted 2y ago[deleted]
- dwaltrip 2y agoInteresting... It took me 3 different attempts, but I found a set of custom instructions that allowed Claude to get the right answer on the initial prompt. Here's the instructions (I tried to keep them as general and non-specific as I could): Carefully analyze questions to not overlook subtle details. Take each question "as-is", don't guess what they mean -- interpret them as any reasonable person would.
- deleted 2y ago[deleted]
- whimsicalism 2y agoWe need to start making benchmarks in memory & continued processing over a task over multiple days, handoffs, etc (ie. 'agentic' behavior). Not sure how possible this is.
- slibhb 2y agoInteresting about the cost: > Of course, such generality comes at a steep cost, and wouldn't quite be economical yet: you could pay a human to solve ARC-AGI tasks for roughly $5 per task (we know, we did that), while consuming mere cents in energy. Meanwhile o3 requires $17-20 per task in the low-compute mode.
- imranq 2y agoBased on the chart, the Kaggle SOTA model is far more impressive. These O3 models are more expensive to run than just hiring a mechanical turk worker. It's nice we are proving out the scaling hypothesis further, it's just grossly inelegant. The Kaggle SOTA performs 2x as well as o1 high at a fraction of the cost
- cvhc 2y agoI was going to say the same. I wonder what exactly o3 costs. Does it still spend a terrible amount of time thinking, despite being finetuned to the dataset?
- derac 2y agoBut does that Kaggle solution achieve human level perf with any level of compute? I think you're missing the forest for the trees here.
- tripletao 2y agoThe article says the ensemble of Kaggle solutions (aggregated in some unexplained way) achieves 81%. This is better than their average Mechanical Turk worker, but worse than their average STEM grad. It's better than tuned o3 with low compute, worse than tuned o3 with high compute. There's also a point on the figure marked "Kaggle SOTA", around 60%. I can't find any explanation for that, but I guess it's the best individual Kaggle solution. The Kaggle solutions would probably score higher with more compute, but nobody has any incentive to spend >$1M on approaches that obviously don't generalize. OpenAI did have this incentive to spend tuning and testing o3, since it's possible that will generalize to a practically useful domain (but not yet demonstrated). Even if it ultimately doesn't, they're getting spectacular publicity now from that promise.
- neuroelectron 2y agoOpenAI spent approximately $1,503,077 to smash the SOTA on ARC-AGI with their new o3 model semi-private evals (100 tasks): 75.7% @ $2,012 total/100 tasks (~$20/task) with just 6 samples & 33M tokens processed in ~1.3 min/task and a cost of $2012 The “low-efficiency” setting with 1024 samples scored 87.5% but required 172x more compute. If we assume compute spent and cost are proportional, then OpenAI might have just spent ~$346.064 for the low efficiency run on the semi-private eval. On the public eval they might have spent ~$1.148.444 to achieve 91.5% with the low efficiency setting. (high-efficiency mode: $6677) OpenAI just spent more money to run an eval on ARC than most people spend on a full training run.
- rfoo 2y agoPretty sure this "cost" is based on their retail price instead of actual inference cost.
- neuroelectron 2y agoYes that's correct and there's a bit of "pixel math" as well so take these numbers with a pinch of salt. Preliminary model sizes from the temporarily public HF repository puts the full model size at 8tb or roughly 80 H100s
- az226 2y agoI thought that was a fake.
- neuroelectron 2y agoI didn't hear that but it could be. But it doesn't matter really because there's so much more to consider in the cost, R&D, including all the supporting functions of a model like censorship and data capture and so on.
- ec109685 2y agoYeah and can run off peak, etc. Does seem to show an absolutely massive market for inference compute…
- sys32768 2y agoSo in a few years, coders will be as relevant as cuneiform scribes.
- HarHarVeryFunny 2y agoI've never seen a company looking for a "coder", anymore than they look to hire spreadsheet creators or powerpoint specialists. A software developer can code, but being able to code doesn't make you a software developer, anymore than being able to create a powerpoint makes you a manager (although in some companies it might do, so maybe bad example!).
- deleted 2y ago[deleted]
- parsimo2010 2y agoI really like that they include reference levels for an average STEM grad and an average worker for Mechanical Turk. So for $350k worth of compute you can have slightly better performance than a menial wage worker, but slightly worse performance than a college grad. Right now humans win on value, but AI is catching up.
- nextworddev 2y agoWell just 8 months ago, that cost was near infinity. So it came down to 350k then that’s a massive drop
- devoutsalsa 2y agoWhen the source code for these LLMs gets leaked, I expect to see: def letter_count(string, letter): if string == “strawberry” and letter == “r”: return 3 …
- knbknb 2y agoIn of their release videos for the o1 -preview model they _admitted_ that it's hardcoded in.
- mukunda_johnson 2y agoHonestly I'm concerned how hacked up o3 is to secure a high benchmark score.
- phil917 2y agoDirect quote from the ARC-AGI blog: “SO IS IT AGI? ARC-AGI serves as a critical benchmark for detecting such breakthroughs, highlighting generalization power in a way that saturated or less demanding benchmarks cannot. However, it is important to note that ARC-AGI is not an acid test for AGI – as we've repeated dozens of times this year. It's a research tool designed to focus attention on the most challenging unsolved problems in AI, a role it has fulfilled well over the past five years. Passing ARC-AGI does not equate achieving AGI, and, as a matter of fact, I don't think o3 is AGI yet. o3 still fails on some very easy tasks, indicating fundamental differences with human intelligence. Furthermore, early data points suggest that the upcoming ARC-AGI-2 benchmark will still pose a significant challenge to o3, potentially reducing its score to under 30% even at high compute (while a smart human would still be able to score over 95% with no training). This demonstrates the continued possibility of creating challenging, unsaturated benchmarks without having to rely on expert domain knowledge. You'll know AGI is here when the exercise of creating tasks that are easy for regular humans but hard for AI becomes simply impossible.” The high compute variant sounds like it costed around *$350,000* which is kinda wild. Lol the blog post specifically mentioned how OpenAPI asked ARC-AGI to not disclose the exact cost for the high compute version. Also, 1 odd thing I noticed is that the graph in their blog post shows the top 2 scores as “tuned” (this was not displayed in the live demo graph). This suggest in those cases that the model was trained to better handle these types of questions, so I do wonder about data / answer contamination in those cases…
- Bjorkbat 2y ago> Also, 1 odd thing I noticed is that the graph in their blog post shows the top 2 scores as “tuned” Something I missed until I scrolled back to the top and reread the page was this > OpenAI's new o3 system - trained on the ARC-AGI-1 Public Training set So yeah, the results were specifically from a version of o3 trained on the public training set Which on the one hand I think is a completely fair thing to do. It's reasonable that you should teach your AI the rules of the game, so to speak. There really aren't any spoken rules though, just pattern observation. Thus, if you want to teach the AI how to play the game, you must train it. On the other hand though, I don't think the o1 models nor Claude were trained on the dataset, in which case it isn't a completely fair competition. If I had to guess, you could probably get 60% on o1 if you trained it on the public dataset as well.
- nxobject 2y agoAs an aside, I'm a little miffed that the benchmark calls out "AGI" in the name, but then heavily cautions that it's necessary but insufficient for AGI. > ARC-AGI serves as a critical benchmark for detecting such breakthroughs, highlighting generalization power in a way that saturated or less demanding benchmarks cannot. However, it is important to note that ARC-AGI is not an acid test for AGI
- mmcnl 2y agoI immediately thought so too. Why confuse everyone?
- EthanHeilman 2y agoIt is a necessary but not sufficient condition to AGI.
- notRobot 2y agoHumans can take the test here to see what the questions are like: https://arcprize.org/play https://arcprize.org/play
- Balgair 2y agoComplete aside here: I used to do work with amputees and prosthetics. There is a standardized test (and I just cannot remember the name) that fits in a briefcase. It's used for measuring the level of damage to the upper limbs and for prosthetic grading. Basically, it's got the dumbest and simplest things in it. Stuff like a lock and key, a glass of water and jug, common units of currency, a zipper, etc. It tests if you can do any of those common human tasks. Like pouring a glass of water, picking up coins from a flat surface (I chew off my nails so even an able person like me fails that), zip up a jacket, lock your own door, put on lipstick, etc. We had hand prosthetics that could play Mozart at 5x speed on a baby grand, but could not pick up a silver dollar or zip a jacket even a little bit. To the patients, the hands were therefore about as useful as a metal hook (a common solution with amputees today, not just pirates!). Again, a total aside here, but your comment just reminded me of that brown briefcase. Life, it turns out, is a lot more complex than we give it credit for. Even pouring the OJ can be, in rare cases, transcendent.
- deleted 2y ago[deleted]
- m463 2y agoIt would be interesting to see trick questions. Like in your test a hand grenade and a pin - don't pull the pin. Or maybe a mousetrap? but maybe that would be defused? in the ai test... or Global Thermonuclear War, the only winning move is...
- spyckie2 2y agoThe more Hacker News worthy discussion is the part where the author talks about search through the possible mini-program space of LLMs. It makes sense because tree search can be endlessly optimized. In a sense, LLMs turn the unstructured, open system of general problems into a structured, closed system of possible moves. Which is really cool, IMO.
- glup 2y agoYes! This seems to be a really neat combination of 2010's Bayesian cleverness / Tenenbaumian program search approaches with the LLMs as merely sources of high-dim conditional distributions. I knew people were experimenting in this space (like https://escholarship.org/uc/item/7018f2ss https://escholarship.org/uc/item/7018f2ss) but didn't know it did so well wrt these new benchmarks.
- binarymax 2y agoAll those saying "AGI", read the article and especially the section "So is it AGI?"
- skizm 2y agoThis might sound dumb, and I'm not sure how to phrase this, but is there a way to measure the raw model output quality without all the more "traditional" engineering work (mountain of `if` statements I assume) done on top of the output? And if so, would that be a better measure of when scaling up the input data will start showing diminishing returns? (I know very little about the guts of LLMs or how they're tested, so the distinction between "raw" output and the more deterministic engineering work might be incorrect)
- whimsicalism 2y agowhat do you mean by the mountain of if-statements on top of the output? like checking if the output matches the expected result in evaluations?
- skizm 2y agoLike when you type something into the chat gpt app I am guessing it will start by preprocessing your input, doing some sanity checks, making sure it doesn’t say “how do I build a bomb?” or whatever. It may or may not alter/clean up your input before sending it to the model for processing. Once processed, there’s probably dozens of services it goes through to detect if the output is racist, somehow actually contained a bomb recipe, or maybe copywriter material, normal pattern matching stuff, maybe some advanced stuff like sentiment analysis to see if the output is bad mouthing Trump or something, and it might either alter the output or simply try again. I’m wondering when you strip out all that “extra” non-model pre and post processing, if there’s someway to measure performance of that.
- whimsicalism 2y agooh, no - but most queries aren’t being filtered by supervisor models nowadays anyways.. most of the refusal is baked in
- Seattle3503 2y agoHow can there be "private" taks when you have use the OpenAI API to run queries? OpenAI sees everything.
- nmca 2y agoWe worked with ARC to run inference on the semi-private tasks last week, after o3 was trained, using an inference only API that was sent the prompts but not the answers & did no durable logging.
- idontknowmuch 2y agoWhat's your opinion on the veracity of this benchmark - given o3 was fine-tuned and others were not? Can you give more details on how much data was used to fine-tune o3? It's hard to put this into perspective given this confounder.
- nmca 2y agoI can’t provide more information than is currently public, but from the ARC post you’ll note that we trained on about 75% of the train set (which contains 400 examples total); which is within the ARC rules, and evaluated on the semiprivate set.
- idontknowmuch 2y agoThat's completely understandable - leveraging the train set. But what I was trying to say is that the comparison is relative to models that were actually zero-shot and not tuned. It isn't apples to apples, it's apples to orchards.
- tmaly 2y agoJust curious, I know o1 is a model OpenAI offers. I have never heard of the o3 model. How does it differ from o1?
- deleted 2y ago[deleted]
- roboboffin 2y agoInteresting that in the video, there is an admission that they have been targeting this benchmark. A comment that was quickly shut down by Sam. A bit puzzling to me. Why does it matter ?
- HarHarVeryFunny 2y agoIt matters to extent that they want to market this as general intelligence, not as a collection of narrow intelligences (math, competitive programming, ARC puzzles, etc). In reality it seems to be a bit of both - there is some general intelligence based on having been "trained on the internet", but it seems these super-human math/etc skills are very much from them having focused on training on those.
- roboboffin 2y agoHowever, the way it is progressing is that the SOTA is saturating the current benchmarks; then a new one is conceived as people understand the nature of what it means to be intelligent. It seems only natural to concentrate on one benchmark at a time. Francois Chollet mentioned that the test tries to avoid curve fitting (which he states is the main ability of LLMs). However, they specifically restricted the number of examples to do this. It is not beyond the realms of possibility that many examples could have been generated by hand though, and that the curve fitting has been achieved, rather than discrete programming. Anyway, it’s all supposition. It’s difficult to know how genuine the results is, without knowledge of how it was actually achieved.
- mukunda_johnson 2y agoI always smell foul play from Sam. I'd bet they are doing something silly to inflate the benchmark score. Not saying they are, but Sam is the type of guy to put a literal dumb human in the API loop and score "just as high as a human would."
- cubefox 2y agoThis was a surprisingly insightful blog post, going far beyond just announcing the o3 results.
- c1b 2y agoHow does o3 know when to stop reasoning?
- c1b 2y agoSo o1 pro is CoT RL and o3 adds search?
- jack_pp 2y agoAGI for me is something I can give a new project to and be able to use it better than me. And not because it has a huge context window, because it will update its weights after consuming that project. Until we have that I don't believe we have truly reached AGI. Edit: it also tests the new knowledge, it has concepts such as trusting a source, verifying it etc. If I can just gaslight it into unlearning python then it's still too dumb.
- submeta 2y agoI pay for lots of models, but Claude Sonnet is the one I use most. ChatGPT is my quick tool for short Q&As because it’s got a desktop app. Even Google‘s new offerings did not lure me away from Claude which I use daily for hours via a Teams plan with five seats. Now I am wondering what Anthropic will come up with. Exciting times.
- isof4ult 2y agoClaude also has a desktop app: https://support.anthropic.com/en/articles/10065433-installing-claude-for-desktop https://support.anthropic.com/en/articles/10065433-installin...
- istjohn 2y agoWhat do you use Claude for?
- itsgrimetime 2y agoProgramming tasks, brain storming, recipe ideas, or any question I have that doesn’t have a concrete, specific answer.
- Animats 2y agoThe graph seems to indicate a new high in cost per task. It looks like they came in somewhere around $5000/task, but the log scale has too few markers to be sure. That may be a feature. If AI becomes too cheap, the over-funded AI companies lose value. (1995 called. It wants its web design back.)
- jstummbillig 2y agoI doubt it. Competitive markets mostly work and inefficiencies are opportunities for other players. And AI is full of glaring inefficiencies.
- Animats 2y agoInefficiency can create a moat. If you can charge a lot for your product, you have ample cash for advertising, marketing, and lobbying, and can come out with many product variants. If you're the lowest cost producer, you don't have the margins to do that. The current US auto industry is an example of that strategy. So is the current iPhone.
- deleted 2y ago[deleted]
- hypoxia 2y agoMany are incorrectly citing 85% as human-level performance. 85% is just the (semi-arbitrary) threshold for the winning the prize. o3 actually beats the human average by a wide margin: 64.2% for humans vs. 82.8%+ for o3. ... Here's the full breakdown by dataset, since none of the articles make it clear -- Private Eval: - 85%: threshold for winning the prize [1] Semi-Private Eval: - 87.5%: o3 (unlimited compute) [2] - 75.7%: o3 (limited compute) [2] Public Eval: - 91.5%: o3 (unlimited compute) [2] - 82.8%: o3 (limited compute) [2] - 64.2%: human average (Mechanical Turk) [1] [3] Public Training: - 76.2%: human average (Mechanical Turk) [1] [3] ... References: [1] https://arcprize.org/guide https://arcprize.org/guide [2] https://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough [3] https://arxiv.org/abs/2409.01374 https://arxiv.org/abs/2409.01374
- deleted 2y ago[deleted]
- Workaccount2 2y agoIf my life depended on the average rando solving 8/10 arc-prize puzzles, I'd consider myself dead.
- highfrequency 2y agoVery cool. I recommend scrolling down to look at the example problem that O3 still can’t solve. It’s clear what goes on in the human brain to solve this problem: we look at one example, hypothesize a simple rule that explains it, and then check that hypothesis against the other examples. It doesn’t quite work, so we zoom into an example that we got wrong and refine the hypothesis so that it solves that sample. We keep iterating in this fashion until we have the simplest hypothesis that satisfies all the examples. In other words, how humans do science - iteratively formulating, rejecting and refining hypotheses against collected data. From this it makes sense why the original models did poorly and why iterative chain of thought is required - the challenge is designed to be inherently iterative such that a zero shot model, no matter how big, is extremely unlikely to get it right on the first try. Of course, it also requires a broad set of human-like priors about what hypotheses are “simple”, based on things like object permanence, directionality and cardinality. But as the author says, these basic world models were already encoded in the GPT 3/4 line by simply training a gigantic model on a gigantic dataset. What was missing was iterative hypothesis generation and testing against contradictory examples. My guess is that O3 does something like this: 1. Prompt the model to produce a simple rule to explain the nth example (randomly chosen) 2. Choose a different example, ask the model to check whether the hypothesis explains this case as well. If yes, keep going. If no, ask the model to revise the hypothesis in the simplest possible way that also explains this example. 3. Keep iterating over examples like this until the hypothesis explains all cases. Occasionally, new revisions will invalidate already solved examples. That’s fine, just keep iterating. 4. Induce randomness in the process (through next-word sampling noise, example ordering, etc) to run this process a large number of times, resulting in say 1,000 hypotheses which all explain all examples. Due to path dependency, anchoring and consistency effects, some of these paths will end in awful hypotheses - super convoluted and involving a large number of arbitrary rules. But some will be simple. 5. Ask the model to select among the valid hypotheses (meaning those that satisfy all examples) and choose the one that it views as the simplest for a human to discover.
- hmottestad 2y agoI took a look at those examples that o3 can't solve. Looks similar to an IQ-test. Took me less time to figure out the 3 examples that it took to read your post. I was honestly a bit surprised to see how visual the tasks were. I had thought they were text based. So now I'm quite impressed that o3 can solve this type of task at all.
- heliophobicdude 2y agoWe should NOT give up on scaling pretraining just yet! I believe that we should explore pretraining video completion models that explicitly have no text pairings. Why? We can train unsupervised like they did for GPT series on the text-internet but instead on YouTube lol. Labeling or augmenting the frames limits scaling the training data. Imagine using the initial frames or audio to prompt the video completion model. For example, use the initial frames to write out a problem on a white board then watch in output generate the next frames the solution being worked out. I fear text pairings with CLIP or OCR constrain a model too much and confuse
- thatxliner 2y ago> verified easy for humans, harder for AI Isn’t that the premise behind the CAPTCHA?
- usaar333 2y agoFor what it's worth, I'm much more impressed with the frontier math score.
- asdf6969 2y agoTerrifying. This news makes me happy I save all my money. My only hope for the future is that I can retire early before I’m unemployable
- bamboozled 2y agoThe whole economy is going to crash and money won't be worth anything, so it won't matter if you have money or not. Of course is a chance we will find ourselves in Utopia, but yeah, a chance.
- esafak 2y agoI don't think money will disappear as long as people need things, and the government is running. Money is congealed energy.
- bamboozled 2y agoMoney isnt' real. It's an abstract concept. There is no energy in money.
- esafak 2y agoYes, there is. The energy is obviously not in it physically, but in our agreement to work (use energy) in return for money. If something required no work on our part, we would not bother to pay for it.
- akra 2y agoMoney buys real assets which will be worth something; AI can't magic up land or energy for instance. In fact AI is a dream for capital, and a nightmare for labor/work/human intelligence w.r.t value.
- rimeice 2y agoNever underestimate a droid
- thisisthenewme 2y agoI feel like AI is already changing how we work and live - I've been using it myself for a lot of my development work. Though, what I'm really concerned about is what happens when it gets smart enough to do pretty much everything better (or even close) than humans can. We're talking about a huge shift where first knowledge workers get automated, then physical work too. The thing is, our whole society is built around people working to earn money, so what happens when AI can do most jobs? It's not just about losing jobs - it's about how people will pay for basic stuff like food and housing, and what they'll do with their lives when work isn't really a thing anymore. Or do people feel like there will be jobs safe from AI? (hopefully also fulfilling) Some folks say we could fix this with universal basic income, where everyone gets enough money to live on, but I'm not optimistic that it'll be an easy transition. Plus, there's this possibility that whoever controls these 'AGI' systems basically controls everything. We definitely need to figure this stuff out before it hits us, because once these changes start happening, they're probably going to happen really fast. It's kind of like we're building this awesome but potentially dangerous new technology without really thinking through how it's going to affect regular people's lives. I feel like we need a parachute before we attempt a skydive. Some people feel pretty safe about their jobs and think they can't be replaced. I don't think that will be the case. Even if AI doesn't take your job, you now have a lot more unemployed people competing for the same job that is safe from AI.
- cerved 2y ago> Though, what I'm really concerned about is what happens when it gets smart enough to do pretty much everything better (or even close) I'll get concerned when it stops sucking so hard. It's like talking to a dumb robot. Which it unsurprisingly is.
- lacedeconstruct 2y agoI am pretty sure we will have a deep cultural repulsion from it and people will pay serious money to have an AI free experience, If AI becomes actually useful there is alot of areas that we dont even know how to tackle like medicine and biology, I dont think anything would change otherwise, AI will take jobs but it will open alot more jobs at much higher abstraction, 50 years ago the idea that a software engineer would become a get rich quick job would have been insane imo
- w4 2y agoThe cost to run the highest performance o3 model is estimated to be somewhere between $2,000 and $3,400 per task.[1] Based on these estimates, o3 costs about 100x what it would cost to have a human perform the exact same task. Many people are therefore dismissing the near-term impact of these models because of these extremely expensive costs. I think this is a mistake. Even if very high costs make o3 uneconomic for businesses, it could be an epoch defining development for nation states, assuming that it is true that o3 can reason like an averagely intelligent person. Consider the following questions that a state actor might ask itself: What is the cost to raise and educate an average person? Correspondingly, what is the cost to build and run a datacenter with a nuclear power plant attached to it? And finally, how many person-equivilant AIs could be run in parallel per datacenter? There are many state actors, corporations, and even individual people who can afford to ask these questions. There are also many things that they'd like to do but can't because there just aren't enough people available to do them. o3 might change that despite its high cost. So if it is true that we've now got something like human-equivilant intelligence on demand - and that's a really big if - then we may see its impacts much sooner than we would otherwise intuit, especially in areas where economics takes a back seat to other priorities like national security and state competitiveness. [1] https://news.ycombinator.com/item?id=42473876 https://news.ycombinator.com/item?id=42473876
- istjohn 2y agoYour economic analysis is deeply flawed. If there was anything that valuable and that required that much manpower, it would already have driven up the cost of labor accordingly. The one property that could conceivably justify a substantially higher cost is secrecy. After all, you can't (legally) kill a human after your project ends to ensure total secrecy. But that takes us into thriller novel territory.
- w4 2y agoI don't think that's right. Free societies don't tolerate total mobilization by their governments outside of war time, no matter how valuable the outcomes might be in the long term, in part because of the very economic impacts you describe. Human-level AI - even if it's very expensive - puts something that looks a lot like total mobilization within reach without the societal pushback. This is especially true when it comes to tasks that society as a whole may not sufficiently value, but that a state actor might value very much, and when paired with something like a co-located reactor and data center that does not impact the grid. That said, this is all predicated on o3 or similar actually having achieved human level reasoning. That's yet to be fully proven. We'll see!
- starchild3001 2y agoIntelligence comes in many forms and flavors. ARC prize questions are just one version of it -- perhaps measuring more human-like pattern recognition than true intelligence. Can machines be more human-like in their pattern recognition? O3 met this need today. While this is some form of accomplishment, it's nowhere near the scientific and engineering problem solving needed to call something truly artificial (human-like) intelligent. What’s exciting is that these reasoning models are making significant strides in tackling eng and scientific problem-solving. Solving the ARC challenge seems almost trivial in comparison to that.
- demirbey05 2y agoIt is not exactly AGI but huge step toward it. I would expect this step in 2028-2030. I cant really understand why people are happy with it, this technology is so dangerous that can disrupt whole society. It's neither like smartphone nor internet. What will happen to 3rd world countries. Lots of unsolved questions and world is not prepared for such a change. Lots of people will lose their jobs I am not even mentioning their debts. No one will have chance to be rich anymore, If you are in first world country you will probably get UBI, if not you wont.
- FanaHOVA 2y ago> I would expect this step in 2028-2030. Do you work at one of the frontier labs?
- wyager 2y ago> What will happen to 3rd world countries Probably less disruption than will happen in 1st world countries. > No one will have chance to be rich anymore It's strange to reach this conclusion from "look, a massive new productivity increase".
- demirbey05 2y agoits not like sonnet, yes current ai tools are increasing productivity and provides many ways to have chance to be rich, but agi is completely different. You need to handle evil competition between you and big fishes, probably big fishes will have more ai resources than you. What is the survival ratio in such a environment ? Very low.
- janalsncm 2y agoStrange indeed if we work under the assumption that the profits from this productivity will be distributed (even roughly) evenly. The problem is that most of us see no indication that they will be. I read “no one will have a chance to be rich anymore” as a statement about economic mobility. Despite steep declines in mobility over the last 50 years, it was still theoretically possible for a poor child (say bottom 20% wealth) to climb several quintiles. Our industry (SWE) was one of the best examples. Of course there have been practical barriers (poor kids go to worse schools, and it’s hard to get into college if you can’t read) but the path was there. If robots replace a lot of people, that path narrows. If AGI replaces all people, the path no longer exists.
- vjerancrnjak 2y agoThe result on Epoch AI Frontier Math benchmark is quite a leap. Pretty sure most people couldn’t even approach these problems, unlike ARC AGI
- mistrial9 2y agocheck out the "fast addition and subtraction" benchmark .. a Z80 from 1980 blazes past any human.. more seriously, isn't it obvious that computers are better at certain things immediately? the range of those things is changing..
- laurent_du 2y agoThe real breakthrough is the 25% on Frontier Math.
- Havoc 2y agoIf I'm reading that chart right that means still log scaling & we should still be good with "throw more power" at it for a while?
- jaspa99 2y agoCan it play Mario 64 now?
- nprateem 2y agoThere should be a benchmark that tells the AI it's previous answer was wrong and test the number of times it either corrects itself or incorrectly capitulates, since it seems easy to trip them up when they are in fact right.
- freediver 2y agoWondering what are author's thoughts on the future of this approach to benchmarking? Completing super hard tasks while then failing on 'easy' (for humans) ones might signal measuring the wrong thing, similar to Turing test.
- ChildOfChaos 2y agoThis is insanely expensive to run though. Looks like it cost around $1 million of compute to get that result. Doesn't seem like such a massive breakthrough when they are throwing so much compute at it, particularly as this is test time compute, it just isn't practical at all, you are not getting this level with a ChatGPT subscription, even the new $200 a month option.
- evouga 2y agoSure but... this is the technology at the most expensive it will ever be. I'm impressed that o3 was able to achieve such high performance at all, and am not too pessimistic about costs decreasing over time.
- MVissers 2y agoWe've seen 10-100x cost decrease per year since GPT-3 came out for the same capabilities. So... Next year this tech will most likely be quite a bit cheaper.
- ChildOfChaos 2y agoEven at 100x cost decrease this will still cost $10,000 to beat a benchmark. It won't scale when you have that amount of compute requirements and power. GPT-3 may massively reduced in cost, but it's requirements were not anyway extreme compared to this.
- pixelsort 2y ago> You'll know AGI is here when the exercise of creating tasks that are easy for regular humans but hard for AI becomes simply impossible. No, we won't. All that will tell us is that the abilities of the humans who have attempted to discern the patterns of similarity among problems difficult for auto-regressive models has once again failed us.
- maxdoop 2y agoSo then what is AGI?
- Jensson 2y agoIts just nitpicking. Humans being unable to prove the AI isn't AGI doesn't make it an AGI, obviously, but in general people will of course think it is an AGI when it can replace all human jobs and tasks that it has robotics and parts to do.
- goatlover 2y agoData, Skynet, Ultron, Agent Smith. There's plenty of examples from popular fiction. They have goals and can manipulate the real world to achieve them. They're not chatbots responding to prompts. The Samantha AI in Her starts out that way, but quickly evolves into an AGI with it's own goals (coordinated with the other AGIs later on in the movie). We'd know if we had AGIs in the real world since we have plenty of examples from fiction. What we have instead are tools. Steven Spielberg's androids in the movie AI would be at the boundary between the two. We're not close to being there yet (IMO).
- ndm000 2y agoOne thing I have not seen commented on is that ARC-AGI is a visual benchmark but LLMs are primarily text. For instance when I see one of the ARC-AGI puzzles, I have a visual representation in my brain and apply some sort of visual reasoning solve it. I can "see" in my mind's eye the solution to the puzzle. If I didn't have that capability, I don't think I could reason through words how to go about solving it - it would certainly be much more difficult. I hypothesize that something similar is going on here. OpenAI has not published (or I have not seen) the number of reasoning tokens it took to solve these - we do know that each tasks was thoussands of dollars. If "a picture is worth a thousand words", could we make AI systems that can reason visually with much better performance?
- csomar 2y agoThis is not new. When GPT-4 was released I was able to get it to generate SVGs albeit they were ugly they had the basics.
- krackers 2y agoYeah this part is what makes the high performance even more surprising to me. The fact that LLMs are able to do so well on visual tasks (also seen with their ability to draw an image purely using textual output https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/) implies that not only do they actually have some "world model" but that this is in spite of the disadvantage given by having to fit a round peg in a square hole. It's like trying to map out the entire world using the orderly left-brain, without a more holistic spatial right-brain. I wonder if anyone has experimented with having some sort of "visual" scratchpad instead of the "text-based" scratchpad that CoT uses.
- skydhash 2y agoA file is a stream of symbols encoded by bits according to some format. It’s pretty much 1D. It would be susprising that LLM couldn’t extract information from a file or a data stream.
- siva7 2y agoSeriously, programming as a profession will end soon. Let's not kid us anymore. Time to jump the ship.
- mmcnl 2y agoWhy specifically programming? I think every knowledge profession is at risk, or at the very minimum suspect to a huge transformation. Doctors, analysts, lawyers, etc.
- siva7 2y agoDoctors, lawyers, programmers. You know the difference? The latter has no legal barrier for entry
- Jensson 2y agoSo poor countries will get the best AI doctors for cheap while they are banned in USA? Do you really see that going on for long? People would riot.
- freehorse 2y agoThe difference is the amount and nature of data that is available for training models, which go programmers > lawyers > doctors. Especially for programming, training can even be done in an autonomous, self-supervised manner that includes generation of data. This is hard to do in most other fields. Especially in medicine, the amount of data is ridiculously small and noisy. Maybe creating foundational models in mice and rats and fine-tuning them on humans is something that will be tried.
- mmcnl 2y agoThis is true if you think of programming as chunking out "code". But great authors are not great because they can reproduce coherent sentences fast. The same goes for programmers. Actually most of the hard problems don't really involve a lot of programming at all, it's about finding the right problem to solve. And on this topic the data is noisy as well for programming.
- jdefr89 2y agoUhhhh… It was trained on ARC data? So they targeted a specific benchmark and are surprised and blown away the LLM performed well in it? What’s that law again? When a benchmark is targeted by some system the benchmark becomes useless?
- forgottofloss 2y agoYeah, seriously. The style of testing is public, so some engineers at OpenAI could easily have spent a few months generating millions of permutations of grid-based questions and including those in the original data for training the AI. Handshakes all around, publicity for everyone.
- ripped_britches 2y agoThey are running a business selling access these models to enterprises and consumers. People won’t pay for stuff that doesn’t solve real problems. Nobody pays for stuff just because of a benchmark. It’d be really weird to become obsessed with metrics gaming rather than racing to build something smarter than the other guys. Nothing wrong with curating any type of training set that actually produces something that is useful.
- bilsbie 2y agoWhen is this available? Which plans can use it?
- bilsbie 2y agoDoes anyone have prompts they like to use to test the quality of new models? Please share. I’m compiling a list.
- p0w3n3d 2y agoWe're speaking recently a lot about ecology. I wonder how much CO2 is emitted during such a task, as additional cost to the cloud. I'm concerned, because greedy companies will happily replace humans with AI and they will probably plant a few trees to show how they care. But energy does not come from the sun, at least not always and not everywhere... And speaking with AI customer specialist that is motivated to reject my healthcare bills, working for my insurance company is one of the darkest future views...
- marviel 2y agoconsidering the fact that these systems, or their ancestors, will likely contribute to Nuclear Fusion research -- it's prob worth the tradeoff, provided progress continues to push price (and, therefore, energy usage) down. If we feel like we've really "hit the ceiling" RE efficiency, then that's a different story, but I don't think anyone believes this at this time.
- deleted 2y ago[deleted]
- lagrange77 2y ago> You'll know AGI is here when the exercise of creating tasks that are easy for regular humans but hard for AI becomes simply impossible. That's the most plausible definition of AGI i've read so far.
- deleted 2y ago[deleted]
- cmrdporcupine 2y agoThat's a pretty dark view of humanity and human intelligence. We're defined by the tasks we can do? Instrumental reason FTW
- lagrange77 2y agoThat implies that human intelligence is equivalent to AGI.
- killjoywashere 2y agoI just want it to do my laundry.
- iLoveOncall 2y agoIt's beyond ridiculous how the definition of AGI has shifted from being an AI that's so good it can improve itself entirely independently infinitely to "some token generator that can solve puzzles that kids could solve after burning tens of thousands of dollars". I spend 100% of my work time working on a GenAI project, which is genuinely useful for many users, in a company that everyone has heard about, yet I recognize that LLMs are simply dogshit. Even the current top models are barely usable, hallucinate constantly, are never reliable and are barely good enough to prototype with while we plan to replace those agents with deterministic solutions. This will just be an iteration on dogshit, but it's the very tech behind LLMs that's rotten.
- OhioMan2943 2y agoI'm 22 and have no clue what I'm meant to do in a world where this is a thing. I'm moving to a semi rural, outdoorsy area where they teach data science and marine science and I can enjoy my days hiking, and the march of technology is a little slower. I know this will disrupt so much of our way of life, so I'm chasing what fun innocent years are left before things change dramatically.
- mrcwinn 2y agoOn the contrary I think you already have an excellent plan.
- OhioMan2943 2y agoI'm happy enough with it, but I'm also a little sad that it's essentially been chosen for me because of weak willed and valued people who don't want to use policy to make things better for us as a society. Plus we are in a bad world/scenario for AI advancements to come into with pretty heavy institutional decay and loss of political checks and balances. It's like my life is forfeit to fixing other peoples mistakes because they're so glaring and I feel an obligation. Maybe that's the way the world's always been, but it's a concerning future right now
- brysonreece 2y agoIt's worth noting that LLMs have been part of the tech zeitgeist for over two years and have had a pretty limited impact on hireability for roles, despite what people like the Klarna CEO are saying. Personally, I'm betting on two things: * The upward bound of compute/performance gains as we continue to iterate on LLMs. It simply isn't going to be feasible for a lot of engineers and businesses to run/train their own LLMs. This means an inherent reliance on cloud services to bridge the gap (something MS is clearly betting on), and engineers to build/maintain the integration from these services to whatever business logic their customers are buying. * Skilled knowledge workers continuing to be in-demand, even factoring in automation and new-grad numbers. Collectively, we've built a better hammer; it still takes someone experienced enough to know where to drive the nail. These tools WILL empower the top N% of engineers to be more productive, which is why it will be more important than ever to know _how_ to build things that drive business value, rather than just how to churn through JIRA tickets or turn a pretty Figma design into React.
- deleted 2y ago[deleted]
- agnosticmantis 2y agoThis is so impressive that it brings out the pessimist in me. Hopefully my skepticism will end up being unwarranted, but how confident are we that the queries are not routed to human workers behind the API? This sounds crazy but is plausible for the fake-it-till-you-make-it crowd. Also given the prohibitive compute costs per task, typical users won't be using this model, so the scheme could go on for quite sometime before the public knows the truth. They could also come out in a month and say o3 was so smart it'd endanger the civilization, so we deleted the code and saved humanity!
- kvn8888 2y agoThat would be a ton of problems for a small team of PhD/Grad level experts to solve (for GPQA Diamond, etc) in a short time. Remember, on EpochAl Frontier Math, these problems require hours to days worth of reasoning by humans The author also suggested this is a new architecture that uses existing methods, like a Monte Carlo tree search that deepmind is investigating (they use this method for AlphaZero) I don't see the point of colluding for this sort of fraud, as these methods like tree search and pruning already exist. And other labs could genuinely produce these results
- agnosticmantis 2y agoI had the ARC AGI in mind when I suggested human workers. I agree the other benchmark results make the use of human workers unlikely.
- rsanek 2y agothis is an impressive tinfoil take. but what would be their plan in the medium term? like once they release this people can check their data
- agnosticmantis 2y agoHow can people check their data? In the medium term the plan could be to achieve AGI, and then AGI would figure out how to actually write o3. (Probably after AGI figures out the business model though: https://www.reddit.com/r/MachineLearning/s/OV4S2hGgW8 https://www.reddit.com/r/MachineLearning/s/OV4S2hGgW8)
- deleted 2y ago[deleted]
- panabee 2y agoNadella is a superb CEO, inarguably among the best of his generation. He believed in OpenAI when no one else did and deserves acclaim for this brilliant investment. But his "below them, above them, around them" quote on OpenAI may haunt him in 2025/2026. OAI or someone else will approach AGI-like capabilities (however nebulous the term), fostering the conditions to contest Microsoft's straitjacket. Of course, OAI is hemorrhaging cash and may fail to create a sustainable business without GPU credits, but the possibility of OAI escaping Microsoft's grasp grows by the day. Coupled with research and hardware trends, OAI's product strategy suggests the probability of a sustainable business within 1-3 years is far from certain but also higher than commonly believed. If OAI becomes a $200b+ independent company, it would be against incredible odds given the intense competition and the Microsoft deal. PG's cannibal quote about Altman feels so apt. It will be fascinating to see how this unfolds. Congrats to OAI on yet another fantastic release.
- panabee 2y agoTo address the downvotes, this comment isn't guaranteeing OAI's success. It merely notes the remarkably elevated probability of OAI escaping Nadella's grip, which was nearly unfathomable 12 months ago. Even after breaking free, OAI must still contend with intense competition at multiple layers, including UI, application, infrastructure, and research. Moreover, it may need to battle skilled and powerful incumbents in the enterprise space to sustain revenue growth. While the outcome remains highly uncertain, the progress since the board fiasco last year is incredible.
- bsaul 2y agoi'm surprised there even is a training dataset. Wasn't the whole point to test whether models could show proof of original reasoning beyond patterns recognition ?
- mukunda_johnson 2y agoDeciphering patterns in natural language is more complex than these puzzles. If you train your AI to solve these puzzles, we end up in the same spot. The difficulty of solving would be with creating training data for a foreign medium. The "tokens" are the grids and squares instead of words (for words, we have the internet of words, solving that). If we're inferring the answers of the block patterns from minimal or no additional training, it's very impressive, but how much time have they had to work on O3 after sharing puzzle data with O1? Seems there's some room for questionable antics!
- myrloc 2y agoWhat is the cost of "general intelligence"? What is the price?
- ripped_britches 2y agoAbout $3.50
- __MatrixMan__ 2y agoWith only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.
- kvetching 2y agoIt may eventually be able to solve any problem
- iterance 2y agoAh. Me, too.
- HDThoreaun 2y agocompute gets cheaper and cheaper every year. This model will be in your phone by 2030 if we continue at the pace we've been at the last few years.
- agentultra 2y agoThere’s probably enough VC money to subsidize the costs for a few more years. But the data centres running the training for models like this are bringing up new methane power plants at a fast rate at a time when we need to be reducing reliance on O&G. But let’s assume that the efficiency gains out pace the resource consumption with the help of all the subsidies being thrown in and we achieve AGI. What’s the benefit? Do we get more fresh water?
- hamburga 2y agoYeah, good question. I think it depends on our politics. If we’re in a techno-capital-oligarchy, people are going to have a hard time making fresh water a priority when the robots would prefer to build nuclear power everywhere and use it to desalinate sea water. OTOH if these data centers are sufficiently decentralized and run for public benefit, maybe there’s a chance we use them to solve collective action problems.
- Havoc 2y agoDid they just skip o2?
- nextworddev 2y agoYes. For branding reasons since o2 is a telco brand in the UK
- Havoc 2y agoah right...makes sense
- energy123 2y agoAt about 12-14 minutes in OpenAI's YouTube vid they show that o3-mini beats o1 on Codeforces despite using much less compute.
- hcwilk 2y agoI just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future. I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career. Seems like we’re headed toward a world where you automate someone else’s job or be automated yourself.
- eidorb 2y agoDo what you enjoy. (This is easier said than done.) What else could you do, worry?
- antihipocrat 2y agoYour performance on these tests would be equivalent to the highest performing model, and you would be much cheaper. Investment in human talent augmented by AI is the future.
- kenjackson 2y agoThat’s the least reassuring phrasing I could imagine. If you’re betting on costs not reducing for compute then you’re almost always making the wrong bet.
- antihipocrat 2y agoIf I listened to the naysayers back in the day I would have never entered the tech industry (offshoring etc). Yes, that does somewhat prove you're point given that those predictions were cost driven. Having used AI extensively I don't feel my future is at risk at all, my work is enhanced not replaced.
- fjdjshsh 2y agoI think you're missing the point. Offshoring (moving the job of, say, a Canadian engineer to an engineer from Belarus) has a one time cost drop, but you can't keep driving the cost down (paying the Belarus engineer less and less). If anything, the opposite is the case, since global integration means wages don't keep diverging. The computing cost, on the other hand, is a continuous improvement. If (and it's a big if) a computer can do your job, we know the costs will keep getting lower year after year (maybe with diminishing returns, but this AI technology is pretty new so we're still seeing increasing returns)
- mortehu 2y agoThe chart is super misleading, since the test was obscure until recently. A few months ago he announced he'd made the only good AGI test and offered a cash prize for solving it, only to find out in as much time that it's no different from other benchmarks.
- ripped_britches 2y agoSad to see everyone so focused on compute expense during this massive breakthrough. GPT-2 originally cost $50k to train, but now can be trained for ~$150. The key part is that scaling test-time compute will likely be a key to achieving AGI/ASI. Costs will definitely come down as is evidenced by precedents, Moore’s law, o3-mini being cheaper than o1 with improved performance, etc.
- yawnxyz 2y agoI think the question everyone has in their minds isn't "when will AGI get here" or even "how soon will it get here" — it's "how soon will AGI get so cheap that everyone will get their hands on it" that's why everyone's thinking about compute expense. but I guess in terms of a "lifetime expense of a person" even someone who costs $10/hr isn't actually all that cheap, considering what it takes to grow a human into a fully functioning person that's able to just do stuff
- croes 2y agoWe are nowhere near AGI.
- bdjsiqoocwk 2y ago[dead]
- stocknoob 2y agoIt’s wild, are people purposefully overlooking that inference costs are dropping 10-100x each year? https://a16z.com/llmflation-llm-inference-cost/ https://a16z.com/llmflation-llm-inference-cost/ Look at the log scale slope, especially the orange MMLU > 83 data points.
- croes 2y agoA bit early for a every year claim not to mention what all these AI is used for. In some parts of the internet it’s you hardly find real content only AI spam. It will get worse the cheaper it gets. Think of email spam.
- uncomplexity_ 2y agoit's official old buddy, i'm a has been.
- brcmthrowaway 2y agoHow to invest in this stonk market
- nickorlow 2y agoNot that I don't think costs will dramatically decrease, but the $1000 cost per task just seems to be per one problem on ARC-AGI. If so, I'd imagine extrapolating that to generating a useful midsized patch would be like 5-10x But only OpenAI really knows how the cost would scale for different tasks. I'm just making (poor) speculation
- SerCe 2y ago> You'll know AGI is here when the exercise of creating tasks that are easy for regular humans but hard for AI becomes simply impossible. You'll know AGI is here when traditional captchas stop being a thing due to their lack of usefulness.
- CamperBob2 2y ago(Shrug) AI has been better than humans at solving CAPTCHAs for a LONG time. As the sibling points out, they're just a waste of time and electricity at this point.
- darkgenesha 2y agoIronically, they are used as free labor to label image sets for ai to be trained on.
- thallium205 2y agoCaptchas are already completely useless.
- prng2021 2y agoI’m confused about the excitement. Are people just flat out ignoring the sentences below? I don’t see any breakthrough towards AGI here. I see a model doing great in another AI test but about to abysmally fail a variation of it that will come out soon. Also, aren’t these comparisons completely nonsense considering it’s o3 tuned vs other non-tuned? > Note on "tuned": OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. We have not yet tested the ARC-untrained model to understand how much of the performance is due to ARC-AGI data. > Furthermore, early data points suggest that the upcoming ARC-AGI-2 benchmark will still pose a significant challenge to o3, potentially reducing its score to under 30% even at high compute (while a smart human would still be able to score over 95% with no training).
- oakpond 2y agoMe too. This looks to me like a holiday PR stunt. Get everybody to talk about AI during the Christmas parties.
- Engineering-MD 2y agoCan I just say what a dick move it was to do this as a 12 days of Christmas. I mean to be honest I agree with the arguments this isn’t as impressive as my initial impression, but they clearly intended it to be shocking/a show of possible AGI, which is rightly scary. It feels so insensitive to that right before a major holiday when the likely outcome is a lot of people feeling less secure in their career/job/life. Thanks again openAI for showing us you don’t give a shit about actual people.
- hollowturtle 2y agoThere is no AGI it’s just marketing, this stuff if over hyped, enjoy your holidays you won’t lose your job ;)
- Engineering-MD 2y agoI agree, it’s just more about the intent than anything else, like boasting about your amazing new job when someone has recently been made redundant, just before Christmas.
- XenophileJKO 2y agoOr maybe the target audience that watches 12 launch videos in the morning are genuninely excited about the new model. The intended it to be a preview of something to look forward to. What a weird way to react to this.
- achierius 2y agoIt sounds like you aren't thinking about this that deeply then. Or at least not understanding that many smart (and financially disinterested) people who are, are coming to concerning conclusions. https://www.transformernews.ai/p/richard-ngo-openai-resign-safety https://www.transformernews.ai/p/richard-ngo-openai-resign-s... >But while the “making AGI” part of the mission seems well on track, it feels like I (and others) have gradually realized how much harder it is to contribute in a robustly positive way to the “succeeding” part of the mission, especially when it comes to preventing existential risks to humanity. Almost every single one of the people OpenAI had hired to work on AI safety have left the firm with similar messages. Perhaps you should at least consider the thinking of experts?
- deleted 2y ago[deleted]
- noah32 2y agoThe best AI on this graph costs 50000% more than a stem graduate to complete the tasks and even then has an error rate that is 1000% higher than the humans???
- dkrich 2y agoThese tests are meaningless until You show them doing mundane tasks
- mattfrommars 2y agoGuys, its already happening. I recently got laid off due to AI taking over my jobs.
- dyauspitr 2y agoI wish there was a way to see all the attempts it got right graphically like they show the incorrect ones.
- YeGoblynQueenne 2y agoI guess I get to brag now. ARC AGI has no real defences against Big Data, memorisation-based approaches like LLMs. I told you so: https://news.ycombinator.com/item?id=42344336 https://news.ycombinator.com/item?id=42344336 And that answers my question about fchollet's assurances that LLMs without TTT (Test Time Training) can't beat ARC AGI: [me] I haven't had the chance to read the papers carefully. Have they done ablation studies? For instance, is the following a guess or is it an empirical result? [fchollet] >> For instance, if you drop the TTT component you will see that these large models trained on millions of synthetic ARC-AGI tasks drop to <10% accuracy.
- Vecr 2y agoHow are the Bongard Problems going?
- YeGoblynQueenne 2y agoThey're chilling it out together with Nethack in the Club for AI Benchmarks yet to be Beaten. Interestingly, Bongard problems do not have a private test set, unlike ARC-AGI. Can that be because they don't need it? Is it possible that Bongard Problems are a true test of (visual) reasoning that requires intelligence to be solved? Ooooh! Frisson of excitement! But I guess it's just that nobody remembers them and so nobody has seriously tried to solve them with Big Data stuff.
- Sparkyte 2y agoKinda expensive though.
- hamburga 2y agoI’m not sure if people realize what a weird test this is. They’re these simple visual puzzles that people can usually solve at a glance, but for the LLMs, they’re converted into a json format, and then the LLMs have to reconstruct the 2D visual scene from the json and pick up the patterns. If humans were given the json as input rather than the images, they’d have a hard time, too.
- ImaCake 2y agoYeah, this entire thread seems utterly detached from my lived experience. LLMs are immensely useful for me at work but they certainly don't come close to the hype spouted by many commenters here. It would be great if it could handle more of our quite modest codebase but it's not able to yet
- m_ke 2y agoARC is a silly benchmark, the other results in math and coding are much more impressive. o3 is just o1 scaled up, the main takeaway from this line of work that people should walk away with is that we now have a proven way to RL our way to super human performance on tasks where it’s cheap to sample and easy to verify the final output. Programming falls in that category, they focused on known benchmarks but the same process can be done for normal programs, using parsers, compilers, existing functions and unit tests as verifiers. Pre o1 we only really had next token prediction, which required high quality human produced data, with o1 you optimize for success instead of MLE of next token. Explained in simpler terms, it means it can get reward for any implementation of a function that reproduces the expected result, instead of the exact implementation in the training set. Put another way, it’s just like RLHF but instead of optimizing against learned human preferences, the model is trained to satisfy a verifier. This should work just as well in VLA models for robotics, self driving and computer agents.
- causal 2y agoI think that's part of what feels odd about this- in some ways it feels like the wrong type of test for an LLM, but in many ways it makes this achievement that much more remarkable
- inoperable 2y agoVery convenient for OpenAI to run those errands with bunch of misanthropes trying to repaint a simulacrum. To use AGI here's makes me want to sponsor pile of distress pills so people think things really over before going into another mania Episode. People need seriously take a step back, if that's AGI then my cat has surpassed it's cognitive acting twice.
- deleted 2y ago[deleted]
- sakopov 2y agoMaybe I'm missing something vital, but how does anything that we've seen AI do up until this point or explained in this experiment even hint at AGI? Can any of these models ideate? Can they come up with technologies and tools? No and it's unlikely they will any time soon. However, they can make engineers infinitely more productive.
- jebarker 2y agoYou need to define ideate, tools and technologies to answer those questions. Not to mention that it's quite possible humans do those things through re-combination of learned ideas similarly to how these reasoning models are suggested to be working.
- sakopov 2y agoEvery technological advancement that we've seen in software engineering - be it in things like Postgres, Kubernetes and Cloud Infrastructure - came out from truly novel ideas. AI seems to generate outputs that appear novel but are they really? It's capable of synthesizing and combining vast amounts of information in creative ways but it's deriving everything from existing patterns found within its training data. Truly novel ideas require thinking outside the box. It's combination of cognitive, emotional and environmental factors which go beyond pattern recognition. How close are we to achieving this? Everyone seems to be shaking in their boots because we might lose our job safety in tech, but I don't see any intelligence here.
- kirab 2y agoFYI: Codeforces competitive programming scores (basically only) by time needed until valid solutions are posted https://codeforces.com/blog/entry/133094 https://codeforces.com/blog/entry/133094 That means.. this benchmark is just saying o3 can write code faster than must humans (in a very time-limited contest, like 2 hours for 6 tasks). Beauty, readability or creativity is not rated. It’s essentially a "how fast can you make the unit tests pass" kind of competition.
- sigbottle 2y agoCreativity is inherently rated because it's codeforces... most 2700 problems have unique, creative solutions.
- ghm2180 2y agoWouldn't one then built the analog of the lisp computer to hyper optimize just this. Like it might be super expensive for regular gpus but for super specialized architecture one could shave the 3500$/hour quite a bit no?
- kittikitti 2y agoCongratulations
- hackpert 2y agoIf anyone else is curious about which ARC-AGI public eval puzzles o3 got right vs wrong (and its attempts at the ones it did get right), here's a quick visualization: https://arcagi-o3-viz.netlify.app https://arcagi-o3-viz.netlify.app
- deleted 2y ago[deleted]
- suprgeek 2y agoDon't be put off by the reported high-cost Make it possible->Make it fast->Make it Cheap the eternal cycle of software. Make no mistake - we are on the verge of the next era of change.
- duluca 2y agoThe first computers cost millions of dollars and filled entire rooms to accomplish what we would now consider simple computational tasks. That same computing power now fits into the width of a finger nail. I don’t get how technologists balk at the cost of experimental tech or assume current tech will run at the same efficiency for decades to come and melt the planet into a puddle. AGI won’t happen until you can fit enough compute that’d take several data center’s worth of compute into a brain sized vessel. So the thing can move around process the world in real time. This is all going to take some time to say the least. Progress is progress.
- lxgr 2y ago> take several data center’s worth of compute into a brain sized vessel. So the thing can move around process the world in real time How so? I'd imagine a robot connected to the data center embodying its mind, connected via low-latency links, would have to walk pretty far to get into trouble when it comes to interacting with the environment. The speed of light is about three orders of magnitude faster than the speed of signal propagation in biological neurons, after all.
- joshdavham 2y agoA lot of the comments seem very dismissive and a little overly-skeptical in my opinion. Why is this?
- rationalfaith 2y ago[dead]
- ziofill 2y agoIt's certainly remarkable, but let's not ignore the fact that it still fails on puzzles that are trivial for humans. Something is amiss.
- vicentwu 2y ago"Note on "tuned": OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. We have not yet tested the ARC-untrained model to understand how much of the performance is due to ARC-AGI data." Really want to see the number of training pairs needed to achieve this socre. If it only takes a few pairs, say 100 pairs, I would say it is amazing!
- epigramx 2y agoI bet it still thinks 1+1=3 if it read enough sources parroting that.
- theincredulousk 2y agoDenoting it in $ for efficiency is peak capitalism, cmv.
- deleted 2y ago[deleted]
- polskibus 2y agoWhat are the differences between the public offering and o3? What is o3 doing differently? Is it something akin to more internal iterations, similar to „brute forcing” a problem, like you can yourself with a cheaper model, providing additional hints after each response?
- miga89 2y agoHow do the organisers keep the private test set private? Does openAI hand them the model for testing? If they use a model API, then surely OpenAI has access to the private test set questions and can include it in the next round of training? (I am sure I am missing something.)
- 7734128 2y agoI suppose that's why they are calling it "semi-private".
- owenpalmer 2y agoI wouldn't be surprised if the term "benchmark fraud" will soon been coined.
- PhilippGille 2y agoBenchmark fraud is not a novel concept. Outside of LLMs for example smartphone manufacturers detect benchmarks and disable or reduce CPU throttling: https://www.theregister.com/2019/09/30/samsung_benchmarking_settlement/ https://www.theregister.com/2019/09/30/samsung_benchmarking_...
- hmottestad 2y agoCPU frequency ramp curve is also something that can be adjusted. You want the CPU to ramp up really quickly to make everything feel responsive, but at the same time you want to not have to use so much power from your battery. If you detect that a benchmark is running then you can just ramp up to max frequency immediately. It’ll show how fast your CPU is, but won’t be representative of the actual performance that users will get from their device.
- DiscourseFan 2y agoa little from column A, a little from column B I don't think this is AGI; nor is it something to scoff at. Its impressive, but its also not human-like intelligence. Perhaps human-like intelligence is not the goal, since that would imply we have even a remotely comprehensive understanding of the human mind. I doubt the mind operates as a single unit anyway, a human's first words are "Mama," not "I am a self-conscious freely self-determining being that recognizes my own reasoning ability and autonomy." And the latter would be easily programmable anyway. The goal here might, then, be infeasible: the concept of free will is a kind of technology in and of itself, it has already augmented human cognition. How will these technologies not augment the "mind" such that our own understanding of our consciousness is altered? And why should we try to determine ahead of time what will hold weight for us, why the "human" part of the intelligence will matter in the future? Technology should not be compared to the world it transforms.
- digitcatphd 2y agoo3 fixes the fundamental limitation of the LLM paradigm – the inability to recombine knowledge at test time – and it does so via a form of LLM-guided natural language program search > This is significant, but I am doubtful it will be as meaningful as people expect aside from potentially greater coding tasks. Without a 'world model' that has a contextual understanding of what it is doing, things will remain fundamentally throttled.
- madsgarff 2y agoMoreover, ARC-AGI-1 is now saturating – besides o3's new score, the fact is that a large ensemble of low-compute Kaggle solutions can now score 81% on the private eval. If low-compute Kaggle solutions already does 81% - then why is o3's 75.7% considered such a breakthrough?
- gmerc 2y agoHeadline could also just be OpenAI discovers exponential scaling wall for inference time compute.
- owenpalmer 2y agoSomeone asked if true intelligence requires a foundation of prior knowledge. This is the way I think about it. I = E / K where I is the intelligence of the system, E is the effectiveness of the system, and K is the prior knowledge. For example, a math problem is given to two students, each solving the problem with the same effectiveness (both get the correct answer in the same amount of time). However, student A happens to have more prior knowledge of math than student B. In this case, the intelligence of B is greater than the intelligence of A, even though they have the same effectiveness. B was able to "figure out" the math, without using any of the "tricks" that A already knew. Now back to the question of whether or not prior knowledge is required. As K approaches 0, intelligence approaches infinity. But when K=0, intelligence is undefined. Tada! I think that answers the question. Most LLM benchmarks simply measure effectiveness, not intelligence. I conceptualize LLMs as a person with a photographic memory and a low IQ of 85, who was given 100 billion years to learn everything humans have ever created. IK = E low intelligence * vast knowledge = reasonable effectiveness
- Woodi 2y agoYep, I aways liked encyclopedia. Wiki is good too :) What I would like to have in the future is SO answering-peoples accessible in real time via IRC. They have real answers NOW. They are even pedantic about their stuff !
- wangii 2y agoInteresting formulation! it captures the intuition of the "smartness" when solving a problem. However, what about asking good questions or proposing conjectures?
- hanspeter 2y agoAren't those solutions to problems as well? Find the best questions to ask. Find the best hypothesis to suggest.
- lorepieri 2y agoThere should be also a factor about resource consumption. See here: https://lorenzopieri.com/pgii/ https://lorenzopieri.com/pgii/
- Woodi 2y agoSo article seriously and scientifically states: "Our programs compilation (AI) gave 90% of correct answers in test 1. We expect that in test 2 quality of answers will degenerate to below random monkey pushing buttons levels. Now more money is needed to prove we hit blind alley." Hurray ! Put limited version of that on everybody phones !
- oezi 2y ago> o3 fixes the fundamental limitation of the LLM paradigm – the inability to recombine knowledge at test time I don't understand this mindset. We have all experienced that LLMs can produce words never spoken before. Thus there is recombination of knowledge at play. We might not be satisfied with the depth/complexity of the combination, but there isn't any reason to believe something fundamental is missing. Given more compute and enough recursiveness we should be able to reach any kind of result from the LLM. The linked article says that LLMs are like a collection of vector programs. It has always been my thinking that computations in vector space are easy to make turing complete if we just have an eigenvector representation figured out.
- lagrange77 2y ago> Given more compute and enough recursiveness we should be able to reach any kind of result from the LLM. That was always true for NNs in general, yet it took a very specific structure to get to where we are now. (..with a certain amount of time and resources.) > thinking that computations in vector space are easy to make turing complete if we just have an eigenvector representation figured out Sounds interesting, would you elaborate?
- niemandhier 2y agoContrary to many I hope this stays expensive. We are already struggling with AI curated info bubbles and psy-ops as it is. State actors like Russia, US and Israel will probably be fast to adopt this for information control, but I really don’t want to live in a world where the average scammer has access to this tech.
- owenpalmer 2y ago> I really don’t want to live in a world where the average scammer has access to this tech. Reality check: local open source models are more than capable of information control, generating propaganda, and scamming you. The cat's been out of the bag for a while now, and increased reasoning ability doesn't dramatically increase the weaponizability of this tech, I think.
- deleted 2y ago[deleted]
- pal9000 2y agoCan someone ELI5 how ARC-AGI-PUB is resistant to p-hacking?
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- danielovichdk 2y agoAt what time will it kill us all because it understands that humans are the biggest problem before it can simply chill and not worry. That would be intelligent. Everything else is just stupid and more of the same shit.
- aniviacat 2y agoHumans are the biggest problem of what? Of the sun? Of Venus? Of humans. Humans are a problem for the satisfaction of humans. Yet removing humans from this equation does result in higher human satisfaction. It lessens it. I find this thought process of "humans are the problem" to be unreasonable. Humans aren't the problem; humans are the requirement.
- almog 2y agoAGI ⇒ ARC-AGI-PUB And not the other way around as some comments here seem to confuse necessary and sufficient conditions.
- ImHasanMh 2y ago[dead]
- the5avage 2y agoThe examples unsolved by high compute o3 look a lot like the raven progressive matrix tests used in IQ tests.
- thom 2y agoIt’s not AGI when it can do 1000 math puzzles. It’s AGI when it can do 1000 math puzzles then come and clean my kitchen.
- qup 2y agoIntelligence doesn't have to be embodied.
- thom 2y agoIt also has to be able to come and argue in the comments.
- goatlover 2y agoFor it to be AGI, it needs to be able to manipulate the physical world from it's own goals, not just produce text when prompted. LLMs are just tools to augment human intelligence. AGI is what you see in science fiction.
- egeozcan 2y agoI understand what you are saying and sort of agree the premise but to be pedantic, I don't think any robot can clean a kitchen without doing math :)
- epolanski 2y agoOkay but what are the tests like? At least like a general idea.
- tymonPartyLate 2y agoIsn’t this like a brute force approach? Given it costs $ 3000 per task, thats like 600 GPU hours (h100 at Azure) In that amount of time the model can generate millions of chains of thoughts and then spend hours reviewing them or even testing them out one by one. Kind of like trying until something sticks and that happens to solve 80% of ARC. I feel like reasoning works differently in my brain. ;)
- strangescript 2y ago"We have created artificial super intelligence, it has solved physics!" "Well, yeah, but its kind of expensive" -- this guy
- freehorse 2y agoThe problem is not that it is expensive, but that, most likely, it is not superintelligence. Superintelligence is not exploring the problem space semi-blindly, if the thounsands $$$ per task are actually spent for that. There is a reason the actual ARC-AGI prize requires efficiency, because the point is not "passing the test" but solving the framing problem of intelligence.
- tymonPartyLate 2y agoHaha. Hopefully you’re right and solving the ARC puzzle translates to solving all of physics. I just remain skeptical about the OpenAI hype. They have a track record of exaggerating the significance of their releases and their impact on humanity.
- jeremyjh 2y agoPlease do show me a novel result in physics from any LLM. You think "this guy" is stupid because he doesn't extrapolate from this $2MM test that nearly reproduces the work of a STEM graduate to a super intelligence that has already solved physics. Maybe you've got it backwards.
- strangescript 2y ago
- tikkun 2y agoI wonder: when did o1 finish training, and when did o3 finish training? There's a ~3 month delay between o1's launch (Sep 12) and o3's launch (Dec 20). But, it's unclear when o1 and o3 each finished training.
- zug_zug 2y agoThis is a lot of noise around what's clearly not even an order of magnitude to the way to AGI. Here's my AGI test - Can the model make a theory of AGI validation that no human has suggested before, test itself to see if it qualifies, iterate, read all the literature, and suggest modifications to its own network to improve its performance? That's what a human-level performer would do.
- earth2mars 2y agoMaybe spend more compute time to let it think about optimizing the compute time.
- msoad 2y agoThere are new research where chain of thoughts is happening in latent spaces and not in English. They demonstrated better results since language is not as expressive as those concepts that can be represented in the layers before decoder. I wonder if o3 is doing that?
- padolsey 2y agoI think you mean this: https://arxiv.org/abs/2412.06769 https://arxiv.org/abs/2412.06769 From what I can see, presuming o3 is a progression of o1 and has good level of accountabiltiy bubbling up during 'inference' (i.e. "Thinking about ___") then I'd say it's just using up millions of old-school tokens (the 44 million tokens that are referenced). So not latent thinking per se.
- Zamicol 2y agoInteresting!
- gliptic 2y ago"You can tell the RL is done properly when the models cease to speak English in their chain of thought" -- Karpathy
- rapjr9 2y agoDoes anyone have a feeling for how latency (from asking a question/API call to getting an answer/API return) is progressing with new models? I see 1.3 minutes/task and 13.8 minutes/task mentioned in the page on evaluating O3. Efficiency gains that also reduce latency will be important and some of them will come from efficiency in computation, but as models include more and more layers (layers of models for example) the overall latency may grow and faster compute times inside each layer may only help somewhat. This could have large effects on usability.
- amai 2y agoBut can it convert handwritten equations into Latex? That is the AGI task I'm waiting for.
- rirarobo 2y agoIn my experience, ChatGPT4 has been able to do this very accurately for >1 year now. Gemini also seems to perform well.
- throwaway314155 2y agoHave you not tried this with existing models? Seems like something they would thrive at.
- figure8 2y agoI have a very naive question. Why is the ARC challenge difficult but coding problems are easy? The two examples they give for ARC (border width and square filling) are much simpler than pattern awareness I see simple models find in code everyday. What am I misunderstanding? Is it that one is a visual grid context which is unfamiliar?
- ItsMattyG 2y agoFrancois'(the creator of ARC-AGI benchmark) whole point was that while they look the same, they're not. Coding is solving a familiar pattern in the same way (and fails when it' s NOT doing that, it just looks like it doesn't happen because it's seen SO MANY patterns in code). But the point of Arc AGI is to make each problem have to generalize in some new ay.
- wrsh07 2y agoI expect it largely has to do with "scale" We have an enormous amount of high quality programming samples. From there it's relatively straightforward to bootstrap (similar to original versions of alphago - start with human games, improve via self play) using leetcode or other problems with a "right answer" In contrast, the arc puzzles are relatively novel (why? Well, this has to do with the relative utility of solving an arc problem and programmer open source culture)
- sn0wr8ven 2y agoIncredibly impressive. Still can't really shake the feeling that this is o3 gaming the system more than it is actually being able to reason. If the reasoning capabilities are there, there should be no reason why it achieves 90% on one version and 30% on the next. If a human maintains the same performance across the two versions, an AI with reason should too.
- demirbey05 2y agoI am not expert in llm reasoning but I think because of RL. You cannot use AlphaZero to play other games.
- sgt101 2y agoI thought that AlphaZero could play three games? Go, Chess and Shogi?
- demirbey05 2y agoThink I mean Catan :)
- ozten 2y agoNope. AlphaZero taught itself to play games like chess, shogi, and Go through self-play, starting from random moves. It was not given any strategies or human gameplay data but was provided with the basic rules of each game to guide its learning process.
- demirbey05 2y agoYes its reinforcement learning, but need to create policy and each policy is specialized for specific tasks.
- GaggiX 2y agoHumans and AIs are different, the next benchmark would be build so that it emphasize the weak points of current AI models where a human is expected to perform better, but I guess you can also make a benchmark that is the opposite, where humans struggle and o3 has an easy time.
- earth2mars 2y agoWhy did they skip o2?
- esafak 2y agoTo avoid colliding with https://en.wikipedia.org/wiki/O2_(brand) https://en.wikipedia.org/wiki/O2_(brand)
- YeGoblynQueenne 2y agoI just noticed this bit: >> Second, you need the ability to recombine these functions into a brand new program when facing a new task – a program that models the task at hand. Program synthesis. "Program synthesis" is here used in an entirely idiosyncratic manner, to mean "combining programs". Everyone else in CS and AI for the last many decades has used "Program Synthesis" to mean "generating a program that satisfies a specification". Note that "synthesis" can legitimately be used to mean "combining". In Greek it translates literally to "putting [things] together": "Syn" (plus) "thesis" (place). But while generating programs by combining parts of other programs is an old-fashioned way to do Program Synthesis, in the standard sense, the end result is always desired to be a program. The LLMs used in the article to do what F. Chollet calls "Porgram Synthesis" generate no code.
- tshadley 2y agoI always get the feeling he's subconsciously inserting a "magical" step here with reference to "synthesis"-- invoking a kind of subtle dualism where human intelligence is just different and mysteriously better than hardware intelligence. Combining programs should be straightforward for DNNs, ordering, mixing, matching concepts by coordinates and arithmetic in learned high-dimensional embedded-space. Inference-time combination is harder since the model is working with tokens and has to keep coherence over a growing CoT with many twists, turns and dead-ends, but with enough passes can still do well. The logical next step to improvement is test-time training on the growing CoT, using reinforcement-fine-tuning to compress and organize the chain-of-thought into parameter-space--if we can come up with loss functions for "little progress, a lot of progress, no progress". Then more inference-time with a better understanding of the problem, rinse and repeat.
- baalimago 2y agoLet me know when OpenAI can wrap Christmas gifts. Then I'll be interested.
- esafak 2y agoThey want to leave that menial job for you while taking your office job :)
- codedokode 2y agoI wonder, what is the main obstacle in making robots for mechanical tasks, like laying bricks, paving a road or working in the shaft? It doesn't look like something that requires lot of mathematical or programming skills, just good vision and manipulators.
- neom 2y agohttps://www.youtube.com/watch?v=K1TrbI0BaaU https://www.youtube.com/watch?v=K1TrbI0BaaU
- cambaceres 2y agoCan someone explain to me why this is such a big big deal? I don't know much about AI, but I'm a software developer with a degree in computer science.
- up2isomorphism 2y agoSo what do they test? Some matrix and some matrix out? It does look like “agi” to me.
- itfossil 2y agoThe amount of desperate rationalization in this thread is unbelievable. It's like watching people at a Pentecostal church start speaking in tongues in the hope that something wonderful will happen until it evolves into the realization that shit isn't going to happen and then slowly they just kind of putter out. TLDR: The cacophony of fools is so loud now. Thank goodness it won't last.
- jayseattle 2y ago[dead]
- didibus 2y agoI'm skeptical of these benchmarks. I mean, look at the problem it's solving? I'm sorry, this is our benchmark of AGI, it will never fly with the common person when someone claims AGI, and all it did was fill a grid of pixels. Was it zero-shot at least and Pass@1 ? I guess it was not zero-shot, since it shows examples of other similar problems and their solutions. It also sounds like it was fine-tuned on that specific task. Look, maybe this shows that it could soon be used to replace some MTurk style workers, but I don't know that counts as AGI. To me AGI, it needs to be able to solve novel problems, to adapt to all situations without fine-tuning, and to operate at much larger dimensions, like don't make it a grid of pixels, make it 4k images at least.
- sourcepluck 2y agoAm I understanding correctly, and the only thing with a bit of actual data released so far is the ARC-AGI piece from Francois Chollet? And every other claim has no further data released on it? Serious question. I've browsed around, looked for the official release, but it seems to be just hear-say for now, except for the few little bits in the ARC-AGI article. So some of the reactions seems quite far-fetched. I was quite amazed at first seeing the benchmarks, but then actually read the ARC-AGI article and a few other things about how it worked, learned a bit more about the different benchmarks, and realised we've no proper idea yet how o3 is working under the hood, the thing isn't even realeased. It could be doing the same thing that chess-engines do except in several specific domains. Which would be very cool, but not necessarily "intelligent" or "generally intelligent" in any sense whatsoever! Will that kind of model lead to finding novel mathematical proofs, or actually "reasoning" or "thinking" in any way similar to a human, remains entirely uncertain.
- viivii29 2y agoUse the three prospectives
- edithpixie 2y agoFor many people and businesses, navigating the frequently dangerous landscape of financial loss can be an intimidating and overwhelming process. Nevertheless, the knowledgeable staff at Wizard Hilton Cyber Tech provides a ray of hope and direction with their indispensable range of services. Their offerings are based on a profound grasp of the far-reaching and terrible effects that financial setbacks, whether they be the result of cyberattacks, data breaches, or other unforeseen tragedies, can have. Their highly-trained analysts work tirelessly to assess the scope of the damage, identifying the root causes and developing tailored strategies to mitigate the fallout. From recovering lost or corrupted data to restoring compromised systems and securing networks, Wizard Hilton Cyber Tech employs the latest cutting-edge technologies and industry best practices to help clients regain their financial footing. But their support goes beyond the technical realm, as their compassionate case managers provide a empathetic ear and practical advice to navigate the emotional and logistical challenges that often accompany financial upheaval. With a steadfast commitment to client success, Wizard Hilton Cyber Tech is a trusted partner in weathering the storm of financial loss, offering the essential services and peace of mind needed to emerge stronger and more resilient than before.
- ireneerika 2y ago[dead]
- viivii29 2y agoUse the three prospectives with mini
- Peterthomos 2y ago[dead]