10 ms·
OpenAI's GPT-6 Astra on ARC-AGI-3
- yusufozkan 1mo agowhat the hell is that score/cost curve lol
- minimaxir 1mo agoDeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.
- Phemist 1mo agoWhat is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?
- minimaxir 1mo agoYes, in theory.
- tedsanders 1mo agoYep. In particular, ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it bend backward.
- Frost1x 1mo agoIt’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper. The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).
- Betelbuddy 1mo ago"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses. Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted." Well I dont know about all of you, but I am celebrating meat based humans...
- LPisGood 1mo agoI think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.
- paxys 1mo agoWhy are you making the assumption that a person's time is worthless? I'd argue that it is the single most valuable resource we all have.
- piloto_ciego 1mo ago99.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.
- slopinthebag 1mo agoThat’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.
- piloto_ciego 1mo agoI agree, like the average human isn't generally intelligent. IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!
- slopinthebag 1mo agoThe average human would have average intelligence
- defrost 1mo agoWhat's the metric on "average human" from a global population of > 8 billion people .. and why would they score mid on a standardised IQ test skewed toward western education / culture?
- slopinthebag 29d agoWho said anything about IQ tests?
- dwohnitmok 1mo ago> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open. Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"
- malfist 1mo agoIs solving a snake like puzzle game in the least number of moves really what defines intelligence?
- GaggiX 1mo agoIf you have never seen the game before probably.
- jawiggins 1mo agoThere's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
- malfist 1mo agoI am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it
- hyperhello 1mo agoThis is like arguing about whether a hot dog is a sandwich (of course it is) or whether the chicken or the egg was first (obviously the egg since all chickens come from eggs). Intelligence is just problem solving in the context of self-awareness. Machines don't have it and never will but they can simulate the process given inputs. You can argue whether humans and animals truly possess self-awareness and in what degree, but the definition of intelligence is as simple as the hot dog debate.
- whattheheckheck 1mo agoIt was defined in Animal Intelligence by George John Ramones in 1882 as "intelligence is the capacity to do the right thing at the right time. It is the ability to respond to the opportunities and challenges presented by a context"
- eli 1mo agoWhy? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.
- mikert89 1mo agoAnything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
- x3haloed 1mo agoYup. Only subjective taste remains.
- GPerson 1mo agoNope that will be commodified in short order.
- tedsanders 1mo agoDisagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
- mikert89 1mo agothese just need more compute: - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn but we both know these examples go against the spirit of my point
- tedsanders 1mo agoPerhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).
- hypfer 1mo agoWhat are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?
- petu 1mo agoOpenAI provides API key with ~unlimited use?
- Frost1x 1mo agoSo, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute. Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going. Alignment++
- manquer 1mo agoYou missed the hard part getting on HN front page , ie. Getting the acceptance of the community / zeitgeist . There is no incentive for OpenAI to subsidize is you if no one reads /reports on your benchmark . They are only going to fund a few that are currently popular . Community acceptance doesn’t automatically mean the best , it is combination of some level of technical quality and the ability of the promoter to socially influence or get support of influencers .
- 6thbit 1mo agoThe instant/no reasoning performed extremely well none 35.2%, $49,791 96.7%, $23,457 35.2% on the standard harness, that's above Opus 5 on high.
- NitpickLawyer 1mo agoSince low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).
- bigbuppo 1mo agoWake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.
- fastball 1mo agoAre we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
- bigzyg33k 29d agoThat counts as 200% on exploitbench to me!
- yomismoaqui 1mo agoNow that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI? Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)
- wise_blood 1mo agorepeating myself, but: once 3 is solved, we would come up with 4. then 5, 6... it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.
- john_alan 1mo agoexactly, this isn't AGI.
- deleted 1mo ago[deleted]
- fxd 1mo ago“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it. But you’d be no nearer to solving consciousness. Given this thought trajectory - what is AGI supposed to be?
- eagerpace 1mo agoI like recursive self improvement instead. It seems like something that is actually quantifiable and kinda “the point” of why consciousness is important to humans.
- fxd 1mo agoSo basically, being able to set it free on some long running goal and it sort of “lives” and autonomously does its own tasks? I wonder at what point consciousness is necessary… that is, if you can have anything like that without it. To the point that solving consciousness (and combining it with intelligence) is what gives you the autonomous, recursive, self-improving thing otherwise it can only drive in the dark and make big mistakes. To your point I think - it’s why we don’t see too many non-conscious advanced biology (it rarely survives against those with it).
- layer8 1mo agoNot sure why you are bringing up consciousness, that’s largely orthogonal to intelligence. AGI is usually taken to mean the capability to match or surpass human intelligence across all conceivable cognitive tasks, as opposed to being limited to certain kinds of tasks, or to not matching the general level of human intelligence in some respect. Intelligence, and hence AGI, doesn’t require consciousness or emotions or sentience.
- fxd 1mo ago[dead]
- an0malous 1mo agoWas OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.
- an0malous 1mo agoActually I can answer my own question: we know that they have had previous access to the tests because they’ve run older models against the same benchmark. I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.
- ajjahs 1mo ago[dead]
- modeless 1mo ago$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.
- WASDx 1mo agoOnce it figures out a puzzle it could probably be instructed to design a specialized harness for Luna to be able to solve other instances of the same puzzle. Minimum wage workers are not solving novel problems.
- at1as 1mo agoI like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken. From https://epoch.ai/latest/announcing-frontiermath-erdos https://epoch.ai/latest/announcing-frontiermath-erdos > Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours > Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself. Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.
- red75prime 1mo agoA very long tail of problems that weren't solved by humans? Sure. It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.
- at1as 1mo agoWhat does this mean? There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems https://github.com/teorth/erdosproblems Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're past that now). New models will exhibit something novel by adding solutions. And the matter of solution is interesting as well (contradiction versus a positive proof). I would venture to predict it'll take years to decades to get to 0 open problems. But I'd be very happy to have this comment look foolish in retrospect as models continue to improve
- 1mo ago
- scotty79 1mo agoHow good are LLMs at doing Mensa tests?
- brokensegue 1mo agoIQ tests? Very good. But most are in the dataset so it's not very meaningful
- lofaszvanitt 29d agoNice, but what about pelicans? No proper pelican means it's sitting on a horse with only a half arse :D.
- unixhero 29d agoHow do you feel about Astra pretty much reaching our current definition of AGI?