14 ms·
Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes
by bluecoconut 2y ago
Efficiency is now key.
~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task.
We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective experience) between 5 second and 5 minutes to solve the task.
(So i'd argue a human is at 0.03USD - 1.67USD per puzzle at 20USD/hr, and they include in their document an average mechancal turker at $2 USD task in their document)
Going the other direction: I am interpreting this result as human level reasoning now costs (approximately) 41k/hr to 2.5M/hr with current compute.
Super exciting that OpenAI pushed the compute out this far so we could see he O-series scaling continue and intersect humans on ARC, now we get to work towards making this economical!
- riku_iki 2y ago> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
- bluecoconut 2y agoThat's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k
- ionwake 2y agosorry to be a noob, but can someone tell me doe sths mena o3 will be unaffordable for a typical user? Will only companies with thousands to spend per query be able to use this? Sorry for being thick Im just confused how they can turn this into an addordable service?
- JohnnyMarcone 2y agoThere are likely many efficiency gains that will be made before it's released, and after. Also they showed o3 mini to be better than o1 for less cost in multiple benchmarks, so there're already improvements there at a lower cost than what available.
- ionwake 2y agoGreat thank you
- HDThoreaun 2y agoThe low compute one did as well as the average person though
- jhrmnn 2y agoThat’s for the low-compute configuration that doesn’t reach human-level performance (not far though)
- riku_iki 2y agoI referred on high compute mode. They have table with breakdown here: https://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough
- EVa5I7bHFq9mnYK 2y agoThat's high EFFICIENCY. High efficiency = low compute.
- gbnwl 2y agoThat's "efficiency" high, which actually means less compute. The 87.5% score using low efficiency (more compute) doesn't have cost listed.
- bluecoconut 2y agothey use some poor language. "High Efficiency" is O3 Low "Low Efficiency" is O3 High They left the "Low efficiency" (O3 High) values as `-` but you can infer them from the plot at the top. Note the $20 and $17 per task aligns with the X-axis of the O3-low
- junipertea 2y agoThe table row with 6k figure refers to high efficiency, not high compute mode. From the blog post: Note: OpenAI has requested that we not publish the high-compute costs. The amount of compute was roughly 172x the low-compute configuration.
- deleted 2y ago[deleted]
- binarymax 2y ago"Note: OpenAI has requested that we not publish the high-compute costs. The amount of compute was roughly 172x the low-compute configuration." The low compute was $17 per task. Speculate 172*$17 for the high compute is $2,924 per task, so I am also confused on the $3400 number.
- bluecoconut 2y ago3400 came from counting pixels on the plot. Also its $20 on for the o3-low via the table for the semi-private, which x172 is 3440, also coming in close to the 3400 number
- xrendan 2y agoYou're misreading it, there's two different runs, a low and a high compute run. The number for the high-compute one is ~172x the first one according to the article so ~=$2900
- Thorrez 2y agoWhat's extra confusing is that in the graph the runs are called low compute and high compute. In the table they're called high efficient and low efficiency. So the high and low got swapped.
- deleted 2y ago[deleted]
- bluecoconut 2y agosome other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculations: The gap in cost is roughly 10^3 between O3 High and Avg. mechanical turkers (humans). Via Pure GPU cost improvement (~doubling every 2-2.5 years) puts us at 20~25 years. The question is now, can we close this "to human" gap (10^3) quickly with algorithms, or are we stuck waiting for the 20-25 years for GPU improvements. (I think it feels obvious: this is new technology, things are moving fast, the chance for algorithmic innovation here is high!) I also personally think that we need to adjust our efficiency priors, and start looking not at "humans" as the bar to beat, but theoretical computatble limits (show gaps much larger ~10^9-10^15 for modest problems). Though, it may simply be the case that tool/code use + AGI at near human cost covers a lot of that gap.
- zamadatix 2y agoI don't follow how 10 random humans can beat the average STEM college grad and average humans in that tweet. I suspect it's really "a panel of 10 randomly chosen experts in the space" or something? I agree the most interesting thing to watch will be cost for a given score more than maximum possible score achieved (not that the latter won't be interesting by any means).
- hmottestad 2y agoMight be that within a group of 10 people, randomly chosen, when each person attempts to solve the tasks at least 99% of the time 1 person out of the 10 people will get it right.
- deleted 2y ago[deleted]
- bcrosby95 2y agoTwo heads is better than 1. 10 is way better. Even if they aren't a field of experts. You're bound to get random people that remember random stuff from high school, college, work, and life in general, allowing them to piece together a solution.
- spencerchubb 2y ago> Super exciting that OpenAI pushed the compute out this far it's even more exciting than that. the fact that you even can use more compute to get more intelligence is a breakthrough. if they spent even more on inference, would they get even better scores on arc agi?
- echelon 2y agoMaybe it's not linear spend.
- lolinder 2y ago> the fact that you even can use more compute to get more intelligence is a breakthrough. I'm not so sure—what they're doing by just throwing more tokens at it is similar to "solving" the traveling salesman problem by just throwing tons of compute into a breadth first search. Sure, you can get better and better answers the more compute you throw at it (with diminishing returns), but is that really that surprising to anyone who's been following tree of thought models? All it really seems to tell us is that the type of model that OpenAI has available is capable of solving many of the types of problems that ARC-AGI-PUB has set up given enough compute time. It says nothing about "intelligence" as the concept exists in most people's heads—it just means that a certain very artificial (and intentionally easy for humans) class of problem that wasn't computable is now computable if you're willing to pay an enormous sum to do it. A breakthrough of sorts, sure, but not a surprising one given what we've seen already.
- mithametacs 2y agoAn algorithm designed for translating between human languages has now been shown to generalize to solving visual IQ test puzzles, without much modification. Yes, I find that surprising.
- freehorse 2y ago> I am interpreting this result as human level reasoning now costs (approximately) 41k/hr to 2.5M/hr with current compute. On a very simple, toy task, which arc-agi basically is. Arc-agi tests are not hard per se, just LLM’s find them hard. We do not know how this scales for more complex, real world tasks.
- SamPatt 2y agoRight. Arc is meant to test the ability of a model to generalize. It's neat to see it succeed, but it's not yet a guarantee that it can generalize when given other tasks. The other benchmarks are a good indication though.
- criddell 2y agoDoes it mean anything for more general tasks like driving a car?
- brookst 2y agoIs every smart person a good driver?
- zarzavat 2y agoLikely yes. Every smart person is capable of being a good driver, so long as you give them enough training and incentive. Zero smart people are born being able to drive.
- fragmede 2y agoThere are different kinds of smarts and not every smart person is good at all of them. Specifically, spacial reasoning is important for driving, and if a smart person is good at all kinds of thinking except that one, they're going to find it challenging to be a good driver.
- 2y ago
- madduci 2y agoLet's see when this will be released to the free tier. Looks promising, although I hope they will also be able to publish more details on this, as part of the "open" in their name
- daxfohl 2y agoI wonder if we'll start seeing a shift in compute spend, moving away from training time, and toward inference time instead. As we get closer to AGI, we probably reach some limit in terms of how smart the thing can get just training on existing docs or data or whatever. At some point it knows everything it'll ever know, no matter how much training compute you throw at it. To move beyond that, the thing has to start thinking for itself, some auto feedback loop, training itself on its own thoughts. Interestingly, this could plausibly be vastly more efficient than training on external data because it's a much tighter feedback loop and a smaller dataset. So it's possible that "nearly AGI" leads to ASI pretty quickly and efficiently. Of course it's also possible that the feedback loop, while efficient as a computation process, isn't efficient as a learning / reasoning / learning-how-to-reason process, and the thing, while as intelligent as a human, still barely competes with a worm in true reasoning ability. Interesting times.
- matusp 2y agoI don't think this is only about efficiency. The model I have here is that this is similar to when we beat chess. Yes, it is impressive that we made progress on a class of problems, but is this class aligned with what the economy or the society needs? Simple turn-based games such as chess turned out to be too far away from anything practical and chess-engine-like programs were never that useful. It is entirely possible that this will end up in a similar situation. ARC-like pattern matching problems or programming challenges are indeed a respectable challenge for AI, but do we need a program that is able to solve them? How often does something like that come up really? I can see some time-saving in using AI vs StackOverflow in solving some programming challenges, but is there more to this?
- edanm 2y agoI mostly agree with your analysis, but just to drive home a point here - I don't think that algorithms to beat Chess were ever seriously considered as something that would be relevant outside of the context of Chess itself. And obviously, within the world of Chess, they are major breakthroughs. In this case there is more reason to think these things are relevant outside of the direct context - these tests were specifically designed to see if AI can do general-thinking tasks. The benchmarks might be bad, but that's at least their purpose (unlike in Chess).
- spamlettuce 2y agookay, but what about literal swe-bench. O3 scored 75% eval
- lugu 2y agoARC is designed to be hard for current models. It cannot be a proxy for how useful they are. It says something else. Most likely those models won't replace human at their tasks in their organization. Instead "we" will design pipeline so that the tasks aligns with the ability of the model and we will put the human at the periphery. Think of how a factory is organised for the robots.
- cle 2y agoEfficiency has always been the key. Fundamentally it's a search through some enormous state space. Advancements are "tricks" that let us find useful subsets more efficiently. Zooming way out, we have a bunch of social tricks, hardware tricks, and algorithmic tricks that have resulted in a super useful subset. It's not the subset that we want though, so the hunt continues. Hopefully it doesn't require revising too much in the hardware & social bag of tricks, those are lot more painful to revisit...
- chefandy 2y agoI think the real key is figuring out how to turn the hand-wavy promises of this making everything better into policy long fucking before we kick the door open. It’s self-evident that this being efficient and useful would be a technological revolution; what’s not self evident is that it wouldn’t benefit the large corporate entities that control even more disproportionately than it does now to the detriment of many other people.
- ein0p 2y agoThis is beta version. By the time they're done with this it'll be measured in single digit dollars, if not cents.
- Macuyiko 2y agoI am not so sure, but indeed it is perhaps also a sad realization. You compare this to "a human" but also admit there is a high variation. And, I would say there are a lot humans being paid ~=$3400 per month. Not for a single task, true, but for honestly for no value creating task at all. Just for their time. So what about we think in terms of output rather than time?