6 ms·
Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "
by mvkel 29d ago
Take it from the mouth of the creator of ARC-AGI:
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
- iterateoften 29d ago2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
- azan_ 29d agoYes, 6 months vs 1 year is huge for technology that has gained wider adoption only recently.
- fn-mote 29d agoAdoption means nothing. 2x gains from a mature technology would be surprising. 2x gains from a new tech would still be called “low hanging fruit” in another setting. I don’t read enough to know in what ways the training / other technical steps have really advanced.
- anvuong 29d agoYou'll also need to compare the amount of compute used now and then, which seems exponential to me.
- abixb 29d ago>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
- z7 29d agoFrançois Chollet wrote in February that he expected ARC-3 to be saturated in "about one year". "Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast." https://x.com/fchollet/status/2022054537293705260 https://x.com/fchollet/status/2022054537293705260
- giancarlostoro 29d agoI feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim? I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful. > AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer. - Sam Altman on AGI
- lenerdenator 29d agoI wonder if Altman's definition also includes taking on the same liability as a coworker would. Probably not.
- giancarlostoro 29d agoWould 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford.
- lenerdenator 29d agoAnd that's the rub, isn't it? If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit. If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.
- alex0015 29d agoWhat do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider. If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.
- gavinray 29d agoHuman brains have difficulty reasoning about exponential growth.
- osigurdson 29d agoThey keep saying that. I'd say it is more like human brains that don't remember high school math have trouble with it.
- saimiam 29d agoIf something at rest is accelerating at 9.8 m/s^2, how long in seconds will it take to reach 10% of c? Answer to the nearest order of magnitude - will it take approximately 1000, 10k, 100k, 1000k seconds? I’m sure you know this is an exponential growth question but have no intuition of the answer.
- MajesticHobo2 29d agoThat is a linear growth problem whose answer is very easy to intuit.
- saimiam 28d agoIt’s not a linear growth problem. Acceleration is quadratic. As per your intuition, how many seconds later will you hit 10% of c?
- redsocksfan45 28d ago[dead]
- yunwal 28d agoAcceleration is not quadratic with respect to speed/velocity, and c is a speed/velocity.
- kccqzy 29d ago
- balefulboy 29d agoWell it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
- nearbuy 29d agoI think you're misunderstanding. Astra is at the top of the official ARC-AGI leaderboard, with an ARC-AGI approved harness. It's not a harness specialized for ARC-AGI. It just does the same thing the regular ChatGPT interface does: keeps conversation history across turns and compacts when it gets too long. Without the harness, it loses its entire context window every move. That's not how humans work and it's not how any real AI service works.
- Forgeties79 29d agoWe are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic. Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation. I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.
- hereme888 29d agoWhy is a >30-min context-length a requirement for AGI, or the naturalness of a human conversation?
- Forgeties79 28d agoDo you descend into repetitive, incoherent babble after a 30min, incredibly focused conversation? Does any human being without some sort of diagnosis? I assume you regularly have conversations that last more than 30 minutes (work meeting, for instance, which gpt could not participate in as an equal voice by any stretch of the imagination). It’s an arbitrary number that felt high enough. If I said 10, people would argue that a frontier model can talk longer than that. But I have absolutely watched ChatGPT fall apart that quickly. I’m sure everyone reading this has. We can nitpick the duration all you want, but we both know it does not take very long for this to occur. It happens particularly fast if you stray from the original topic and/or aren’t using a frontier model.
- hereme888 28d agoI work exclusively with OpenAI cloud models, and most high-quality agent harnesses auto-compact and does pretty well sticking to the plan of action. Papers that research the definition of AGI never include context length, and I myself don't see the logic for it either.
- tomjen3 29d agoAre there plans for ARC 4?
- mvkel 27d agoThere are plans up through ARC 7 on the roadmap
- zit-hb 29d agoAbsolutely. You should never underestimate the compounding effect such a release can have. Having the right tools to create new tools.