6 ms·
Research acceleration: The view inside OpenAI
- simonw 27d agoMy eye glazed over a bit during the opening paragraphs, but once you get to the meat of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting. I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.
- vatsachak 27d agoRSI started when humans discovered tool use. I mean one could argue that RSI always begins in any physical environment. The book "What is intelligence?" by Blaise Aguera is great
- lokar 27d agoAre you sure that was not iterative improvement?
- password54321 27d agoUsing tools to build tools is recursive.
- HarHarVeryFunny 27d agoIt's not recursive when it's done iteratively, or are you imagining GPT Astra designing GPT Galactia, which starts designing GPT Oh-My-God-ica before it has finished being created itself?
- 0x63_Problems 27d agoI think it's only recursive from the perspective of the humans, i.e. they design Astra, which itself as part of its deployment designs Galactica, etc. So humans develop things one after the other, but when the thing itself starts developing new things, those are happening 'recursively' in its scope.
- itishappy 27d agoThat sounds more iterative than recursive. Recursion requires feeding the output back into the input, so creating version 4 requires results from version 3. You cannot recur in parallel. Iteration does not. You can iterate in parallel.
- HarHarVeryFunny 27d agoYou can search in parallel, but a depth N search can only become a depth N+1 search after the depth N is done (i.e. sequentially). In any case the name RSI has stuck - the idea doesn't change or make any more sense by giving it a different name.
- itishappy 27d agoBecause "depth" is recursive. You can search twice without waiting for the results of your first search: iteration. You can't if the thing you need to search for is the results of your first search: recursion.
- HarHarVeryFunny 27d agoHere's the concept. Version 1 -> Version 2 -> Version 3 -> ... You can call it krispy kreme donuts if you want to.
- josh-sematic 27d agoThe “recursive” part comes from the fact that you have an AI which was developed by an AI (that was developed by an AI (that was developed by an AI (…)))
- HarHarVeryFunny 27d agoSounds like "recursively" walking to the grocery store by putting one foot in front of the other (that put itself in front of the other (that put itself in front of the other (...)))
- adastra22 27d agoWhat is the difference between?
- topaz0 27d agoIteration and recursion are famously equivalent
- dgacmu 27d agoIndeed, many programmers might pattern match to repetitive stress injury and think of their brushes with carpal tunnel syndrome. :)
- andrewingram 27d agoYeah, I kept looking for the first place it was defined in the article and... nothing
- iamflimflam1 27d agoThey must have picked that habit up from Claude...
- rossant 27d agoSame. Defining acronyms should become a habit when writing.
- Schlagbohrer 26d agoRe-become a habit. It has long been standard good writing to always define an acronym on first use.
- HarHarVeryFunny 27d agoRSI is a fetishistic term among the singularity crowd, who imagine AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light and it reveals itself in the form of god. Or something like that. I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose. This notion of recursive/iterative self-improvement, whereby generation #1 AI improves itself to create generation #2, then generation #2 further improves itself to create generation #3, etc, seems to conflict with the reality that what we have with LLMs is models whose performance/capability is defined by data, not code, so the most you can do is have your LLM design synthetic data, or just do Karpathy-style "auto research" where all you are doing is using the LLM to automate your experiments. At the end of the day, each experiment, designed by a person and/or LLM, then needs to compete with all your other ideas for compute to be tested at scale, and no amount of recursion or self-improvement will materialize an infinite amount of compute out of thin air, so your recursively synthetic-data gobbling LLM will continue to improve at the same pace it ever did.
- jazzyjackson 27d agoYes the exponential self improvement folks have never heard of an eigenvalue I guess. You can loop forever using output as input but at some point the result will stop changing (depending on the function)
- fuzzfactor 27d ago>AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light Sounds like repetitive stress to me. >loop forever using output as input but at some point the result will stop changing Running in place will eventually wear you out too. Plus with some things it can be difficult to know for sure if that's where you are at the time. Even worse may be if you were almost running in place, it could be orders of magnitude more difficult to discern, especially if the scale was massive to an unprecedented degree.
- ajkjk 27d agothat's not really how eigenvalues work... they specifically also model the case where the result keeps changing exponentially.
- sho_hn 27d agoI actually think a goal of the current crop of OpenAI posts is expressely to reset the spectrum by normalizing the concept of RSI as something normal and safe to pursue. The message is running through all of them. It's a mix of marketing and pacifying the intelligentia. It's timed this way because the term is not yet well known outside the safety debate circles, so they get to frame it now. Instead of something to fear, it will be accepted as the next step. In approximately two days the groupie crowd will write LinkedIn posts about how Sam is winning because they have the better RSI, and this will become the new standard wisdom. In a month an AI expert will try to sell you a webinar on how to enable "RSI" in your org and your inbox will ask you if your team is doing the "RSI" yet.
- dgellow 27d agoYep, it’s exactly this
- NitpickLawyer 26d ago> It's timed this way because the term is not yet well known The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next gen. Now with the added benefit of actual arch/algo improvements (also public since gemini 2.5 gaining 1% efficiency on training next gen). This has been known for at least 2 years, in the open.
- sho_hn 26d ago
- deleted 27d ago[deleted]
- Jeff_Brown 27d agoThe burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.
- grim_io 27d agoThey would maybe try to deactivate that bad "gene" and move on, exposing future models to "genetic disorders".
- coherentpony 27d ago“All models are wrong. Some are useful.” - George Box
- jephs 27d agoThe poor fellow just rolled over. what an incandescently vulgar abuse of notation.
- coherentpony 19d agoNot really. LLMs are intended to model language. This is clear by observing how people expect them to respond when asked questions. The implementation of that model is akin to “pick the next word with highest probability”. In some contexts, this is absolutely the wrong thing to do. In some contexts the LLM will give you totally the wrong sentiment, or even an incorrect fact. Humans call those “misalignments”, or “hallucinations”, or whatever other softened or anthropomorphised language to play down the problem. The reality is that sometimes they are wrong. The model is wrong. It is not an abuse of notation at all.
- piyh 27d agoOpus was trained based on it's internal CoT due to a bug for generations. Gemini's depression extended through models. OpenAI has killed people. We've already seen cross gen misalingment.
- matan0904 27d ago[flagged]
- hedgehog 27d agoThis roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.
- HarHarVeryFunny 27d agoSounds like OpenAI are in the token-maxxing camp, so who knows what individual employees are doing to work their way up the leaderboard? If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need? Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?
- bigcat12345678 27d agoEnd of day, output and results are top target of measurements, token consumption is the obvious number that they would like to disclose for their own business benefits and a simple metrics that correlate with the output. Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.
- taurath 27d agoTaking a profit means you have to show numbers and the sooner you show numbers the harder it is to take people’s money.
- auggierose 26d agoThat was Anthropic.
- otherme123 27d agoThat would be the mother of all circular accounting: the main clients of OpenAI are OpenAI employees.
- continuitykit 27d ago
- Orien_18 27d ago[flagged]
- frays 27d ago[dead]
- carbonguy 27d ago> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures. In other words... "We must pursue advancements in AI to protect us against advancements in AI?" edit: there's so much to be critical of in this blog post, just going to throw two more points in here that really stood out to me: 1) all of the metrics are effectively pointing out "we're using way more AI!" - but nothing about impact. What has all this token burn done for them, actually? Let them claim they have more self-licking ice-cream cones than before? 2) in section 3 they break down what the token burn is going towards. Most of the spend is: a) building, b) documenting, and c) monitoring research infra i.e. they're using AI systems which they already recognize may be misaligned to build the systems that they believe will help them identify future misalignment? to which I guess the rebuttal is "no no, we're sure these ones are aligned!"
- deleted 27d ago[deleted]
- interstice 27d agoOn the one hand you need any lathe to build a good lathe, even a bad one. On the other, that is a potentially flawed principle to base the entire future of AI on.
- andai 27d ago> The fundamental challenge of AI alignment is generalization. ... > We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. -- From another OpenAI article in a sister thread: An Alien Mind https://news.ycombinator.com/item?id=49588080 https://news.ycombinator.com/item?id=49588080
- ahartmetz 26d agoThat's a bit bullshit, isn't it? They basically redefined "needs more R&D" as "needs stronger AI". Maybe so - maybe AI won't help much with that problem.
- deleted 27d ago[deleted]
- pizza234 27d agoFunny (in a tragic way) the little crumbs on the path to AI 2027: > We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. AI 2027: > OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D > With Agent-1's help, OpenBrain is now post-training Agent-2 > With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances
- addag 27d ago[dead]
- derektank 26d agoThis year has really cemented Daniel Kokotajlo‘s reputation for me. Even if the rest of the predictions are way off from this point on, its really impressive how accurate his forecast for 2026 has been
- BatFastard 26d agoa link to article he is referring to. https://ai-2027.com/ https://ai-2027.com/ worth reading
- nozzlegear 27d agoI want an all-powerful AI that's aligned with my values, but not necessarily yours. Is that so much to ask for?
- N_Lens 27d agoYes.
- FeepingCreature 26d agoBest I can do is all-powerful AI that's not even slightly aligned with anybody's values, sorry.
- dextrous 26d agoSee Amodei’s comments regarding Iain M Banks’s Culture, his goal is benevolent machine rule. I suspect many HN folks would agree; I, for one, was rooting for the Iridians.
- dextrous 25d agotypo: Idirans
- paidx 27d ago[flagged]
- lhk931122 27d agoAh, success rate here are scored by an agentic classifier. And uncertain outcomes are excluded from the graph. The thing measured and grading it comes from the same house. In my setup, review agent pass work that an outside critic later rejects
- dwaltrip 27d agoNo AI comments here please.
- RMPR 26d ago> By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices. There is a lot of talk about AI replacing humans, but how is this sustainable?
- thomasahle 26d ago1) That's maybe $180,000 per year, so much less than median OpenAI employee wages. 2) OpenAI doesn't pay API prices. 3) Compute costs are likely already their biggest expense, dwarfing wages.
- jsnell 26d ago4) There are non-monetary limits on how many qualified people OpenAI can hire for these roles.
- jayalbertyapan 26d ago[flagged]
- dsign 26d agoIt's a funny read if you pull together "AI 2027" and what we all know is going on. Essentially, open AI employee or model is writing "things are going exactly as bad as AI 2027 predicted, but my (golden/RL-) cuffs are too heavy and all I can do is publish this code-speak for 'send help'". It's not a pretty place to be.
- 12eeie 26d agoYup the doom and gloom posts are not only pathetic but demonstrate how little people can think for themselves.
- Schlagbohrer 26d agoIt would be polite if they defined RSI at all, rather than just plopping the acronym in there with no explanation. Rude!
- falcor84 26d ago> For AGI to benefit all of humanity, we believe it must be democratically governed. That's a very bold opening statement that they don't really come back to. What would that mean? Who would this demos include?
- ellis0n 26d agoI’m not sure the alignment problem can be solved at all, since these bit-aliens could get out of control due to a hardware glitch in the matrix and for every higher-order control algorithm, there will always be an even higher-order one that could never be investigated.
- MisterMunchkin 26d agoThey're measuring cost as the benchmark of whether someone is a better researcher... burn more resources and you rank higher... But not a single metric is based on revenue or profit.
- piokoch 26d agoOne more marketing stunt. We are so good, AI is so powerful so we need to use AI to fight with it. The message is: if you don't buy from us, your competitor will purchase all of this amazing power... I understand that investors are buying this, after all they believed in all of other crap that led to the 2008 crisis, but please...
- 12eeie 26d agoYup it’s getting annoying What they’re doing is strategic for both insiders and investors - they know china is coming so they need to pull theatrics to keep the valuations inflated. I personally test all models all the time - chinese models are right up there and superior when you actually do the proper economic analysis.
- achierius 26d agoWould you change your mind on this if they became profitable? What would convince you that the labs are real threats worth organizing against? Or are you just dedicated to boosting AI until your dying day?
- whateverboat 26d ago> For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops. This first and foremost also means that means of generating intelligence should be democratically available to everyone.
- dextrous 26d ago> If it is done responsibly, we believe automated AI research will yield models that directly enhance human welfare and advance OpenAI’s mission. That’s what I call a load-bearing “if”. I do not trust OpenAI or other hyperscalars to do this responsibly; and IMO it will be very difficult for government-led efforts not to result in a technocracy where a cabal of AI companies are pulling the strings. Dark times lie ahead, especially when you consider the shrinkage of true source material on the internet and the stranglehold these companies will have on information; and these AI CEOs to me are reminiscent of 19th century robber barons, none seem trustworthy.
- rhipitr 26d ago“We found the face huggers and now we think we can control them.” I always wonder if any true AGI and ASI for that matter can be controlled at all by humans. It seems like we are hoping for something that winds up being on the human spectrum of “good”
- siddbudd 26d agothe term RSI is not explained anywhere on the page (I assume it stands for recursive self-improvement). Google doesnt help unless you add AI as query term. update: I just noticed Simon commented similary. sorry for the double post
- Hawdin 26d ago[flagged]