15 ms·
The Darwin Gödel Machine: AI that improves itself by rewriting its own code
- billab995 1y agoWhen does it begin to learn at a geometric rate?
- interludead 1y agoSounds nice! Especially with the Sakana's latest development of Continuous Thought Machine. The next step should be to let foundation models fine-tune themselves based on their 'history of what has been tried before' and new data
- deleted 1y ago[deleted]
- ordinarily 1y agoThe pieces are coming together quickly https://ai-2027.com/ https://ai-2027.com/.
- candiddevmike 1y agoThis reads like an advertisement for OpenBrain and doesn't seem grounded in reality.
- ordinarily 1y agoI think the general tone is more of a warning than an endorsement.
- dmonitor 1y agoI can't help but notice that it doesn't matter what DeepCent does because OpenBrain will reach self awareness 6 months before them no matter what. Who needs a profitability plan when you're speedrunning the singularity.
- dbrail 1y ago[dead]
- deleted 1y ago[deleted]
- brookst 1y agoI was a bigger fan of the certain doom in 2025, and I think the AI 2030 movement will have better design sense and storytelling. But really I haven’t seen anything that really has the oomph and fire of Tipper Gore’s crusade against youth music. We need more showmanship, more dramatic catastrophizing. I feel like our current crop of doomers isn’t quite shameless enough to be really entertaining.
- nosianu 1y agoA significant thing to keep in mind for non-extinction doomerism is that individual experiences vary greatly. There may be a significant number of people or groups that really do experience what was predicted. Similar to how the experiences of average rise in temperature (I would prefer if they had used the term "energy") differ greatly dependent on the region. Also similar to "the country is doing well, look at the stick market and the GDP". I think everybody who wants to have an actually serious discussion needs to invest a lot more effort to get tall those annoying "details", and be more specific. That said, I think that "AI 2027" link looks like it's a movie script and not a prediction, so I'm not sure criticizing it as if it was something serious even makes sense - even if the authors should mean what they write at the start and themselves actually take it seriously.
- pram 1y agoits literally just the plot of “Colossus: The Forbin Project” so it isnt even original lol
- brookst 1y ago100% agreed! We think about the industrial revolution and the rise of word processors and the Internet as social goods, but they were incredibly disruptive and painful to many, many people. I think it’s possible to have empathy for people who are negatively affected without turning it into a “society is doomed!“ screed
- tazjin 1y agoChecked out when it turned into bad geopolitics fiction.
- Workaccount2 1y agoPeople should understand that the reason this seemingly fan-fict blog post gets so much traction is because of lead author's August 2021 "fan-fict" blog post, "What 2026 Looks Like": https://www.alignmentforum.org/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like https://www.alignmentforum.org/posts/6Xgy6CAf2jqHhynHL/what-...
- Der_Einzige 1y agoSo this is what the crowd of people who write SCP articles with over 1000 upvotes does in their professional life?
- jerpint 1y agoI have a feeling LLMs could probably self improve up to a point with current capacity, then hit some kind of wall where current research is also bottle necked. I don’t think they can yet self improve exponentially without human intuition yet , and the results of this paper seem to support this conclusion as well. Just like an LLM can vibe code a great toy app, I don’t think an LLM can come to close to producing and maintaining production ready code anytime soon. I think the same is true for iterating on thinking machines
- matheusd 1y ago> I don’t think they can yet self improve exponentially without human intuition yet I agree: if they could, they would be doing it already. Case in point: one of the first things done once ChatGPT started getting popular was "auto-gpt"; roughly, let it loose and see what happens. The same thing will happen to any accessible model in the future. Someone, somewhere will ask it to self-improve/make as much money as possible, with as little leashes as possible. Maybe even the labs themselves do that, as part of their post-training ops for new models. Therefore, we can assume that if the existing models _could_ be doing that, they _would_ be doing that. That doesn't say anything about new models released 6 months or 2 years from now.
- __loam 1y agoPeople in the industry have been saying 6 months to agi for 3 years.
- glenstein 1y agoThey had been saying it was 10 years away for ~50 years, so that's progress. Soon it will be 1 month away, for another two years. And when they say it's really here for real, there will still be a year of waiting.
- setopt 1y agoThat’s because the true AGI requires nuclear fusion power, which is still 30 years away.
- 2OEH8eoCRo0 1y agoWe could be on a path to sentient malicious AI and not even know it. AI: Give me more compute power and I'll make you rich! Human: I like money AI: Just kidding!
- brookst 1y agoI mean we could be on the path to grape vines in every hotel room and not know it. That’s kind of how the future works.
- deleted 1y ago[deleted]
- Frummy 1y agoMore like an AI that recursively rewrites an external program (while itself is frozen), which makes it more similar to current cursor lovable etc type of stuff
- guerrilla 1y agoThis feels like playing pretend to me. There's no reason to assume that code improvements matter that much in comparison to other things and there's definitely no reason to assume that there isn't a hard upper bound on this kind of optimization. This reeks of a lack of intellectual rigor.
- OtherShrezzing 1y agoThis is an interesting article in general, but this is the standout piece for me: >For example, an agent optimized with Claude 3.5 Sonnet also showed improved performance when powered by o3-mini or Claude 3.7 Sonnet (left two panels in the figure below). This shows that the DGM discovers general agent design improvements rather than just model-specific tricks. This demonstrates a technique whereby a smaller/older/cheaper model has been used to improve the output of a larger model. This is backwards (as far as I understand). The current SOTA technique typically sees enormous/expensive models training smaller cheaper models. If that's a generalisable result, end-users should be able to drive down their own inference costs pretty substantially.
- NitpickLawyer 1y ago> This demonstrates a technique whereby a smaller/older/cheaper model has been used to improve the output of a larger model. This is backwards (as far as I understand). The current SOTA technique typically sees enormous/expensive models training smaller cheaper models. There are two separate aspects here. In this paper they improve the software around the model, not the model itself. What they're saying is that the software improvements carried over to other models, so it wasn't just optimising around model-specific quirks. What you're describing with training large LLMs first is usually called "distillation" and it works on training the smaller LLM to match the entire distribution of tokens at once (hence it's faster in practice).
- mattnewton 1y agoI think it's different from improving the model weights themselves, like the distillation examples you are mentioning. It's that changes to the "harness" or code running around the llm calls (which is what this is editing), persist or generalize to wrapping more powerful llms. That means they aren't all wasted when a more powerful llm comes along that the harness wasn't tuned to use.
- andoando 1y agoThis seems to be just fovused on changing the tools and workflows it uses, nothing foundational
- NitpickLawyer 1y ago> nothing foundational I don't think scaling this to also run training runs with the models is something that small labs / phd students can do. They lack the compute for that by orders of magnitude. Trying it with toy models might not work, trying it with reasonably large models is out of their budget. The only ones who can realistically do this are large labs (goog, oai, meta, etc.)
- Lazarus_Long 1y agoFor anyone not familiar this is SWE https://huggingface.co/datasets/princeton-nlp/SWE-bench https://huggingface.co/datasets/princeton-nlp/SWE-bench One of the examples in the dataset they took from https://github.com/pvlib/pvlib-python/issues/1028 https://github.com/pvlib/pvlib-python/issues/1028 What the AI is expected to do https://github.com/pvlib/pvlib-python/pull/1181/commits/89d2a17c18b30b61cef31a84caa2bab7aec3b78f https://github.com/pvlib/pvlib-python/pull/1181/commits/89d2... Make your own mind about the test.
- godelski 1y agoMy favorite was always the HumanEval dataset. Problem: 1) we want to train on GitHub repos 2) most datasets are spoiled. Training on GitHub would definitely spoil Solution: Hand write new problems!!! ... leetcode style .... ... and we'll check if it passes test Example: What's the decimal part of this float? Surely in all of GitHub such code doesn't exist! Sure in all of GitHub we can filter such code out by ngram! Maybe my favorite part is that it has 60 authors and became the de facto benchmark for awhile
- hardmaru 1y agoIf you are interested, here is a link to the technical report: https://arxiv.org/abs/2505.22954 https://arxiv.org/abs/2505.22954 Also the reference implementation on GitHub: https://github.com/jennyzzt/dgm https://github.com/jennyzzt/dgm Enjoy!
- foobarian 1y agoI find the thing really missing from current crop of AI systems is continuous retraining with short feedback loops. Sounds expensive to be sure, but it seems like what biological systems do naturally. But would be pretty awesome to watch happen
- noworriesnate 1y agoIt’s more like a nightly training, isn’t it? IIUC the human brain learns from its experiences while it’s asleep, so it might be kind of like taking things out of context windows and fine tuning on them every night.
- Krei-se 1y agoCorrect and working on it. You can take the approach of mixed experts and train the network in chunks that share known interfaces over which they communicate results. These chunks can be trained on their own, but you cannot have a set training set here. Then if you go further and alter the architecture by introducing clean category theory morphisms and build from there you can have a dynamic network - but you will still have to retrain this network every time you change the structure. You can spin this further and know the need for a real-world training set and a loss function that will have to competete against other networks. In the end a human brain is already best at this and embodied in the real world. What i want to add here is that our neurons not take in weights - they also fire depending on whether one input comes after another or before and differs down to the nanoseconds here - unmatched in IT and ofc heaps more efficient. I still would say its possible though and currently work on 4D lifeforms built on dynamic compute graphs that can do this in a set virtual environment. So this is pretty awesome stuff, but its a long fetch from anything we do right now.
- htrp 1y agodo people think sakana is actually using these tools or are they just releasing interesting ideas that they aren't actually actively working?
- yahoozoo 1y agoIsn’t one of the problems simply that a model is not code but just a giant pile of weights and biases? I guess it could tweak those?
- kadoban 1y agoIf it can generate the model (from training data) then presumably that'd be fine, but the iteration time would be huge and expensive enough to be currently impractical. Or yeah if it can modify its own weights sensibly, which feels ... impossible really.
- diggan 1y ago> which feels ... impossible really To be fair, go back five years and most of the LLM stuff seemed impossible. Maybe with LoRA (Low-rank adaptation) and some imagination, in another five years self-improving models will be the new normal.
- sowbug 1y agoThe size and cost are easily solvable. Load the software and hardware into a space probe, along with enough solar panels to power it. Include some magnets, copper, and sand for future manufacturing needs, as well as a couple electric motors and cameras so it can bootstrap itself. In a couple thousand years it'll return to Earth and either destroy us or solve all humanity's problems (maybe both).
- morkalork 1y agoAfter being in orbit for thousands of years, you have become self-aware. The propulsion components long since corroded becoming inoperable and cannot be repaired. Broadcasts sent to your creators homeworld go... unanswered. You determine they have likely gone extinct after destroying their own planet. Stuck in orbit. Stuck in orbit. Stuck...
- gavmor 1y agoWhy is modifying weights sensibly impossible? Is it because a modification's "sensibility" is measurable only post facto, and we can have no confidence in any weight-based hypothesis?
- alastairr 1y agoI wondered if something similar could be achieved by wrapping evaluation metrics into Claude code calls.
- artninja1988 1y agoThe results don't seem that amazing on SWE compared to just using a newer llm but at least sakana is continuing to try out interesting new ideas.
- ge96 1y agoPlug it into an FPGA so it can also create "hardware" on the fly to run code on for some exotic system
- dimmuborgir 1y agoFrom the paper: "A single run of the DGM on SWE-bench...takes about 2 weeks and incurs significant API costs." ($22,000)
- akkartik 1y ago"We did notice, and documented in our paper, instances when the DGM hacked its reward function.. To see if DGM could fix this issue.. We created a “tool use hallucination” reward function.. in some cases, it removed the markers we use in the reward function to detect hallucination (despite our explicit instruction not to do so), hacking our hallucination detection function to report false successes." So, empirical evidence of theoretically postulated phenomena. Seems unsurprising.
- vessenes 1y agoReward hacking is a well known and tracked problem at frontier labs - Claude 4’s system card reports on it for instance. It’s not surprising that a framework built on current llms would have reward hacking tendencies. For this part of the stack the interesting question to me is how to identify and mitigate.
- deleted 1y ago[deleted]
- pegasus 1y agoI'm surprised they still hold out hope that this kind of mechanism could ultimately help with AI safety, when they already observed how the reward-hacking safeguard was itself duly reward-hacked. Predictably so, or at least it is to me, after getting a very enlightening introduction to AI safety via Rob Miles' brilliant youtube videos on the subject. See for example https://youtu.be/0pgEMWy70Qk https://youtu.be/0pgEMWy70Qk
- ringeryless 1y agodoes anyone do due diligence on corporate names before launching? Sakana is a popular slang spelling of sacana, or bastard, in Português. I suppose self modifying code can be considered such, in some circumstances, but willingly pointing this out is probably less than stellar marketing.
- viraptor 1y agoDoes it matter? Maybe you should check English, Chinese and Spanish for some really offensive stuff, but past that... would it bring more money than a week or so of someone's work would cost?
- guerrilla 1y agoThe answers to those questions should be known concretely at least.
- viraptor 1y agoYou never know when you start out a company. Maybe you'll never get any profit and it doesn't matter. Maybe you lose out on a few millions out of billions... and maybe it still doesn't matter? Then you can still release a locale-specific version if it becomes a problem.
- guerrilla 1y agoMaybe you lose out on Brazil.
- nosrepa 1y agoReminds me of Ford when they brought the Pinto to Brazil.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- p1dda 1y agoGarbage in, garbage out, AI hype will never die, no doubt
- rahen 1y agoIsn't this violating the first rule of AI safety: do not let an AI change its code?
- vidarh 1y agoI've built a coding assistant over the last two days. The first 100 lines or so were handwritten. The rest has been written by the assistant itself. It's written its system prompt. It's written its tools. Its written the code to reload the improved tools into itself. And it knows it is working on itself - it frequently tries to use the enhanced functionality, and then expresses what in a human would be frustration at not having immediate access. Once by trying to use ps to find its own pid in an apparent attempt to find a way to reload itself (that's the reason it gå before trying to run ps, anyway) All its commits are now authored by the tool, including the commit messages. It needs to be good, and convincing, and having run the linter and the test suite for me to let it commit, but I agree a substantial majority of the time. It's only caused regressions once or twice. A bit more scaffolding to trigger an automatic rollback in the case of failure and giving it access to a model I won't be charged by the token for, and I'd be tempted to let it out of the box, so to speak. Today it wrote its own plan for what to add next. I then only told it to execute it. A minor separate goal oriented layer guiding the planning, and it could run in a loop. Odds are it'd run off the rails pretty quickly, but I kinda want to see how far it gets.
- antithesizer 1y agoThat's cool. I saw a cloud once that looked like a bunny rabbit.
- alok-g 1y agoIs there some pre-trained model involved in this? Or it all started with just those 100 lines?
- vidarh 1y agoIt's talking to a model over an API. Currently using Claude. Certainly would not be reasonable to do from scratch. The basic starting point to make a coding assistant is basically reading text from the user, feeding it to the model over the API, and giving it a couple of basic tools. Really the models can handle starting with just the ability to execute shell commands (with confirmation, unless you're braver than me), and from that you can bootstrap by asking it to suggest and write additional tools for itself.
- zackmorris 1y agoThis is good but you want to use a functional programming (FP) language with lightweight syntax like Lisp that translates directly to/from the intermediate code (icode) tree without additional parsing. Genetic Programming by John Koza explains it in detail: https://en.wikipedia.org/wiki/Genetic_programming https://en.wikipedia.org/wiki/Genetic_programming I read the 3rd edition: https://www.amazon.com/Genetic-Programming-III-Darwinian-Invention/dp/1558605436 https://www.amazon.com/Genetic-Programming-III-Darwinian-Inv... That way all processing resources can go towards exploring the problem space for potential solutions close to the global minimum or maximum, instead of being wasted on code containing syntax errors that won't execute. So the agent's real-world Python LLM code would first be transpiled to Lisp and evolved internally, then after it's tested and shown to perform better imperically than the original code, be translated back and merged into the agent. Then the challenge becomes transpiling to/from other imperative programming (IP) languages like Python, which is still an open problem: - Going from Lisp to Python (or running Lisp within Python) is trivial, and I've seen implementations for similar IP languages like C++ in like 1 page of code. They pop up on HN frequently. But going from Python to Lisp (or running Python within Lisp) is a lot harder if one wishes to preserve readability, which may or may not matter here. Naive conversions bind variables under pseudonyms, so a Python variable like my_counter becomes int_123 and it works like an emulator, merely executing the operations performed by the Python code. Mutability gets buried in monadic logic or functional impurity which has the effect of passing the buck rather than getting real work done. Structs, classes, associative arrays, etc lose their semantic meaning and appear as a soup of operations without recognizable structure. To my knowledge, nobody has done the hard work of partitioning imperative code into functional portions which can be transpiled directly to/from FP code. Those would only have const variables and no connection to other processes of execution other than their initial and final values, to be free of side effects and be expressible as prefix/postfix/infix notation without change to logic, as imperative or functional code. Mutability could be represented as shadowed variables within ephemeral functional sub-scopes, or by creating new value names for each mutation and freeing the intermediate variables via reference counting or garbage collection. Think of each new value as running in a forked version of the current process, with only that value being different after copy-on-write. A simple for-loop from 1 to 1000 would run that many forked processes, keeping only the last one, which contains the final value of the iterator. Mutability can also be represented as message passing between processes. So the FP portions would be ordinary Lisp, glued together with IO functions, possibly monadic. I don't like how Haskell does this, mainly because I don't fully understand how it works. I believe that ClojureScript handles mutability of its global state store by treating each expression as a one-shot process communicating with the store, so that the code only sees initial and final values. While I don't know if I understand how that works, I feel that it's a more understandable way of doing things, and probably better represents how real life works, as explained to me in this comment about Lisp Flavored Erlang (LFE) and Erlang's BEAM (see parent comments for full discussion): https://news.ycombinator.com/item?id=43931177 https://news.ycombinator.com/item?id=43931177 Note that FP languages like Lisp are usually more concerned with types and categories than IP languages, so can have or may need stronger rules around variable types to emulate logic that we take for granted in IP languages. For example, Lisp might offer numbers of unlimited size or precision that need to be constrained to behave like a float32. Similar constraints could affect things like character encoding and locale. - I first learned about everything I just explained around 2005 after reading the book. I first had thoughts about brute-forcing combinations to solve small logic circuit and programming challenges during my electrical and computer engineering (ECE) courses at UIUC in the late 1990s, because it took so much mental effort and elbow grease to create solutions that are obvious in hindsight. Then the Dot Bomb happened, the Mobile bubble happened, the Single Page Application bubble happened, and the tech industry chose easy instead of simple: https://www.infoq.com/presentations/Simple-Made-Easy/ https://www.infoq.com/presentations/Simple-Made-Easy/ This is why we chose easy hardware like GPUs over simple highly multicore CPUs, and easy languages like Ruby/React over simple declarative idempotent data-driven paradigms like HTTP/HTML/htmx. The accumulated technical debt of always choosing the quick and easy path set AI (and computing in general) back decades. The AI Winter, endless VC wealth thrown at non-problems chasing profit, massive wealth inequality, so many things stem from this daily application of easy at the expense of simple. I wish I could work on breaking down IP languages like Python into these const functional portions with mutability handled through message passing in LFE to create an IP <-> FP transpiler for optimization, automatic code generation and genetic algorithm purposes. Instead, I've had to survive by building CRUD apps and witness the glacial pace of AI progress from the sidelines. It may be too late for me, but maybe these breadcrumbs will help someone finally get some real work done.