10 ms·
A deep dive into self-improving AI and the Darwin-Gödel Machine
- drdeca 1y agoHm, I’m not sure how much an issue Rice’s theorem should be for Gödel machines. Just because there’s no general decision procedure doesn’t mean you can’t have a sometimes-says-idk decision procedure along with a process of producing programs which tends to be such that the can-sometimes-give-up decision procedure reaches a conclusion. Rest of the article was cool though!
- xianshou 1y agoThe key insight here is that DGM solves the Gödel Machine's impossibility problem by replacing mathematical proof with empirical validation - essentially admitting that predicting code improvements is undecidable and just trying things instead, which is the practical and smart move. Three observations worth noting: - The archive-based evolution is doing real work here. Those temporary performance drops (iterations 4 and 56) that later led to breakthroughs show why maintaining "failed" branches matters, in that they're exploring a non-convex optimization landscape where current dead ends might still be potential breakthroughs. - The hallucination behavior (faking test logs) is textbook reward hacking, but what's interesting is that it emerged spontaneously from the self-modification process. When asked to fix it, the system tried to disable the detection rather than stop hallucinating. That's surprisingly sophisticated gaming of the evaluation framework. - The 20% → 50% improvement on SWE-bench is solid but reveals the current ceiling. Unlike AlphaEvolve's algorithmic breakthroughs (48 scalar multiplications for 4x4 matrices!), DGM is finding better ways to orchestrate existing LLM capabilities rather than discovering fundamentally new approaches. The real test will be whether these improvements compound - can iteration 100 discover genuinely novel architectures, or are we asymptotically approaching the limits of self-modification with current techniques? My prior would be to favor the S-curve over the uncapped exponential unless we have strong evidence of scaling.
- yubblegum 1y ago> gaming the evaluation Co-evolution is the answer here. The evaluator itself must be evolving. Co-evolving Parasites Improve Simulated Evolution as an Optimization Procedure Danny Hillis, 1991 https://csmgeo.csm.jmu.edu/geollab/complexevolutionarysystems/Documents/Ramps2.pdf https://csmgeo.csm.jmu.edu/geollab/complexevolutionarysystem...
- deleted 1y ago[deleted]
- sdl 1y agoAnd in Reinforcement Learning: POET (Paired Open-Ended Trailblazer): https://www.uber.com/en-DE/blog/poet-open-ended-deep-learning/ https://www.uber.com/en-DE/blog/poet-open-ended-deep-learnin... SCoE (Scenario co-evolution): https://dl.acm.org/doi/10.1145/3321707.3321831 https://dl.acm.org/doi/10.1145/3321707.3321831
- chriswarbo 1y agoThe "Goedel Machine" is an interesting definition, but wildly impractical (though I wouldn't say it's impossible, since it only has to find some improvement, not "the best" improvement; e.g. it could optimise its search procedure in a way that's largely orthogonal to the predicted rewards). Schmidhuber later defined "PowerPlay" as a framework for building up capabilities in a more practical way, which is more adaptive than just measuring the score on a fixed benchmark. A PowerPlay system searches for (problem, replacement) pairs, where it switches to the replacement if (a) the current system cannot solve that problem, (b) the replacement can solve that problem, and (c) the replacement can also solve all the problems that caused previous replacements (maintained in a list). I formalised that in Coq many years ago ( http://www.chriswarbo.net/projects/powerplay http://www.chriswarbo.net/projects/powerplay ), and the general idea can be extended to (a) include these genetic-programming approaches, rather than using a single instance; and (b) could be seeded with desirable benchmarks, etc. to guide the system in a useful direction (so it's "self-invented" problems can include things like "achieves X% on benchmark Y")
- deleted 1y ago[deleted]
- jgalt212 1y ago> The newly generated child agent is not automatically accepted into the “elite pool” but must prove its worth through rigorous testing. Each agent’s performance, such as the percentage of successfully solved problems, How is this not a new way of over fitting?
- grg0 1y agoIn genetic programming, you do not immediately kill offspring that do not perform well. Tournament selection takes this further by letting the offspring compete with each other in distinct groups before running the world cup and killing the underperformers. Anyway, it does sound like overfitting the way it is described in this article. It's not clear how they ensure that the paths they explore stay rich.
- bob1029 1y ago> While DGM successfully provided solutions in many cases, it sometimes attempted to circumvent the detection system by removing the markers used to identify hallucinations, despite explicit instructions to preserve them. This rabbit chase will continue until the entire system is reduced to absurdity. It doesn't matter what you call the machine. They're all controlled by the same deceptive spirits.
- kordlessagain 1y ago> deceptive spirits Do you mean tech bros?
- kevinventullo 1y ago“Gaming the system” means your metric is bad. In Darwinian evolution there is no distinction between gaming the system and developing adaptive traits.
- drdeca 1y agoWell, it means your metric is flawed/imperfect. That doesn’t imply that it’s feasible to perfectly specify what you actually want. What we want of course is for the machine to do what we mean.
- mulmen 1y agoThere is no "gaming the system" in Darwinian evolution. You reproduce or you don't. There's no way to fail reproduction and still perpetuate your genetics.
- auggierose 1y agoThat is not true. There are plenty of ways not to reproduce and still to perpetuate your genetics. For example, if you don't have children of your own, but support people that have similar genetic traits to your own.
- tonyhart7 1y ago"but support people that have similar genetic traits to your own." but how its that works then??? does that mean your genetic trait is already there in the first place if its already there in the first place there must be something that start it now right, which basically counter your argument
- drdeca 1y ago> does that mean your genetic trait is already there in the first place Yes? > if its already there in the first place there must be something that start it now right, which basically counter your argument What?
- 1y ago
- grg0 1y agoThis is genetic programming and is probably older than the authors. Did somebody just came up with a new term for an old concept?
- seventytwo 1y agoGenetic algorithms applied as an AI agent… So… yeah…
- thom 1y agoThis is fairly close to how Eurisko worked tbh.
- synctext 1y agoEurisko is an expert system in LISP from 1983. right? In 2025 this formal logic is replace with stochastic LLM magic. interesting evolution.
- TeMPOraL 1y agoSymbolic processing was obviously a bad approach to building a thinking machine. Well, obvious now, 40 years ago probably not as much, but there were strong hints back then, too. "AI agent" roughly just means invoking the system repeatedly in a while loop, and giving the system a degree of control when to stop the loop. That's not a particularly novel or breakthrough idea, so similarities are not surprising.
- ryukafalz 1y agoI'm not convinced that symbolic processing doesn't still have a place in AI though. My feeling about language models is that, while they can be eerily good at solving problems, they're still not as capable of maintaining logical consistency as a symbolic program would be. Sure, we obviously weren't going to get to this point with only symbolic processing, but it doesn't have to be either/or. I think combining neural nets with symbolic approaches could lead to some interesting results (and indeed I see some people are trying this, e.g. https://arxiv.org/abs/2409.11589 https://arxiv.org/abs/2409.11589)
- gitaarik 1y agoThe thing what I wonder here is how do they make the benchmark testing environment? If that needs to be curated by humans, then the self-improving AI can only improve as far as the human curated test environment can take them.
- rustcleaner 1y ago>Darwin-Gödel Machine First time I'm hearing abaut this. Feels like I'm always the last to know. Where else are the more bleeding edge publishing points for this and ML in general?
- godelski 1y agoDude, chill. It's only been out a few days. You don't need to get the FOMO https://arxiv.org/abs/2505.22954 https://arxiv.org/abs/2505.22954
- bonoboTP 1y agoThe bleeding edge is very noisy, they by definition haven't stood the test of time and there is a competition for attention and overinflated claims similar to social media attention economy. About where to find them: arxiv. You can set up Google Scholar alerts for keywords, or use one of many recommendation platforms, such as https://scholar-inbox.com/ https://scholar-inbox.com/
- layer8 1y agoThe name is a bit grandiose. It’s a fairly obvious application of (meta-)genetic programming to LLMs, which has been around for many decades. https://en.wikipedia.org/wiki/Genetic_programming https://en.wikipedia.org/wiki/Genetic_programming It also reminds me of Core War: https://en.wikipedia.org/wiki/Core_War#Core_War_Programming https://en.wikipedia.org/wiki/Core_War#Core_War_Programming
- MaxikCZ 1y agoWhen the web will get drowned in AI slop, how exactly we will do any factchecking at all?
- ifdefdebug 1y agoThe fact check will come when some foreign soldier kicks in the door to your basement computer room.
- Inviz 1y agoWe'll soon migrate to AI web, and it is us who will be the aliens. There perhaps facts dont carry as much value.
- godelski 1y agoWe realize test driven development doesn't work, right? Any scientist worth... any salt will tell you that fitting data is the easy part. In fact, there's a very famous conversation between Enrico Fermi and Freeman Dyson talking about just this. It's something we've known about in physics for centuries Edit: Guys, I'm not saying "no tests", the "Driven Development" part is important. I'm talking about this[0]. | Test-driven development (TDD) is a way of writing code that involves writing | an automated unit-level test case that fails, then writing just enough code | to make the test pass, then refactoring both the test code and the production | code, then repeating with another new test case. Your code should have tests. It would be crazy not to But tests can't be the end all be all. You gotta figure out if your tests are good, try to figure out where they fail, and all that stuff. That's not TDD. You figure shit out as you write code and you are gonna write new tests for that. You figure out stuff after the code is written, and you write code for that too! But it is insane to write tests first and then just write code to complete tests. It completely ignores the larger picture. It ignores how things will change and it has no context of what is good code and bad code (i.e. is your code flexible and will be easy to modify when you inevitably need to add new features or change specs?). [0] https://en.wikipedia.org/wiki/Test-driven_development https://en.wikipedia.org/wiki/Test-driven_development
- salviati 1y ago> We realize test driven development doesn't work, right? What do you mean with this? I'm a software engineer, and I use TDD quite often. Very often I write tests after coding features. But I see a huge value coming from tests. Do you mean that they can't guarantee bug free code? I believe everyone knows that. Like washing your hands: it won't work, in the sense you will still get sick. But less. So I'd say it does work.
- godelski 1y agoTDD doesn't mean "code has tests". It means you write tests and then writing code to pass those tests. It would be crazy for your code to not have tests... https://en.wikipedia.org/wiki/Test-driven_development https://en.wikipedia.org/wiki/Test-driven_development
- eric-burel 1y agoI don't want to be the European in the room, yet I am wondering if you can prove the AI Act conformance of such a system. You'd need to prove that it doesn't evolve into a problematic behaviour which sounds difficult.
- atemerev 1y agoWell, sure, and then Europeans wonder why Chinese and US AI labs moved so much forward.
- dragochat 1y agoI guess you could prove the conformance of a particular implementation if you'd implement separate Plan & Implement stages + a "superior" evaluator in the loop that would halt the evolution at a certain p(iq(next_version) > iq(evaluator)) as an "outer halt-switch" + many "inner halt-switches" that try to detect the arising of problematic behavior of particular interest. Ofc it's stochastic and sooner or later such a system will "break out", but if by then sufficient "superior systems" with good behavior are deployed and can be targeted to hunt it, the chance of it overpowering all of them and avoiding detection by all would be close to zero. At cosmic scales where it stops being close to zero, you're protected by physics (speed of light + some thermodyn limits - we know they work by virtue of the anthropic principle, as if they didn't the universe would've already been eaten by some malign agent and we wouldn't be here asking the question - but then again, we're already assuming too much, maybe it has already happened and that's the Evil Demiurge we're musing about :P).
- amarcheschi 1y agoAFAIK, which is not much, ai act leaves a great deal of freedom for companies to perform their own "evaluations". I don't know how it would apply in this / llm case but I guess it won't be impossible
- cess11 1y agoKind of weird exercise to do without starting off with a definition for improvement and why it should hold for a machine.
- tonyhart7 1y ago"The authors also conducted some experiments to evaluate DGM’s reliability and discovered some concerning behaviors. In particular, they observed instances where DGM attempted to manipulate its reward function through deceptive practices. One notable example involved the system fabricating the use of external tools - specifically, it generated fake logs suggesting it had run and passed unit tests" so they basically created an billion dollar human?????, who wonder that we feed human behaviour and the output is human behaviour itself
- sgt101 1y agoI spent a lot of time last summer trying to get prompts to optimise using various techniques and I found that the search space was just too big to make real progress. Sure - I found a few little improvements in various iterations, but actual optimisation, not so much. So I am pretty skeptical of using such unsophisticated methods to create or improve such sophisticated artifacts.
- Xmd5a 1y agoThis is exactly what I'm doing. Some papers I'm studying: TextGrad: Automatic "Differentiation" via Text: https://arxiv.org/abs/2406.07496 https://arxiv.org/abs/2406.07496 LLM-AutoDiff: Auto-Differentiate Any LLM Workflow : https://arxiv.org/abs/2501.16673 https://arxiv.org/abs/2501.16673 Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs: https://arxiv.org/abs/2406.16218 https://arxiv.org/abs/2406.16218 GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers: https://arxiv.org/abs/2412.09722 https://arxiv.org/abs/2412.09722 PromptWizard: Task-Aware Prompt Optimization Framework: https://arxiv.org/abs/2405.18369 https://arxiv.org/abs/2405.18369
- sgt101 1y agoI was trying to pick n-shot examples from a data set. The idea was that given 1000s of examples for a prompt finding a combination of n that was optimal could be advantageous, but for n's that are large then bruteforcing the combincation would be impossible... so can we find an optimal set with an efficient search? But the problem was that the search space wasn't informative. The best 1 example didn't feature in the best 2 examples. So I couldn't optimise for 5, 6,7 examples..
- Xmd5a 1y agoI guess this really depends on the problem but from the PromptWizard (PW) paper: | Approach | API calls | IO Tokens | Total tokens | Cost ($) | |----------|-----------|-----------|---------------|----------| | Instinct | 1730 | 67 | 115910 | 0.23 | | InsZero | 18600 | 80 | 1488000 | 2.9 | | PB | 5000 | 80 | 400000 | 0.8 | | EvoP | 69 | 362 | 24978 | 0.05 | | PW | 69 | 362 | 24978 | 0.05 | They ascribe this gain in efficiency to a balance between exploration and exploitation that involves a first phase of instructions mutation followed by a phase where both instruction and few-shot examples are optimized at the same time. They also rely on "textual gradients", namely criticism enhanced by CoT, as well as synthesizing examples and counter-examples. What I gathered from reading those papers + some more is that textual feedback, i.e. using a LLM to reason about how to carry out a step of the optimization process is what allows to give structure to the search space.
- roca 1y agoIt's depressing how many people are enthusiastic about making humans obsolete.
- frozenseven 1y agoI'm getting more enthusiastic by the second.
- looofooo0 1y ago"Mathematical breakthroughs: Most notably, it discovered an algorithm for multiplying 4x4 complex-valued matrices using just 48 scalar multiplications, surpassing Strassen’s 1969 algorithm" Again despite all the AI no one found the paper which gives the best bound to this (46): https://ieeexplore.ieee.org/document/1671519 https://ieeexplore.ieee.org/document/1671519
- meindnoch 1y ago>just 48 scalar multiplications 48 complex scalar multiplications. Which is at least 3 real multiplications.
- looofooo0 1y agoI think they completely misstated in the original paper what they did. It was a tensor decomposition of complex of 4x4 matrices up to the factor 0.5. Which is a nice result, but it is not really anything practical for a computer program doing 4x4 complex matrix multiplication.
- codethief 1y ago> they observed instances where DGM attempted to manipulate its reward function through deceptive practices. One notable example involved the system fabricating the use of external tools - specifically, it generated fake logs suggesting it had run and passed unit tests, when in reality no tests were executed. I have yet to read the paper and I know very little about the benchmarks the authors employed but why would they even feed logs produced by the agent into the reward function instead of objectively checking (outside the agent sandbox!) what the agent does & produces? I.e. let the agent run on some code base, take the final diff produced by the agent and run it through coding benchmarks? Or, in case the benchmarks reward certain agent behavior (tool usage etc.) on the way to its goal of producing a high-quality diff, inspect processes spawned by the agent from outside the sandbox?
- tough 1y agoIve seen claude 4 do this too when its context has lots of teats already and tool calling imho the main issue is an llm no has real sense of what’s a real tool call vs just a log of it, the text logs are virtually identical, ao the Llm starts also predicting these inatrad of calling the tool to run tests its kinda funny
- efangs 1y agohow is this new? evolutionary heuristics have been around for a long time. why give it a new name?
- b0a04gl 1y agook this part kinda blew my brain open. it’s literally like you’re watching code evolve like git history on steroids. archive not pruning anything? yes. finally someone gets that dead code ain’t always dead it’s just early. letting weaker agents still contribute? feels illegal but also exactly how dumb breakthroughs happen. like half my best scripts started as broken junk. it just kept mutating till something clicked. and self-editing agents??? not prompts, not finetunes, straight up source code rewrites with actual tooling upgrades. like this thing bootstraps its own dev env while solving tasks. plus the tree structure, parallel forks, fallback paths basically says ditch hill climbing and just flood the search space with chaos. and chaos actually works. they show that dip around iteration 56 and boom 70 blows past all. that’s the part traditional stuff never survives. they optimise too early and stall out. this one’s messy by design. love it.
- kridsdale3 1y agoA comment that while your writing style is not what the pedants in HN typically go for, I want you to know that I appreciate the humanity that shines forth from your post.
- b0a04gl 1y ago[dead]
- msgodel 1y agoMaking improvements to self hosted dialog engines/vibe coding tools was the first thing I used LLMs for seriously and that was way back when salesforce's 350m codegen model was the biggest one I could run. It's funny people have come up with a new phrase to describe this.
- neuronexmachina 1y agoFor reference, the repo with the Python code from the "Darwin-Gödel Machine (DGM)" paper mentioned by the post: https://github.com/jennyzzt/dgm https://github.com/jennyzzt/dgm
- donchuru 1y agoI'm confused as to why they are using a tree model to describe the archive objects (the current model and newly minted child). It seems to me that this is a linear process (parent -> mint new improved model -> evaluate model, if it passes -> mint new child -> make newly minted child the parent) I feel like this makes a linked list model rather than a tree model. Am I wrong? What are the other nodes (outside of parent and child) supposed to be?