6 ms·
AlphaEvolve: Gemini-powered coding agent scaling impact across fields
- baq 5mo agoRSI is here on the hardware level and on software level. Sprinkle with a couple algorithmic breakthroughs and results are nigh unimaginable.
- marcus_ai 5mo ago[flagged]
- kadam2576 5mo ago[flagged]
- stalfie 5mo agoWell, if the evaluation infrastructure is something humans could have had access to before, and that the agents key "skill" is just that it's a more patient and scalable worker, I would still argue that this "comes from the agent". Humans get bored, inpatient, or run out of time, and so often give up in what they perceive to be a decent "local minima". Early verification harnesses using gpt-4 for optimizing robot reward functions succeeded quite well on the fact that the LLM just kept going (link below). As long as it is too boring for a human to use the same evaluation infrastructure, this is still an agent skill. https://arxiv.org/abs/2310.12931 https://arxiv.org/abs/2310.12931
- deleted 5mo ago[deleted]
- pingou 5mo agoAI improving itself (or at least the architecture it runs on), the singularity is near as they say. Do we have other examples of AI being used to improve the LLMs, apart for the creation of synthetic data and the testing of the models?
- mkw5053 5mo agoI feel like the most viral lately is https://github.com/karpathy/autoresearch https://github.com/karpathy/autoresearch
- deleted 5mo ago[deleted]
- NitpickLawyer 5mo ago> Do we have other examples of AI being used to improve the LLMs Yes, last year when they revealed AlphaEvolve they used a previous gemini model to improve kernels that were used in training this gen models, netting them a 1% faster training run. Not much, but still.
- lewtun 5mo agoShameless plug: https://huggingface.co/spaces/smolagents/ml-intern https://huggingface.co/spaces/smolagents/ml-intern It’s a simple harness around Opus, but with tight integration to Hugging Face infra, so the agent can read papers, test code and launch experiments
- westurner 5mo agoWhat are the benchmarks for this, in terms of costs of computation and error; cost to converge? Re: hyperparameter tuning and autoresearch: https://news.ycombinator.com/item?id=47444581 https://news.ycombinator.com/item?id=47444581 Parameter-free LLMs would be cool
- dinfinity 5mo ago> AI improving itself This is the thing to look for in 2027, imho. All the big AI labs have big projects working on research agents, also specifically into improving AI (duh) and I expect a lot of that to get out of the experimental phases this year. Next year they actually get to do a lot of work and I think we will see the first big effective architectural change co-invented by AI.
- maxothex 5mo agoWhat I'm most curious about is how this translates to messy, real-world codebases without well-defined metrics. Most production software isn't chip design or kernel optimization - it's business logic with unclear success criteria. The infrastructure story is impressive, but I'd love to see how they handle domains where the evaluation function itself is ambiguous.
- alecco 5mo agoAre Googlers themselves happy using Gemini coding agent instead of Claude Code or Codex? (no snark, I'm really asking)
- carbocation 5mo agoLast month, Steve Yegge suggested that they are not: https://xcancel.com/Steve_Yegge/status/2043747998740689171 https://xcancel.com/Steve_Yegge/status/2043747998740689171
- PunchTornado 5mo agoThis couldn't be further from the truth
- NitpickLawyer 5mo ago> He says the problem is that they can't use Claude Code because it's the enemy, and Gemini has never been good enough to capture people's workflows like Claude has, so basically agentic coding just never really took off inside Google. They're all just plodding along, completely oblivious to what's happening out there right now. This is a bunch of gabagoo. Wrong on so many layers, it's not even worth reading further. a) goog has agentic coding in both antigravity & cli forms. While it is not at the level of cc + opus, it's still decent. b) goog has their own versions of models trained on internal code c) goog has claude in vertex, and most definitely can set it up in secure zones (like they can for their clients) so they'd be able to use claude (at cost) within their own projects.
- aleksiy123 5mo agoAgreed, however imo there is def some problems unique to Google which is making the internal experience less than ideal. Hoping they can figure it out sooner rather than later.
- jatora 5mo agoAntigravity's workflow is so slow and buggy with permissions, it is unuseable compared to cc/codex. the only part that is nice is that it allows usage of Opus. Gemini CLI is an absolute joke. I dont know if its the harness, or the Gemini models' poor instruction following, but more than 50% of my sessions gemini turns insane and ends up in thought loops. In another 25% it does far more updating than I asked or is reasonable. This is why next to no-one talks about either of them. Does Antigravity's agent manager even work yet without crashing and showing zero conversations? In typical google fashion they released AG and appear to have then set it on autopilot with a skeleton crew of less than 1 developer. Clear issues have not been fixed since day 1. Some settings just do not work. Permissions are not respected.
- brkn 5mo agoI would be interested to see how exactly the agent helped. How was it used, where did it lead to the given improvement and in how far would it have taken a human to come to the same solution.
- j2kun 5mo agoThe blog post has many links to papers and preprints discussing this exact question.
- Lt_Riza_Hawkeye 5mo agoThe CANOS arxiv link says absolutely nothing about AlphaEvolve, Gemini, or LLMs. It seems to use purely traditional ML models. If AE did in fact write a quick script to test different configurations in order to optimize the results, they don't seem to have bothered to write about it. I can't read the Nature paper about DeepConsensus, but from the summary, it doesn't really explain what role AE had in improving DC. It would be nice to be able to read about what role it actually played, and whether it used traditional or novel methods of performing it
- armanj 5mo agoseems like `karpathy/autoresearch` on steroids
- momojo 5mo agoThis reminds me of Antirez's "Don't fall into the anti-AI hype" [0] In a sentence: These foundation models are really good at optimizing these extremely high level, extremely well defined problem spaces (ie multiply matrices faster). In Antirez's case, it's "make Redis faster". There have been two reactions: "Oh it would never work for me" and "I have seen months of my life accomplished in an hour", and I think they're both right. I think we should be excited for Antirez, (who has since been popping off [1]), and I think the rest of us should rest easy knowing that LLM's can't (and maybe were never meant to) tackle the tacit-knowledge-filled, human-system-centric, ambiguously-defined-problem-space jobs most mortals work. [0] https://antirez.com/news/158 https://antirez.com/news/158 [1] https://antirez.com/news/164 https://antirez.com/news/164
- dinfinity 5mo ago> I think the rest of us should rest easy knowing that LLM's can't [...] What if (when?) (AI-assisted) research moves AI beyond LLMs? Do you think that can't happen?
- kubb 5mo agoNot in the next decade. Won't get funded.
- dinfinity 5mo agoPrivate investment in the US has grown from 100 billion in 2024 to almost 300 billion USD in 2025 [0]. Add public investments worldwide and private investments in at least China and Europe. I'm pretty sure money is not going to be the blocker. [0] https://hai.stanford.edu/ai-index/2026-ai-index-report https://hai.stanford.edu/ai-index/2026-ai-index-report
- kubb 5mo agoThe money will go to LLMs.
- deleted 5mo ago[deleted]
- dandaka 5mo agoHow many times we have to hear again about Erdös problems? :) It sounds like a great achievement for humanity at first, but after a while they keep coming back!
- j2kun 5mo agoThere are only some 700 open Erdos problems left, so when they're all solved you can finally rest.
- pilooch 5mo agoAlphaEvolve couples map-elites with LLMs. It's an key step in machine learning, in the vein of DQN for reinforcement learning. AE brings diversity from the genetic algorithms community to large scale optmized deep learning and RL models. It is a mandatory step for moving forward. The approach is clean and simple, while generic. The only caveats is the per optimization problem definition of the map élites dimensions. But surely, this will get tackled somehow over the next few years. If you don't know about map-elites, go look up Jean-Baptiste Mouret' s work and talks, it's both very interesting and universal.
- stijntonk 5mo agoI wish that Google would focus on bringing their Gemini 3.x models to GA, and provide enough capacity such that one not constantly has to fight with 429 errors. It often feels like they do not want me to develop applications for corporate clients using their Vertex API. It is just such a shame, given that their models were so great for document analysis etc.
- VladVladikoff 5mo agoAre you doing it on a free plan? I noticed they serve way more 429s on the free plan.
- stijntonk 5mo agoNo, for clients we use paid Vertex AI accounts. We often need to host workloads in an EU region, which rules out “global” models (and probably better capacity). In the past, we used a wrapper that round-robined across multiple projects to get enough quota. Luckily, many of our workloads are workflow-style tasks, so we can simply keep retrying on 429s. Fun fact: for one of their services, I think it was Stitch, I noticed that my paid key kept hitting quota, while the free worked fine. That blew my mind.
- BoxedEmpathy 5mo agoI've been seeing the same in my product; 429s in vertex. We generally avoid any Google AI for the most part because it's so unreliable.
- stijntonk 5mo agoWhat is your use case, and what do you use instead? For analysis of large complex docs I find it pretty great, and fair priced.
- AndrewKemendo 5mo agoFrom the comments it seems that this community (mostly career software people) is starting to move into a new phase of grief about the median software engineer losing their hoped for permanent place in society. -2021-2024 was Denial -2024-2025 was Anger and Bargaining -2026 seems to be some combo of anger, bargaining and acceptance depending mostly on your class/age
- artninja1988 5mo agoI think we are still in the denial phase.
- nmitchko 5mo agoA fantastically simple solution to improving algorithms, I wish I had this years ago in activation engineering: https://blog.n.ichol.ai/llm-activation-engineering-an-easy-foray https://blog.n.ichol.ai/llm-activation-engineering-an-easy-f... How do I access AlphaEvolve?
- Yokohiii 5mo agoThis is just a flex post. Be a billion dollar company or get out.
- charleshn 5mo agoThey'll likely make it available at some point, but for now one can use OpenEvolve [0] which is not quite as good but should be a good start to use the same LLM-driven evolutionary framework. [0] https://github.com/algorithmicsuperintelligence/openevolve https://github.com/algorithmicsuperintelligence/openevolve
- arian_ 5mo agoWe went from 'AI will replace programmers' to 'AI will help programmers' to 'AI writes code while other AI reviews it' in about 18 months. At this rate the humans are just providing the electricity.
- deleted 5mo ago[deleted]
- svieira 5mo ago> In advertising and marketing, WPP used AlphaEvolve to refine AI model components, navigating complex, high-dimensional campaign data and achieving 10% accuracy gains over their competitive manual model optimizations. Ah good, we're getting closer and closer to Venus, Inc. every day. /s
- agluszak 5mo agoMeanwhile Gemini CLI has been broken for months! https://github.com/google-gemini/gemini-cli/issues/22141 https://github.com/google-gemini/gemini-cli/issues/22141
- guybedo 5mo agoand yet Gemini still can't code
- kaueg 5mo agoWelcome to HN @berlianta; TIL green username === new user in HN; Stories posted by new users are called noobstories [1]; [1]: https://news.ycombinator.com/noobstories https://news.ycombinator.com/noobstories
- berlianta 5mo agoNo need for a welcome message, just stick to the topic. Thank you
- ainch 5mo agoThe AI CEOs love to pontificate about AI curing cancer, but it seems like DeepMind is the only one actively working on these research problems, while OpenAI/Anthropic largely chase enterprise/coding revenue.
- igorpcosta 5mo agoThere's not a lot of opportunities in this space yet. This is the closest we can get to High degree solver kinda of problem. There are only 3 companies doing this to date: Google, Sakana AI and Autohand AI.
- cpard 5mo agoAll the *Evolve publications have very impressive results but from the time I’ve spent on the information published I feel that the attention goes to the LLMs and the AI side of things, although the outcomes reported are in almost all cases the result of very well designed environments for both the LLM and the evolutionary algorithm to work well. This paper here is a great example of that and it’s worth a reading. Magellan: Autonomous Discovery of Novel Compiler Optimization Heuristics with AlphaEvolve https://arxiv.org/abs/2601.21096 https://arxiv.org/abs/2601.21096
- sjhalani7 5mo agoThis is crazy- the fact that it is helping with stuff like quantum too is huge!
- zkmon 5mo agoAn issue I have been noticing with claude is, for simple tasks, it gives extremely bloated code and artifacts, which sometime does not even work. Gemini balances it quite well, by giving a working solution with the exact amount of code and minimal complexity, that is easier to manage. The only thing I go to Claude these days, is for front-end code (HTML). Here also, it gives too much of CSS code (60% of the file size), but I'm OK with that as it gives a bit of polished look, though heavy on file size.