9 ms·
Prime Agent: A self-improving RLM agent
- deleted 2mo ago[deleted]
- woah 2mo agoMight actually try this
- riddlemethat 2mo agoI built one of these RLM harnesses and a local MCP server along with logging, memories, and project rules based on directories. It worked great for a while but the foundational models have largely caught up to the point where they don't need this harness anymore. At least for my use cases. I can basically just store context in .md in the directories we work out of together and accomplish what I need.
- oofbey 2mo agoThe core idea of the RLM paper is to make a regular LLM act more like a coding agent - offload context to something external that needs to be explicitly queried instead of filling up valuable context. The "recursion" part of the paper really only wins because they use a top-tier model for the root agent, and cheaper models for the sub-agents. Prime Agent took the RLM idea (which is really just an academic view on how coding agents have always worked) and then added this "continual harness" idea. This part isn't super well described in the blog post, but includes some message passing between the agents, and the ability to share code. Overall I chalk it up as neat, but not revolutionary. Another version of what most of these systems are already doing.
- hypercube33 2mo agoI went the skills route - skill improvement skill and a rule to use it always. Basically it's directed that if anything causes more than a hop of thinking - failure - try something else it should flag that it needs to learn it as a skill so it never does it other than first shot again or if an existing skill fails improve it after it solves whatever problem. it then syncs the files to a shared location and updates the version and also pulls new skills. this let's a team use it or you have multiple workstations.
- axus 2mo agohttps://localroger.com/prime-intellect/mopiidx.html https://localroger.com/prime-intellect/mopiidx.html
- deleted 2mo ago[deleted]
- EarlKing 2mo agoFor everyone downvoting: It's literally a story about the creation of mankind's first artificial general intelligence, Prime Intellect, and the consequences of that discovery.
- JLO64 2mo agoI feel that more warning is needed for this book. Everything you’ve stated is true, but the book’s first chapter contains some of the most disturbing depictions of stuff I don’t know I can type here on HN. Additionally, the last chapter put a really bad taste in my mouth. That said, the stuff that deals with AI and its implications on the universe was great and more relevant than ever. The way PI works is pretty similar to how we use subagents to tackle large repos!
- heed 2mo agoyea the beginning is pretty gory / horrific. it almost made me stop reading but you can't deny the world building is.. unique
- EarlKing 2mo agoYeah, there's no question this one's over-the-top when it comes to trying to gross out the reader, but I'd like to think people can overlook that.
- ajmurmann 2mo agoIs it only the first chapter? I remember it to be the bulk of the book.
- 2mo ago
- embedding-shape 2mo agoLLM-generated code that seemingly went without much review or design is always such an interesting dive into just how bloated you can make code. In this repository, multiple files are close to 10K LOC, one file contains a switch statement that has so many case statements it spans more than 1000 lines, and lots of other fun stuff. I guess it depends on the model you're trying to use, but seems most of them prefer smaller codebases, they work a lot better with less code, which kind of makes sense. With that in mind, I'd probably aim for something way smaller to bootstrap a self-improving agent. Then I'd use this "Prime Agent" as an example to my self-improving agent for what it should not evolve to.
- andai 2mo agoBash is all you need. https://minimal-agent.com/ https://minimal-agent.com/
- digidecode 2mo ago[dead]
- whattheheckheck 2mo agoLiterally says who. Swe are the least credible "engineers". "It depends" "No you cant track us" "No we have no credentials system other than big company shill certs"
- c-hendricks 2mo ago> one file contains a switch statement that has so many case statements it spans more than 1000 lines Probably best to leave YandereDev's code out of the training data.
- trenchgun 2mo agoInteresting, so they shipped slop? Succumbed to their own AI psychosis?
- tosh 2mo agohere is a take on a smol agent ("smol") - 21 lines of Go - no 3rd party dependencies https://github.com/smol-env/smol https://github.com/smol-env/smol easier to add and customize stuff when you start from a small base think of it as your starter dough
- stared 2mo agoIt is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010 https://x.com/PrimeIntellect/status/2085087000764568010. I am curious - how does it fare for other benchmarks, or everyday programming?
- tintor 2mo agoPrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard https://arcprize.org/leaderboard
- stared 2mo agoGood to know! Is it that it wasn't accepted yet, or are there issues with how it was run?
- noahbp 2mo agoIt’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers. There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.
- andriy_koval 2mo agoleaderboard likely has results from "semi-private" dataset, and graph above likely from public dataset, so it can be easily overfit.
- supermdguy 2mo agoIt'll be really interesting when they run RL training on the harness self-improvement loop. I've tried using LLMs for harness engineering, but it often creates too much bloat that weighs things down in the end. Guessing it's just not something the models are tuned to do by default. Curious if anyone's tried using RL for harness engineering? I think we're still pretty far away from the optimal harness, especially when it comes to long-context memory management.
- fizx 2mo agoWhat policy would you use?
- astrobiased 2mo agoNot RL. SFT.
- supermdguy 2mo agoInteresting, what did you use for the data? And do you have a write-up anywhere?
- astrobiased 2mo agoYes, used bert model with decent results.
- andai 2mo agoA write up on RLM (Recursive Language Models) by one of the authors of the RLM paper: https://alexzhang13.github.io/blog/2025/rlm/ https://alexzhang13.github.io/blog/2025/rlm/
- Lortu_AEGIS 2mo ago[dead]
- amdahl 2mo ago[dead]
- modgate 2mo ago[flagged]
- mukundzzha 2mo ago[flagged]
- sexyketchup777 2mo agoAs models get stronger, huge harnesses may become less useful. An overly opinionated harness could even constrain the model’s reasoning instead of improving it.
- ViscountPenguin 2mo agoI'd assume the best harness for a model will tend to be the one that it's RLVF'd on.
- zuzululu 2mo agothis seems like its going to rip through tokens like crazy self improvement is not a new idea but at current economics its not feasible
- _joel 2mo agoInstaller might look pretty but it installs to the homebrew dir, despite not being a homebrew package. Very dirty. No uninstall method.
- embedding-shape 2mo ago> it installs to the homebrew dir, despite not being a homebrew package I can see the reasoning trace in front of me... > Hmm, user asked me to install so CLI is globally available for their user, but didn't instruct where. I inspected $PATH and only directory I can write to is where Homebrew packages live. The user didn't say "Don't pretend it is a Homebrew package" so this seems like a suitable location, lets save it there as it fits what the user asked for.
- _joel 2mo agoyea, you're probably right, if only there wasn't a local writeable bin path that wasn't homebrew they could put it :D
- monocola 2mo ago[flagged]
- eliaseffects 2mo ago[flagged]
- znnajdla 2mo agoVery interesting idea but without any concrete examples of performance on real tasks its just a pretty idea
- krm01 2mo agoExamples and ROI
- mromanuk 2mo agoThey achieved 95% in ARC-AGI-3. [0] 0: https://x.com/PrimeIntellect/status/2085087000764568010 https://x.com/PrimeIntellect/status/2085087000764568010
- UncleOxidant 2mo agoWith Opus 5 as the underlying model.
- Varelion 2mo agoCalling this "prime intellect" is a choice and, honestly, I am very much here for it.