19 ms·
We put a coding agent in a while loop
- taw1285 1y agoThis is so amazing. Are there any resources or blogs on how people do this for production services? In my case, I need to rewrite a big chunk of my commerce stack from Ruby to Typescript.
- ghuntley 1y agoNice. Check out https://ghuntley.com/ralph https://ghuntley.com/ralph to learn more about Ralph. It's currently building a Gen-Z esoteric programming language and porting the standard library from Go to the Cursed programming language. The compiler is working, I'm just finishing up the touches of the standard library before launching. The language is called Cursed.
- sfarshid 1y agoThanks Geoff, Ralph was our inspiration to do this! We were curious to see if we can do away with IMPLEMENTATION_PLAN.md for this kind of task
- gregpr07 1y agoAGI was just 1 bash for loop away all this time I guess. Insane project.
- dhorthy 1y agowas deeply unsettling among other things
- ghuntley 1y agoIt is, isn't it mate? Shit, I stumbled upon Ralph back in February and it shook me to the core.
- cogogo 1y agoNot that I want to be shaken but what is Ralph? A quick search showed me some marketing tools but that cant be what you are referring to is it?
- ghuntley 1y agoRalph is a technique. The stupidest technique possible. Running an agent in a while true loop. https://ghuntley.com/ralph https://ghuntley.com/ralph
- cogogo 1y agoLess flippantly that was sort of my thought. I’m probably a paranoid idiot and I’m not really sure I can articulate this idea properly but I can imagine a less concise but broader prompt and an agent configured in a way it has privileges you dont want it to have or a path to escalate them and its not quite AGI but its a virus on steroids - like a company or resource (think utilities) killer. I hope Im just missing something but these models seem pretty capable of wreaking all kinds of havoc if they just keep looping and have access nobody in their right mind wants.
- rukuu001 1y agoJust need to add ID.md, EGO.md and SUPEREGO.md and we're done.
- apwell23 1y ago[flagged]
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- thebiglebrewski 1y agoThe agent terminating its own process was hilarious
- ghuntley 1y agoIt's why I called it Ralph. Because it's just not all there, but for some strange reason it gets 80% of there pretty well. With the right observational skills, you can tune it into 81, then 82, then 83, then 84. But there's always gaps, always holes. It's a lovable approach, a character, just like Ralph Wiggum.
- wrs 1y agoI’ve done a few ports like this with Claude Code (but not with a while loop) and it did work amazingly well. The original codebase had a good test suite, so I had it port the test suite first, and gave it some code style guidance up front. Then the agent did remarkably well at doing a straight port from one imperative language to another. Then there’s some purely human work to get it really done — 80-90% done sounds about right.
- giantg2 1y agoThere's a lot of "it kind of worked" in here. If we actually want stuff that works, we need to come up with a new process. If we get "almost" good code from a single invocation, you just going to get a lot of almost good code from a loop. What we likely need is a Cucumberesque format with example tables for requirements that we can distill an AI to use. It will build the tests and then build the code to to pass the tests.
- ghuntley 1y agoStrangely enough, TLA+ and other formal proofs work very well for driving Ralph.
- giantg2 1y agoI would consider that expected but not strange. The thing blocking adoption is that most devs/people find those formal languages difficult or boring. That's even true of things like Cucumber - it's boring and most organizations care little for robust QA.
- VincentEvans 1y agoThere will be a a new kind of job for software engineers, sort of like a cross between working with legacy code and toxic site cleanup. Like back in the day being brought in to “just fix” a amalgam of FoxPro-, Excel-, and Access-based ERP that “mostly works” and only “occasionally corrupts all our data” that ambitious sales people put together over last 5 years. But worse - because “ambitious sales people” will no longer be constrained by sandboxes of Excel or Access - they will ship multi-cloud edge-deployed kubernetes micro-services wired with Kafka, and it will be harder to find someone to talk to understand what they were trying to do at the time.
- Jtsummers 1y agoSuperfund repos.
- throwup238 1y agoNow that's an open source funding model governments can get behind.
- binary132 1y agoA lot of big open source repos need to be given the superfund treatment
- notachatbot123 1y agoA whole bigger lot of closed source software needs to be given the superfund treatment!
- cruffle_duffle 1y agoWhat makes you so sure it will have a repo? I don’t recall the last time Claude suggested anything about version control :-)
- mcny 1y agoClaude will give what you asked for. My sensible chuckle moment was when I asked it to create a demo asp net web API and it did everything but add the authorize tag or any kind of authentication. I asked what was missing and until i mentioned it, it didn't mention authentication or authorization at all.
- NitpickLawyer 1y ago> After finishing the port, most of the agents settled for writing extra tests or continuously updating agent/TODO.md to clarify how "done" they were. In one instance, the agent actually used pkill to terminate itself after realizing it was stuck in an infinite loop. Ok, now that is funny! On so many levels. Now, for the project itself, a few thoughts: - this was tried before, about 1.5 years ago there was a project setup to spam github with lots of "paper implementations", but it was based on gpt3.5 or 4 or something, and almost nothing worked. Their results are much better. - surprised it worked as well as it did with simple prompts. "Probably we're overcomplicating stuff". Yeah, probably. - weird copyright / IP questions all around. This will be a minefield. - Lots of SaaS products are screwed. Not from this, but from this + 10 engineers in every midsized company. NIH is now justified.
- ghuntley 1y ago> - weird copyright / IP questions all around. This will be a minefield. Yeah, we're in weird territory because you can drive an LLM as a Bitcoin mixer over intellectual property. That's the entire point/meaning behind https://ghuntley.com/z80 https://ghuntley.com/z80. You can take something that exists, distill it back to specs, and then you've got your own IP. Throw away the tainted IP, and then just run Ralph over a loop. You are able to clone things (not 100%, but it's better than hiring humans).
- rasz 1y ago>and then you've got your own IP. except you dont
- heavyset_go 1y ago> then you've got your own IP. AI output isn't copyrighted in the US.
- miohtama 1y agoHe is referring to taking AI output and making it your company's property.
- bigmattystyles 1y agoStarting to think of this quote more and more: "This business will get out of control. It will get out of control and we'll be lucky to live through it." https://www.youtube.com/watch?v=YZuMe5RvxPQ&t=22s https://www.youtube.com/watch?v=YZuMe5RvxPQ&t=22s
- ramraj07 1y agoThe irony is that everyone did live through that business. So what youre saying is we will live through this too!
- cptroot 1y agoAre you feelin' lucky?
- hugh-avherald 1y agoIf there's anything Tom Clancy has taught me it's that everything works out in the end.
- rkachowski 1y ago> In one instance, the agent actually used pkill to terminate itself after realizing it was stuck in an infinite loop. The alexandrian solution to the halting problem.
- arresin 1y agoNo it did not.
- vntok 1y agoDo you have more current information than the authors who say it did?
- MagMueller 1y agoI would love to fix my docs with this. I have them in the main browser-use repo. What do you recommend that the agent does never push to main browser-use, but only to its own branch?
- dhorthy 1y agoYeah you can easily tweak this to push to a branch or a fork or something in the generated prompt.md
- kh_hk 1y agoI am honestly surprised how we went from almost OCD TDD and type purism, to a "it kinda works" attitude to software.
- baq 1y agoalways has been, the difference is now the 'it compiles, ship it' loop is 10x-100x faster than 2 years ago
- precompute 1y agoFaster development speeds make people implicitly believe they won't be accountable for the results of their actions.
- kuschku 1y agoThere's always been both sides. One creating a foundation of absolutely stable, reliable code, methodically learning from every mistake. This code lives for many decades to come. The other building throwaway projects as fast as possible, with no regard to specs, constraints, reliability or even legality. They use ecery trick in the book, and even the ones that aren't yet. They've always been much faster than the first group. Except AI now makes the second group 10× faster yet again.
- Dilettante_ 1y agoLiterally just read a blogpost[1] about this. Gist: The two ebb and flow in waves. "It kinda works" produces innovation, OCD hones the artifacts until it runs out of material and the cycle continues. [1]https://worksonmymachine.ai/p/safe-is-what-we-call-things-later https://worksonmymachine.ai/p/safe-is-what-we-call-things-la...
- bwestergard 1y agoThere are always two major results from any software development process: a change in the code and a change in cognition for the people who wrote the code (whether they did so directly or with an LLM). Python and Typescript are elaborate formal languages that emerged from a lengthy process of development involving thousands of people around the world over many years. They are non-trivially different, and it's neat that we can port a library from one to the other quasi-automatically. The difficulty, from an economic perspective, is that the "agent" workflow dramatically alters the cognitive demands during the initial development process. It is plain to see that the developers who prompted an LLM to generate this library will not have the same familiarity with the resulting code that they would have had they written it directly. For some economic purposes, this altering of cognitive effort, and the dramatic diminution of its duration, probably doesn't matter. But my hunch is that most of the economic value of code is contingent on there being a set of human beings familiar with the code in a manner that requires writing having written it directly. Denial of this basic reality was an economic problem even before LLMs: how often did churn in a development team result in a codebase that no one could maintain, undermining the long-term prospects of a firm?
- tikhonj 1y agoThere's a classic Peter Naur paper about this from 1985: "Programming as Theory Building" https://pages.cs.wisc.edu/~remzi/Naur.pdf https://pages.cs.wisc.edu/~remzi/Naur.pdf
- metadat 1y agoDiscussed 7 months ago (45 comments): https://news.ycombinator.com/item?id=42592543 https://news.ycombinator.com/item?id=42592543 Great read overall, an interesting challenge to the conception that at its core, programming is about producing code.
- grimgrin 1y agofound a copy that isn't a scanned paper: https://gist.github.com/dpritchett/fd7115b6f556e40103ef https://gist.github.com/dpritchett/fd7115b6f556e40103ef
- beefnugs 1y ago"At one point we tried “improving” the prompt with Claude’s help. It ballooned to 1,500 words. The agent immediately got slower and dumber. We went back to 103 words and it was back on track." Isn't this the exact opposite of every other piece of advice we have gotten in a year? Another general feedback just recently, someone said we need to generate 10 times, because one out of those will be "worth reviewing" How can anyone be doing real engineering in such a: pick the exact needle out of the constantly churning chaos-simulation-engine that (crashes least, closest to desire, human readable, random guess)
- dhorthy 1y agoHmm what sorts of advice in the last year are you referring to? Like the “run it ten times and pick the best one” thing? Or something else? I kind of agree that picking from 10 poorly-promoted projects is dumb. The engineering is in setting up the engine and verification so one agent can get it right (or 90% right) on a single run (of the infinite ish loop)
- jjani 1y ago> Hmm what sorts of advice in the last year are you referring to? They're almost certainly referring to first creating a fleshed out spec and then having it implement that, rather than just 100 words.
- mistrial9 1y agothe core might be - the difference between an LLM context window, and an agent's orders in a text. LLM itself is a core engine, running in an environment of some kind (instruct vs others?). Agents on the other hand, are descendants of the old Marvin Minsky stuff in a way.. it has objectives and capacities, at a glance. LLMs are connected to modern agents because input text is read to start the agent.. inner loops are intermediate outputs of LLM, in language. There is no "internal code" to this set of agents, it is speaking in code and text to the next part of the internal process. There are probably big oversights or errors in that short explanation. The LLM engine, the runner of the engine, and the specifics of some environment, make a lot of overlap and all of it is quite complicated. hth
- nis0s 1y agoWhy is this flagged?
- cluckindan 1y agoNow I want to put one of these in a loop, give it access to some bitcoin, and tell it to come up with a viable strategy to become a billionaire within the next month.
- hoppp 1y agoI wanted to know how much it cost? I would be scared to run this without knowing the exact cost. Its not a good idea to do it without a payment cap for sure, its a new way to wake up with a huge bill the next day.
- bckr 1y ago$800
- debazel 1y agoThey did mention how much they spent here: https://github.com/repomirrorhq/repomirror/blob/main/repomirror.md#numbers https://github.com/repomirrorhq/repomirror/blob/main/repomir... > We spent a little less than $800 on inference for the project. Overall the agents made ~1100 commits across all software projects. Each Sonnet agent costs about $10.50/hour to run overnight.
- rogerrogerr 1y agoDoes anyone else get dull feelings of dread reading this kind of thing? How do you combat it?
- dhorthy 1y agocombat how? (And yes, yes I do)
- rogerrogerr 1y agoCombat the feelings, I guess. Not really sure.
- zdwolfe 1y agoYes, and so far I haven't been able to combat it.
- shaky-carrousel 1y agoBy being there when FrontPage was released. This is just the same, all over again.
- swader999 1y agoFrontPage with Clippy of to the corner yelling 'You're absolutely right!'
- bitexploder 1y agoStoicism. Dichotomy of control. Is this something you can control? If no, don’t dread. If yes, do something. Often, all you have firmly in your grasp are things inside of your brain. Catch the negative thought. Acknowledge it. Move on. Do not dwell. Take proactive steps to be ready in your career. You do tech ling enough and you live through multiple cycles like this.
- rogerrogerr 1y agoAll appreciated, thanks. Any thoughts on what those proactive steps would be? I'm early-career (26yo, "senior" software dude at a defense outfit).
- rozab 1y agoThese people are weird. The blog post that inspired this has this weird iMessage screenshot, like a shitty investment grift facebook ad: https://ghuntley.com/ralph/ https://ghuntley.com/ralph/ Apparently one of the lucky few who learned this special technique from Geoff just completed a $50k contract for $297. But that's not all! Geoff is generous to share the special secret prompt that unlocked this unbelievable success, if only we subscribe to his newsletter! "This free-for-life offer won't last forever!" I am sceptical.
- imiric 1y agoI can't tell whether this "technique" is serious or a joke, and/or if it's some elaborate grift. In any case, the writing style of that entire blog is off-putting. Gibberish from a massive ego.
- ghuntley 1y agoIt's both serious and a joke. The seriousness is that it works (to point) and the implications to our profession as software developers. The joke is just how stupid it is. Refer to the original story link above for proof of outcomes.
- ofjcihen 1y agoThis reply is actually making this more confusing.
- efitz 1y agoIn one instance, the agent actually used pkill to terminate itself after realizing it was stuck in an infinite loop. That is pretty awesome and not something I would have expected from an agent; it hints (but does not prove) that it has some awareness of its own workings.
- taberiand 1y agoIt hints that a suitable auto completion of the input prompt is to output a pkill command
- salomonk_mur 1y agoWe, too, are just auto-complete, next-token machines.
- seba_dos1 1y agoWe are auto-complete next-token machines, but plastic and attached to many other not less important subsystems, which is a crucial difference.
- efitz 1y agoI honestly think that partially-OSS SaaS is in for a rocky road; many popular paid or freemium tools are likely to be rewritten by AI and published as OSS with permissive licenses over the next year or two. I also think that the same capability will largely invalidate the GPL, as people point agents at GPL software and write new software that performs the same function as OSS with more permissive licenses. My reasoning is this: the reason that people use OSS versions of software that has restrictive licensing terms, is because it’s not worth the effort to them to rewrite. Corporations certainly, but also individuals, will be able to use similar approaches to what these people used, and in a day or two come back to a mostly-functional (but buggy) new software package that does most of what the original did, but now you have a brand new software that you control completely and you are not beholden to or restricted by anyone. Next time someone tries to pull an ElasticSearch license trick on AWS, AWS will just point one or a thousand agents at the source and get a brand new workalike in a week written in their language du jour, and have it fully functional in a couple of months. Doesn’t circumvent patent or trademark issues but it’ll be hard to assert that it’s not a new work, esp. if it’s in an entirely different language. Just something I’ve been thinking about recently, that LLM agents change the game when it comes to software licensing.
- popcorncowboy 1y ago> partially-OSS SaaS is in for a rocky road Agent-in-a-loop gets you remarkably far today already. It's not straightforward to "rip" capability even when you have the code, but we're getting closer by the week to being able to go "Project X has capability Y. Use [$approach] and port this into our project". This HAS to put a fat question mark over the viability of any SaaS that makes their code visible.
- ath3nd 1y agoAnd I hired a cleaning lady, paid her £200, and when I came back, the house was clean. The difference is that I did not write a blog post about it, nor did I got overly excited about it as if I had just discovered sliced bread, nor did I harbor any illusions that it was me who did anything of value. Next, I will write a while loop filling my disk with files of random sizes and with random byte content inside. I will update you on the progress when I am back tomorrow. I do expect great results and a nicely filled disk!
- andyferris 1y agoPeople keep saying that Gemini 2.5 Pro can solve some problem that Sonnet 4 cannot, or that GPT5 can solve a problem that Gemini 2.5 Pro cannot, or that Sonnet 4 can solve some problem that GPT5 cannot. There was a blog article about mixing together different agents into the same conversation, taking turns at responses and improving results/correctness. But it takes a lot of effort to make your own claude-code-clone with correct API for each provider and prompts tuned for those models and tool use integrated etc. And there's no incentive for Anthropic/OpenAI/Google to write this tool for us. OTOH it would be relatively easy for the bash loop to call claude code, codex CLI, etc in a loop to get the same benefit. If one iteration of one tool gets stuck, perhaps another LLM will take a different approach and everything can get back on track. Just a thought.
- yyhhsj0521 1y ago> it takes a lot of effort to make your own claude-code-clone Maybe we could try write that into a markdown file, and let Claude code at it for one night in a while loop
- ilijavanil 1y agohttps://github.com/albertvucinovic/chat.sh https://github.com/albertvucinovic/chat.sh
- rane 1y ago> People keep saying that Gemini 2.5 Pro can solve some problem that Sonnet 4 cannot Most definitely can. It's insane how well just telling Claude to ask help from Gemini works in practice. https://github.com/raine/consult-llm-mcp https://github.com/raine/consult-llm-mcp Disclaimer: made it
- billylo 1y agoThank you. Will try it today.
- ofjcihen 1y agoAs a security professional who makes most of my money from helping companies recover from vibe coded tragedies this puts Looney Toons style dollar signs in my eyes. Please continue.
- discordance 1y agoWould love to hear more about your work and how you have tapped into that market if you're keen to share. Even if it's just anecdotes about vibe-in-production gone wrong, that would be really entertaining.
- ofjcihen 1y agoAbsolutely. Before vibe coding became too much of a thing we had the majority of our business coming from poorly developed web applications coming from off shore shops. That’s been more or less the last decade. Once LLMs became popular we started to see more business on that front which you would expect. What we didn’t expect is that we started seeing MUCH more “deep” work wherein the threat actor will get into core systems from web apps. You used to not see this that much because core apps were designed/developed/managed by more knowledgeable people. The integrations were more secure. Now though? Those integrations are being vibe coded and are based on the material you’d find on tutorials/stack etc which almost always come with a “THIS IS JUST FOR DEMONSTRATION DONT USE THIS” warning. We also see a ton of re-compromised environments. Why? They don’t know how to use CICD and just recommit the vulnerable code. Oh yeah, before I forget, LLMs favor the same default passwords a lot. We have a list of the ones we’ve seen (will post eventually) but just be aware that that’s something threat actors have picked up on too. EDIT: Another thing, when we talk to the guys responsible for the integrations or whatever was compromised a lot of the time we hear the excuse “we made sure to ask the LLM if it was secure and it said yes”. I don’t know if they would have caught the issue before but I feel like there’s a bit of false comfort where they feel like they don’t have to check themselves.
- ofjcihen 1y agoOH MAN I almost forgot. We’ve had a few of these stem from custom LLM agents. The most hilarious one we’ve seen was one that you could get to print its instructions pretty easily. In the instructions was a bit about “DON’T TALK ABOUT FILES LABELED X”. No guardrails other than that. A little creative prompting got it to dump all files labeled X.
- x3haloed 1y agoNice! I've been thinking that we need something like this for a while. I didn't realize it could be so simple! I've been looking into other techniques as well like making a little hibernation/dehydration framework for LLMs to help them process things over longer periods of time. The idea is that the agent either stops working or says that it needs to wait for something to occur, and then you start completions again upon occurrence of a specific event or passage of some time. I have always figured that if we could get LLMs to run indefinitely and keep it all in context, we'd get something much more agentic.
- stevage 1y agoIt'd be pretty interesting to do this with no predermined goal. Get ai to find a project to work on,zand just work on it for a while until it thinks it's done, then start on the next one.
- billylo 1y agoGreat idea. And host them on github-next-2025.com or maybe use it as a new benchmark to see how models/tools progress.
- phplovesong 1y agoDamn. I can not even start to grasp the slop-level on this one. I guess a software devs future is to read slop commits and prs, and somehow try to unfuck what the ai did generate. I rather be on the pigfarm shoveling pigshit and castrate bulls.
- dorgo 1y agonaaa, you just run "unfuck it" in a loop..
- eisbaw 1y agohttps://gist.github.com/eisbaw/8edc58bf5e6f9e19418b2c00526ccbe0 https://gist.github.com/eisbaw/8edc58bf5e6f9e19418b2c00526cc... produced https://github.com/eisbaw/CMake-Nix https://github.com/eisbaw/CMake-Nix and it works
- lionkor 1y agoI recently tried vibe-coding a pretty simple program. All I can say is that I'm horrified at people doing this. Not only did it produce extremely inadequate solutions (not for a lack of trying), but also these solutions were BARELY fulfilling the requirements, and nothing else. At one point, I gave it a scenario which demonstrated a common failure case, such an important one that it would have broken horribly in production. Its reaction was to make hundreds of changes, one of which, hidden behind hundreds of other changed lines, was to HARDCODE the special case which I had shown it. Of course, that test then passed, and I assumed it had fixed the problem. It was only much later that I discovered this special-case handling. It was not caught during multiple rounds of AI code review. Another instance of such a fuck-up was that the AI insisted on fixing tests which were failing, which it had written, but it kept continuously failing to do so. It ended up making hundreds of changes across various functions, sometimes related, sometimes unrelated, and never figured out that the test itself was not relevant and made no sense after a recent refactor. The AI completely failed to consider, after many rounds of back and forth and trying, to take a single step back and look at the function itself, instead of the line that was failing. This happens every time I touch AIs and try to let them do work autonomously, regardless of which AI it is. People who think these AIs do a good job are the same people who would get chewed up during a 5 minute code review by a senior. I am genuinely afraid for the horseshit quality ""work"" people who use AI extensively are outputting. I use AIs as a way to be more productive; if you use it to do your job for you, I pray for the people who have to use your software.
- missingdays 1y ago> was to HARDCODE the special case which I had shown it. Happened to me as well while trying out GPT-5. My prompt was something like "fix this test", where the test contained a class Foo. It gave me the solution in the form of "if element.class == 'Foo': return null". Gave me a laugh at least
- octodoctor 1y agoYeah, I’ve run into that too. When you let the AI "drive" completely, it tends to patch symptoms instead of reasoning about the system. I wouldn’t trust it to autonomously fix production code either. Where it does shine for me is in the grindy parts: refactoring, writing boilerplate, scaffolding new components, or even surfacing edge cases I hadn’t thought about. I’m building FreeDevTools, and I still do the design + final decision-making myself. The AI just helps me move faster across SEO, styling, bug-fixing, backend/frontend glue code, etc. Basically, I treat it more like a junior pair programmer, useful for speed, but absolutely not a replacement for review, testing, or architectural thinking. https://github.com/HexmosTech/FreeDevTools https://github.com/HexmosTech/FreeDevTools
- deafpolygon 1y agoMakes me wonder if we will see books, documentation, etc, written in this way. Imagine fiction books entirely written by AI, prompted on by humans.
- franze 1y agoI coded (well code directed) Floktoid https://floktoid.franzai.com/ https://floktoid.franzai.com/ via Claude Code in the Cloud (a cheap Hetzner Server). whenever it goes idle it self prompts it to "Continue, if nothing else to do, read CLAUDE.md and continue from there." max 5 times per hour, if this is reached I get an email to check (I hardly get emails). see the repo to judge code quality
- nkmnz 1y agohi simon, will the vue3 version of assistant ui be maintained? that would be awesome! p.s.: funny to meet again here. Last time we met was 2022 in Berlin! congrats to your journey so far!!
- ManlyBread 1y agoSeems like the agent takes a lot of liberties when it comes to porting stuff: https://github.com/search?q=repo%3Arepomirrorhq%2Fbetter-use+%22For+Now%22&type=code https://github.com/search?q=repo%3Arepomirrorhq%2Fbetter-use... None of these issues seem to be documented outside these files.
- fergie 1y agoI'm curious: does Typescript make sense as a language for machines?
- curtisszmania 1y ago[dead]
- leeroihe 1y agoWhy does anyone over 22 "compete" in hackathons?.... you're literally just giving shareholders and clout seekers work for free...
- sunir 1y agoI have been developing long lived self-directing agent loops. Longest with problem solving has been about 4 hours. Longest without problem solving has been nearer 8 hours until it was done. The biggest problem is simply what we think is clear is confusing to the AIs. They seem like they speak English fluently but they are aliens. You need to force them to active listen first and write out what they understand then reload them with a clean context with the written understanding and confirm. Ideation is also mostly limited to synthesis. So it’s better to work on problems that get progressively more complete towards a known objective rather than problems that require exploration.
- ghuntley 1y ago> So it’s better to work on problems that get progressively more complete towards a known objective rather than problems that require exploration. Yes. The longest I've had a self-directing agent loop running is a cumulative of three months. One goal, one purpose. Every now and then I modify the prompt in the background, and the agent picks up the updated prompt on the next loop.
- mring33621 1y agovery interesting! how's it doing? are you making progress toward your goal?
- _mocha 1y agoI'm retired from the industry, and posts like these take me back to the early days of cybersecurity (where people memorized scripts). Talking with my nephews and nieces, I can already tell that many new grads struggle with fundamentals—things like choosing the right data types and containers for short-lived strings, understanding how memory allocation works, or even marginally improving a basic hashing function. I worry the next decade will bring an influx of undertrained engineers.
- klysm 1y agoI agree understanding how memory allocation works, but not sure I would agree that understanding how to _improve_ a basic hashing function is very important.
- Dilettante_ 1y agoAs the "ceiling" grows, the "floor" of what's considered "fundamentals" moves in the same direction.
- ponector 1y ago>>I worry the next decade will bring an influx of undertrained engineers. How can it be the other way? No one is investing into education of their developers, also trying to save some money on them. Get cheap fresh grads and make them develop new stuff!
- reedlaw 1y agoIronic to see this juxtaposed with another front-page story, "We put agentic AI browsers to the test – They clicked, they paid, they failed". The closing thoughts in the linked article ("feeling the AGI" and "very beginning of the exponential takeoff curve") leave me feeling skeptical considering this project prompted agents to port existing code into another language. Impressive, but it doesn't lead me to believe a singularity event is imminent.
- joshmlewis 1y agoOne of the biggest nuggets people need to take away from this: > At one point we tried “improving” the prompt with Claude’s help. It ballooned to 1,500 words. The agent immediately got slower and dumber. We went back to 103 words and it was back on track. Keep your prompts / agent instructions short. Focus on the wide view, not specifics.
- narmiouh 1y agoThe problem with simple prompts it it gives AI the most leeway to do what it thinks is important, which may work in case of just convert from one language to another (even in those cases it took its own liberties - good or bad). This won't work when devising a new application from scratch or where you are expecting consistent agentic output to be usable in a predictable situation.
- MagMueller 1y agoWe could do a hackathon where its only allowed to change 1 line.
- wnolens 1y ago> Each Sonnet agent costs about $10.50/hour to run overnight. When expressed like that, I can't help but see it as a wage figure.
- didibus 1y agoThat's already what an agent is though. It's LLM invocation inside a loop where the exit condition is supposed to be some goal for the agent to have met, which you generally provide some heuristic or deterministic criteria so it can assert that the goal is reached or not. I'm not sure about Claude Code, but with Amazon Q, if you prompt it in a vibe coded way, and give it a goal like that and say to keep going until the goal criteria is met (which could be running tests and passing them, or running a sub-agent that evaluates if the goal is met). Then I've seen it go for like 2 hours before it ended.
- wedn3sday 1y agoIn the tribunals to come, whoever implemented the --dangerously-skip-permissions flag will be prosecuted as a war criminal.