4 ms·
This is very cool, somewhat inspiring, and (personally) very informative: I didn't actually know what "agentic" AI use was, but this did an excellent job (incid
by RangerScience 1y ago
This is very cool, somewhat inspiring, and (personally) very informative: I didn't actually know what "agentic" AI use was, but this did an excellent job (incidentally!) explaining it.
Might poke around...
What makes something a good potential tool, if the shell command can (technically) can do anything - like running tests?
(or it is just the things requiring user permission vs not?)
- tough 1y ago> What makes something a good potential tool, if the shell command can (technically) can do anything - like running tests? Think of it as -semantic- wrappers so the LLM can -decide- what action to take at any given moment given its context, the user prompt, and available tools names and descriptions. creating wrappers for the most used basic tools even if they all pipe to terminal unix commands can be useful. also giving it speicif knowledge base it can consult on demand like a wiki of its own stack etc
- notpushkin 1y agoAlso it’s safer than just giving unrestricted shell access to an LLM.
- tough 1y agothat too, ideally autonomous agents will be only spawnrd in their own secure environments using docker or vm’s or posix / unix security but yeah
- radanskoric 1y agoThanks, sharing my learnings on how coding agents work was my main intention with the article. Personally I was a bit surprised by how much of the "magic" is coming directly from the underlying LLM. The shell command can run anything really. When I tested it, it asked me multiple times to run the tests and then I could see it fixing the tests in iterations. Very interesting to observe. If I was to improve this to be a better Ruby agent (which I don't plan to do, at least not yet), I would probably try adding some Rspec/Minitest specific tools that would parse the response and present it back to the LLM in a cleaned up format.
- elif 1y agoWhy stop there? Give it a capybara tool and make it a full TDD agent
- radanskoric 1y agoThat's a very neat idea, maybe even add something like browser-use to allow it to implement a Rails app and try it out automatically. I think you should try it. :) I'm being serious. This sounds like a fun project but I have to turn my attention to other projects for the near future. This was more of an experiment for me, but it would be cool to see someone try out that idea.
- RangerScience 1y agoDo you know of examples of other agents with more defined tools, to use as inspiration/etc? (Like - what would it look like to clean up test results for an LLM?)