3 ms·
This is a very interesting introduction to a blog post, but... I'm somehow missing the actual blog post. How does this stuff work in practice? What are some con
by derdi 3mo ago
This is a very interesting introduction to a blog post, but... I'm somehow missing the actual blog post. How does this stuff work in practice? What are some concrete examples? How does one get from JavaScript tokenizing things in a commit hook to validating that the LLM didn't disable tests it didn't agree with, or any other helpful property?
- gritzko 3mo agoI am the author. I am trying to limit one post to one page. Most people here are reading reasoning all day, I am afraid. Might get tired. I also aspire to make one post a day. To be continued.
- chickensong 3mo agoYou need to always be looking for what can be done deterministically and what can't. If it can, write a script or whatever is needed to make that happen. Your agent can help you figure this out. The agent becomes a glue layer for all your scripts. Use LLM judgement as an extra layer on top of a mechanical baseline. > validating that the LLM didn't disable tests it didn't agree with Provide a test runner and force the agent to call it. Have it emit something if you want evidence.
- derdi 3mo agoI know how to write a test that verifies other tests. But I can do that already, I'm wondering what Beagle would add. I also know that if the meta-test is writable to the agent, it will change it if it feels like it wants to get rid of some other test. Even if it can't change the meta-test, it can hollow out existing tests to make them pass trivially. I don't think nondeterminism is the problem. The problem is following rules: If I tell the agent not to change tests, it can conveniently "forget" about this. It doesn't much matter if it forgets deterministically. The problem is that it can forget at all.
- chickensong 3mo agoMy bad, the article was fairly general and I thought your question was general as well. Having followed some of the links now, I think your question still stands. The only way I've found to really force rules is via hooks, and even then I think it's just prompt injection? Maybe some kind of hook/checksum thing to ensure you're running unchanged tests, but as you're pointing out, the agents can get sneaky and do weird stuff if they have write ability.