3 ms·
Ungrounded LLM outputs are a bit like your dreams. Without anything to test hypotheses against, stuff can pop in and out of existence and physics is just advice
by cadamsdotcom 2mo ago
Ungrounded LLM outputs are a bit like your dreams. Without anything to test hypotheses against, stuff can pop in and out of existence and physics is just advice.
Ground your LLM. Tests, documentation, give it many ways to run the thing its reasoning about. It needs to be able to test its hypotheses on its own.
Take yourself out of that loop so you only find out once it's sure.
- tra3 2mo agoLove LLMs gonna keep using them. It feels like your suggested approach is expensive, in terms of tokens. I feel (second time I say this) that when I steer the process I get pretty good results vs my coworkers that let the LLMs run away. I do have data on our token usage, not much in terms of quality of the delivery. I keep thinking about the c compiler implementation that anthropic shared earlier in the year that had all the requirements you mention and arguably wasn’t that great.
- vkazanov 2mo agoThr thing is that both you and your agent should have a way to verify the solution. OBVIOUSLY, the compiler experiment was just a cringe pr stunt. But it has a point: everything works better with a good testing loop, and compilers always have one by thr nature of the work they do
- cadamsdotcom 2mo ago> expensive, in terms of tokens. No amount of tokens can come close to my hourly rate.
- mrtesthah 2mo agoDo you steer your agents by manually running every single test and linter and reporting the results back to them?
- amelius 2mo agoBut this is exactly what the AI labs should be doing ...
- moffkalast 2mo agoAnd they are, at least for Claude I know it writes random mocks and tests in its virtual env even in the web version, cause it sometimes annoyingly includes them in the final result. It's the only reason it produces anything that runs.
- dgellow 2mo agoThat’s exactly why I don’t believe LLMs will cure cancer anytime soon, make terrible lawyers, shouldn’t be trusted for medical decisions, etc. software and maths are some really the niches where we have great, battle tested, reliable validation tools. That’s not the case for “softer” domains
- cadamsdotcom 2mo agoSoftware is a very "spiky domain"; things either work or fail, and there is sharp delineation between and easy verification. Hm. Two orthogonal properties! This sounds like a 2x2 matrix! Let's swap hard/easy around & explore the 4 possibilities... There are domains with sharp delineation and hard verification; they are not at risk until AI gets much better. Humans operate in these domains by applying tremendous deep thought and subjective judgment - our superpower. Domains with soft delineation and easy verification are most at risk: "it's a picture of a cat" remains true through a wide range of perturbations - eg. skewing the image or moving it across a pixel or correcting its white balance or even changing the cat. AI music? Lots of domains already solved by AI here but they're also not that meaty. My prediction is the next interesting stuff will happen where verification is hard but there's no sharp delineation. It's the world of "I'll know it when I see it". Good customer service?
- dgellow 2mo agoEh, that’s a very interesting way to differentiate, I will steal your explanation next time I have that discussion, if you don’t mind!
- gwerbin 2mo agoThe only people who think LLMs would make good lawyers are the people selling LLMs. The more practical among us recognize that LLMs are our amazing tools for searching through and making sense of large amounts of text with a high level of sophistication, which can significantly enhance the productivity of a human lawyer.