4 ms·
But that means they're good enough for the short them. This is why I find it such an interesting question to wield, but also, it is why my wording was poor in t
by notashelf 2mo ago
But that means they're good enough for the short them. This is why I find it such an interesting question to wield, but also, it is why my wording was poor in the post.
- deleted 2mo ago[deleted]
- deathanatos 2mo agoShort term can, IME, be very short. I've seen people generate, say, a bash script with an LLM. It's generated: short term, the problem is "solved": we've generated a bash script. … but does it work? Someone comes along, reviews it, "this is garbage, and does not do what it says it purports to do". Perhaps it even gave an output: the script computed … something, but it's just GIGO. But that "check if this works" friction is the same friction that is what people try to avoid by generating it with an LLM in the first place. If you're too lazy to write the script, you're practically by definition too lazy to verify it.
- aprdm 2mo agoYou can have another agent write the tests and verify the former agent ? This is pretty basic stuff. Makes me question if people are actually trying to use AI
- habinero 2mo agoNow you have two problems lol, in that you don't know if the tests are any good or actually test the thing in question. Sooner or later you run out of turtles to put on the stack.
- ileonichwiesz 2mo agoThat’s a solved problem, you just add another agent to check if the tests are any good, and one more to oversee the test-checker, and one more…
- TheOtherHobbes 2mo agoIf this is your working environment, it sounds like quite an unusual place. I literally can't imagine generating a script with an LLM without testing it at all. Bash is one of those situations where LLMs can do really well. No human on Earth can remember all of the commands and even fewer humans can remember all of the switches for all of the commands. This kind of remembering, searching, and assembling is exactly what LLMs are good at - as long as you're not writing a gigantic build system with hundreds of moving parts, in which case you should probably be using something more streamlined anyway.
- lmm 2mo ago> I literally can't imagine generating a script with an LLM without testing it at all. Then you're extremely unimaginative as well as unusually fastidious. Certainly someone - several someones - are generating lots of scripts and not testing them, given the PRs I'm seeing.