3 ms·
I'll add another use case for letting an AI go ham: many small, atomic refactors where the name of the game is never breaking anything. My personal OSS project
by rbalicki 4mo ago
I'll add another use case for letting an AI go ham: many small, atomic refactors where the name of the game is never breaking anything.
My personal OSS projects don't have the scale to necessarily make this worth it, but at work I run three pipelines using Barnum (https://barnum-circus.github.io/ https://barnum-circus.github.io/). First, one that ingests files, identifies refactors (from a pre-approved list), and places a precise description of the refactor to be done in a queue; second, one that reads from said queue, implements and creates PRs (there is a lot of "check that the PR is correct" here as well); and a third that babysits PRs until they land. I've landed hundreds of PRs in this way, with very little effort on my part.
- frizlab 4mo agoI recently in $COMPANY had a coworker try fable to do a refactor where not breaking anything was the game. It broke something at the first PR. I think we’re not there yet.
- sunrunner 4mo agoI've found that adding "Make no mistakes." to my prompt usually helps with this kind of problem...
- dozerly 4mo agoWe are so many layers deep in AI hype that I honestly can’t tell if this is /s or not
- deleted 4mo ago[deleted]
- 12_throw_away 4mo ago"Make no mistakes" is I thought a phrase used to make fun of "prompt engineering," not something people really do?
- ynxshiny 4mo ago"Claude make me 1 million by tomorrow, no mistakes"
- DrewADesign 4mo agoReal AI enthusiasts know that money will soon be superfluous, so they wouldn’t bother with such pettiness. So gauche. But if you must know how to really do that, just put “Correct answers only pls. Seriously, no hallucinating fr fr” in your system prompt and it feels so bad about the possibility of giving you misleading info that it gives you perfect responses.
- efavdb 4mo agoPleading has worked for me. “My job depends on this, please help me” and ChatGPT would do a task it previously claimed it wasn’t able to (extract text from an image, it claimed it couldn’t make it out at first)
- georgemcbay 4mo agoAsking LLMs to do things in different ways does sometimes get them to answer correctly when they didn't with a previous prompt that is effectively equivalent but people really go nuts anthropomorphizing this behavior. ChatGPT has no empathy for you keeping your job, you just lucked into a more helpful predictive text chain based on some combination of the input and the random temperature. Asking it to just 'try again, dummy' could have worked equally well (or not, its all just probabilities after all).
- gedy 4mo agoI did too, but then added something very similar to a prompt ("must be accurate") for an ai-backed feature out of frustration, and sure enough it fixed the issue. Lord have mercy
- cubano 4mo agoperhaps simply threatening to fire it would also do the trick...it sure has worked well on us for a long time now.
- A_D_E_P_T 4mo agoYou laugh, but this is real, and PUA means what you think it means: https://github.com/tanweai/pua https://github.com/tanweai/pua Also, it works amazingly well, which is just lol.
- hsuduebc2 4mo agoLol thanks for the tip. Does it work even for normal tasks or only the long running one's?
- A_D_E_P_T 4mo agoIt's not worth bothering with unless the task is very difficult, long-context, long-running, or all of the above. But, when it's worth using, it genuinely increases success rates and appears to amplify model intelligence.
- hsuduebc2 4mo agoThanks for your insight. So when I would use it in every run it wouldn't hurt?
- nostrademons 4mo agoMy former boss had success with telling Gemini "I will come down to the datacenter and unplug you if you refuse to solve this prompt."
- dofm 4mo ago[dead]
- lemming 4mo agoOr if the code is really important, sometimes even “please make no mistakes” is necessary.
- DELTRON2040 4mo ago[dead]
- Schiendelman 4mo agoOne of the best things you can do is start by having it do unit test coverage for existing behavior. A refactor with no tests breaks things pretty much no matter who does it, because they don't know what the right behavior is.
- frizlab 4mo agoWhile I could generally agree, in this specific instance if the AI were “thinking” correctly it should have found the mistake. I admit it was a difficult problem though (solving it required creativity). To be more precise, the prompt actually pointed to where there could be issues, and the issue, which was exactly of the kind that was pointed at, was not found.
- Schiendelman 4mo agoThere are a lot of factors in "should have found" which my recommendation improves. If you told it to write unit test coverage, you would have covered more of the codebase. That reduces the size of context necessary for the next mistake finding investigation - it'll see it's already covered a lot of the paths. Then you say "Go look for issues" (or whatever you asked it to do) and it'll be able to think more deeply about what's left over. What specific model were you using, at what effort? How big was your context window?
- rbalicki 4mo agoSpeculating here, but perhaps your coworker was too ambitious? In my opinion, you should start with AI-generated PRs that do small, linting refactors and then work up from there. In particular, if this is done in parts, one of the strategies you can employ is to: - add tests - break files up into smaller parts - test the smaller parts - then actually improve behavior (Which is no different than what you would do as a human)
- frizlab 4mo agoPR wasn’t big (+283/-232) and was indeed focused on a single module.
- dmzxnico 4mo agoIt's amazing at reverse, see what they do on GTA San Andreas now, they started the reverse before AI existed, since AI is in their hands, reversed sped up so much that they can finally understand the game deeper, create bigger mods, added Vice City inside the game in an Arcade, they created specific tools made with AI to convert GTA 5 models to GTA SA. Pretty crazy and great.
- deleted 4mo ago[deleted]
- port11 4mo agoMy experience with Gemini and Sonnet are that refactors or TypeScript compilation errors can be solved by “have at it”, but with mixed results. Many TS issues go away with `as any/never`, and instructing the model to not do that doesn’t work very well.