9 ms·
I work at OpenAI (not on Codex) and have used it successfully for multiple projects so far. Here's my flow: - Always run more than one rollout of the same prom
by avital 1y ago
I work at OpenAI (not on Codex) and have used it successfully for multiple projects so far. Here's my flow:
- Always run more than one rollout of the same prompt -- they will turn out different
- Look through the parallel implementations, see which is best (even if it's not good enough), then figure out what changes to your prompt would have helped nudge towards the better solution.
- In addition, add new modifications to the prompt to resolve the parts that the model didn't do correctly.
- Repeat loop until the code is good enough.
If you do this and also split your work into smaller parallelizable chunks, you can find yourself spending a few hours only looping between prompt tuning and code review with massive projects implemented in a short period of time.
I've used this for "API munging" but also pretty deep Triton kernel code and it's been massive.
- deleted 1y ago[deleted]
- owebmaster 1y agoCan it be used to fix bugs? Because the ChatGPT web app is full of them and I don't think they are getting fixed. Pasting big amounts of text freezing the tab is one of them.
- dimal 1y agoBugs? Those are grubby human work. Seriously, everyone should get good at fixing bugs. LLMs are terrible at it when it’s slightly non-obvious and since everyone is focusing on vibe coding, I doubt they’ll get any better.
- jampekka 1y agoThe Android app is even worse.
- owebmaster 1y agoIf that is what the best unlimited AI can deliver we are safe for at least 10 years more.
- ionwake 1y agoYou guys are doing great work, codex too, keep at it.
- th0ma5 1y agoDo you find yourself ditching on the things when they change something important with the new prompt? I don't get how people aren't absolutely exhausted by actually implementing this prompt messing advice when I thought there were studies saying small seemingly insignificant changes greatly change the result, hide blind spots, and even having a prompt for engineering a better prompt has knock on increases in instability. Do people just have a higher tolerance for doing work that is not related to the problem than I do? Perhaps I only work on stuff there is no prior example for, but every few days I read someone's anecdote on here and get discouraged in all new ways.
- avital 1y agoNot to downplay the issue you raise but I haven't noticed this. Every iteration I make on the prompts only make the request more specified and narrow and it's always gotten me closer to my desired goal for the PR. (But I do just ditch the worse attempts at each iteration cycle) Is it possible that reasoning models combined with the actual interaction with the real codebase makes this "prompt fragility" issue you speak of less common?
- th0ma5 1y agoNo, I've played with all the reasoning models and they just make the noise and weirdness even worse. When I dig into every little issue, it's always something incredibly bespoke. Like the actual documentation that's on the internet is out of date for the library that was installed and the API changed, the way the one library works in one language is not how it works in the other language, just all manner of surprising things. I really learned a lot about the limits of digital representation of information.
- csmpltn 1y ago> "Look through the parallel implementations, see which is best (even if it's not good enough), then figure out what changes to your prompt would have helped nudge towards the better solution." How can non-technical people tell what's "best"? You need to know what you're doing at this point, look for the right pitfalls, inspect everything in detail... this right here is the entire counter-argument for LLMs eliminating SWE jobs...
- throwuxiytayq 1y agoI don’t think anyone expects software engineers will disappear and get replaced by janitors trained to proompt. I’m sure experts will stick around until the singularity curve starts looking funny. It’s probably gonna suck to enter the industry from now on, though.
- dingnuts 1y ago> I don’t think anyone expects software engineers will disappear holy gaslighting Christ have some links, lots of people think that https://www.reddit.com/r/ITCareerQuestions/comments/126v3pm/since_coding_is_probably_becoming_obsolete_what/ https://www.reddit.com/r/ITCareerQuestions/comments/126v3pm/... https://medium.com/technology-hits/the-death-of-coding-why-coding-will-be-obsolete-in-5-years-ad24234bba73 https://medium.com/technology-hits/the-death-of-coding-why-c... https://medium.com/@TheRobertKiyosaki/are-programmers-obsolete-59c9c65f8992 https://medium.com/@TheRobertKiyosaki/are-programmers-obsole... https://www.forbes.com/sites/hessiejones/2024/09/21/the-automation-takeover-are-software-engineers-becoming-obsolete/ https://www.forbes.com/sites/hessiejones/2024/09/21/the-auto... and on and on, endless thinkpieces about this. Certainly SOMEONE, someone with a lot of money, thinks software engineers are imminently replaceable. > until the singularity curve starts looking funny. well there's absolutely no evidence whatsoever that we've made any progress to bringing about Kurzweil's God so I think regardless of what Sam Altman wants you to believe about "general AI" or those thinkpieces, experts are probably okay.
- cdolan 1y agoI think you are correct that people say this, but its absurd that they are saying it in the first place. Coding/engineering/etc is all problem solving in a strucutred manner. That skill is not going anywhere
- ivraatiems 1y agoHow much faster is this than simply writing the code yourself?
- avital 1y agoEasily 5-10x or even more in certain special cases (when it'd take me a lot of upfront effort to get context on some problem domain). And it can do all the "P2"s that I'd realistically never get to. There was a day where I landed 7 small-to-medium-size pull requests before lunch. There are also cases where it fails to do what I wanted, and then I just stop trying after a few iterations. But I've learned what to expect it to do well in and I am mostly calibrated now. The biggest difference is that I can have agents working on 3-4 parallel tasks at any given point.
- deleted 1y ago[deleted]
- atonse 1y agoThis has been my experience too. Certain tickets that would’ve taken me hours (and in one case, days), I’ve been able to finish in minutes. Other tasks take maybe the same amount of time. But just autocomplete saves micro-effort all day long.
- thearn4 1y agoI end up asking the same question when experimenting with tools like Cursor. When it can one-shot a small feature, it works like magic. When it struggles, and the context gets poisoned and I have to roll back commits and retry part of the way through something, it hits a point where it was probably easier for me to just write it. Or maybe template it and have it finish it. Or vice versa. I guess the point being that best practices have yet to truly be established, but totally hands-off uses have not worked well for me so far.
- sunnybeetroot 1y agoWhy commit halfway through implementing something with Cursor? Can you not wait until it’s created a feature or task that has been validated and tests written for it?
- yieldcrv 1y agohow much would this cost you if you didn't work at OpenAI?
- macrolime 1y agoSounds like you're manually doing something that could form the basis of further reinforcement learning. Nudging the UI slightly for this exact flow could generate good training data.