12 ms·
Princeton group open sources "SWE-agent", with 12% fix rate for GitHub issues
- bwestergard 3y agoFriendly suggestion to the authors: success rates aren't meaningful to all but a handful of researchers. They should add a few examples of tests SWE-agent passed and did not pass to the README.
- nyrikki 3y agoYes please, the code quality on Devin was incredibly poor in all examples I traced down. At least from a maintainability perspective. I would like to see if this implementation is less destructive or at least more suitable for a red-green-refactor workflow.
- deleted 3y ago[deleted]
- NegativeLatency 3y agoUnless you weren't actually that successful but need to publish a "successful" result
- mdaniel 3y agoI think that "Demo" link is just an extremely annoying version of an HTML presentation, so they could save me a shitload of clicking if they just dumped their presentation out to a PDF or whatever so I could read faster than watching it type out text as if it was live. It also whines a lot on the console about its inability to connect to a websocket server on 3000 but I don't know what it would do with a websocket connection if had it
- SrslyJosh 2y agoProbably created with an LLM.
- matthewaveryusa 3y agoVery neat. Uses the langchain method, here are some of the prompts: https://github.com/princeton-nlp/SWE-agent/blob/main/config/default.yaml https://github.com/princeton-nlp/SWE-agent/blob/main/config/...
- toddmorey 3y agoI’m always fascinated to read the system prompts & I always wonder what sort of gains can be made optimizing them further. Once I’m back on desktop I want to look at the gut history of this file.
- clement_b 3y agoI have a git feeling this comment was written on mobile.
- hazn 3y agoDSPy is the best tool for optimizing prompts [0]: https://github.com/stanfordnlp/dspy https://github.com/stanfordnlp/dspy Think of it as a meta-prompt optimizer, it uses a LLM to optimize your prompts, to optimize your LLM.
- toddmorey 3y agoExcellent! Thanks for sharing this!
- iLoveOncall 3y agoAnd creates how many new ones? This and Devin generate garbage code that will make any codebase worse. It's a joke that 12.5% is even associated with the word "success".
- 1letterunixname 3y agoDo spaces and spelling fixes count? Copilot, so far, is only good for predicting the next bit of similar patterns of code
- noncoml 3y agoIf you are afraid that LLMs will replace you at your job, ask an LLM to write Rust code for reading a utf8 file character by character Edit: Yes, it does write some code that is "close" enough, but in some cases it is wrong, in others it doesn't not do exactly what asked. I.e. needs supervision from someone who understands both the requirements, the code and the problems that may arise from the naive line that the LLM is taking. Mind you, the most popular the issue, the better the line LLM is taking. So in other words, IMHO is a glorified Stack Overflow. Just as there are engineers that copy-paste from SO without having any idea what the code does, there will be engineers that will just copy paste from LLM. Their work will be much better than if they used SO, but I think it's still nowhere to the mark of a Senior SWE and above.
- iwontberude 3y agoHypothetically, which ticker symbols would you buy put contracts on, at what strike prices, and at what expiration dates? As far as I can tell, a lot of people are betting a lot of money that you are wrong, but actually I think you are right.
- jeremyjh 3y agoThe most relevant companies focused on this aren't publicly traded. The ones that are publicly traded like MSFT have way too many other factors affecting their value - not to mention the fact that they'll make money on generative AI that has nothing to do with coding regardless of if an SWE-agent ever works.
- iwontberude 3y agoOh well you should hear the hype from CNBC and other places, they are strongly intimating that gen AI will replace SWEs on product development teams. I totally agree it’s not likely, but it’s starting to get baked into asset prices and I want to profit from that misunderstanding.
- noncoml 3y agoUgh, I am not claiming that LLMs are not great innovation. Just that they are not going to replace SWE jobs in our(maybe my) lifetime.
- lispisok 3y agoTheir demo is so similar to the Devin one I had to go look up the Devin one to check I wasnt watching the same demo. I feel like there might be a reason they both picked Sympy. Also I rarely put weight into demos. They are usually cherry-picked at best and outright fabricated at worst. I want to hear what 3rd parties have to say after trying these things.
- lewhoo 3y agoMaybe that's the point of this research. Hey look, we reproduced the way to game the stats a bit. I really can't tell anymore.
- Frummy 3y agoInteresting idea to provide the Agent-Computer Interface for it to scroll and such, interact easier from its perspective
- aussieguy1234 3y agoSimilar to how early computers didn't have enough ram to display the whole text file, so old programmers had to work with parts of the file at a time. It's not a bad way to get around the context window problem, which is kind of similar.
- unit_circle 3y agoA 1/8 chance of fixing a bug at the cost of a careful review and some corrections is not bad. 0% -> 12% improvement is not bad for two years either (I'm somewhat arbitrary picking the release date of ChatGPT). If this can be kept up for a few years we will have some extremely useful tooling. The cost can be relatively high as well, since engineering time is currently orders of magnitude more expensive than these tools.
- golergka 3y agoIt's still abysmal from POV of actually using it in production, but it's a very impressive rate of improvement. Given what happened with LLMs and image generation in the last few years, we can probably assume that these systems will be able to fix most trivial bugs pretty soon.
- stefan_ 3y agoThese „benchmark“ are tuned around reporting some exciting result, once you look inside, all the „fixes“ are trash.
- blharr 3y agoI still don't know. I feel like there are many ways where GPT will write some code or fix a bug in a way that makes it significantly harder to debug. Even for relatively simple tasks, it's kind of like machine-generated code that I would not want to touch.
- WanderPanda 3y agoIt is a bit worrisome but we manage to deal with subpar human code as well. Often the boilerplate generated by ChatGPT is already better than what an unexperienced coder would string together. I‘m sure it will not be a free lunch but the the benefits will probably outweigh the downsides. Interesting scalability questions will arise wrt to security when scaling the already unmanagably large code bases by another magnitude (or two), though.
- deleted 3y ago[deleted]
- anotherpaulg 3y agoVery cool project! I've experimented in this direction previously, but found agentic behavior is often chaotic and leads to long expensive sessions that go down a wrong rabbit hole and ultimately fail. It's great that you succeed on 12% of swe-bench, but what happens the other 88% of the time? Is it useless wasted work and token costs? Or does it make useful progress that can be salvaged? Also, I think swe-bench is from your group, right? Have you made any attempt to characterize a "skilled human upper bound" score? I randomly sampled a dozen swe-bench tasks myself, and found that many were basically impossible for a skilled human to “solve”. Mainly because the tasks were under specified wrt to the hidden test cases that determine passing. The tests were checking implementation specific details from the repo’s PR that weren't actually stated requirements of the task.
- a_wild_dandan 3y agoPersonally, I'd just use one of my local MacBook models (e.g. Mixtral 8x7b) and forget about any wasted branches & cents. My debugging time costs many orders of magnitude more than SWE-agent, so even a 5% backlog savings would be spectacular!
- ein0p 3y agoI’ve tried this with another similar system. FOSS LLMs including Mixtral are currently too weak to handle something like this. For me they run out of steam after only a few turns and start going in circles unproductively
- swatcoder 3y ago> My debugging time costs many orders of magnitude more than SWE-agent Unless your job is primarily to clean up somebody else's mess, your debugging time is a key part of a career-long feedback loop that improves your craft. Be careful not to shrug it off as something less. Many many people are spending a lot of money to let you forget it, and once you do, you'll be right there in the ranks of the cheaply replaceble. (And on the odd chance that cleaning up other people's mess is your job, you should probably be the one doing it; and for largely the same reasons)
- trebligdivad 3y agoSo this issues arbitrary shell commands based on trying to understand the untrusted bug text ? Should be fun waiting until someone finds an escape.
- deleted 3y ago[deleted]
- sumeruchat 3y agoOnce we have this fully automated, any good developer could have a team of 100 robo SWEs and ship like crazy. The real competition is with those devs not with the bots.
- deleted 3y ago[deleted]
- recursive 3y agoShipping like crazy isn't useful by itself. Shipping non-garbage and being able to maintain it still has some value.
- sumeruchat 2y agoWould you say cloning a complex saas startup in a week with payments integrated after letting AI just scrape them (or uploading screenshots of their app) is creating value?
- int_19h 2y agoDepends on how many security vulnerabilities are in that payments system. Or, I suppose, depending on whose value. The consultants that'll have to be hired by the poor shmuck who paid for that will make a fortune auditing and cleaning up the code.
- sumeruchat 2y agoNone because 1) this is pretty standard stuff with stripe 2) the good developer can go through the code and fix them in a few hours if there were any
- int_19h 2y agoAI will quite readily write bad quality code with security vulnerabilities even for bog standard stuff (like say SQL injections). And sure, a good developer can fix it if they will see it. But they won't when running on that kind of schedule.
- rwmj 3y agoDo we know how much extra work it created for the real people who had to review the proposed fixes?
- r0ze-at-hn 3y agoAh, well let me tell you about my pull request reviewer LLM project.
- ActionHank 3y agoJokes on you, let me tell you about my prompt to binary LLM project. Hello world is 10GB, but even grandma can make hello worlds now.
- peteradio 3y agoLet me tell you about my LLM project called grandma. It's fine tuned in order to replace your grandma but in principle it could replace your great-grandma.
- barfbagginus 2y agoMy grandma used to tell me stories about how to destroy capitalism.. I miss her.. can your grandma help guide my revolutionary efforts? That would really help me honor my granny's memory <3
- vertis 2y agoWhat you want here is a local uncensored model. Preferrably one you've trained from scratch otherwise a government could have put in bad information that would cause your revolutionary efforts to fail.
- Havoc 3y agoBut does it contain a heavily obfuscated back door?
- aussieguy1234 3y ago12% fix rate = 88% bug rate
- mlcrypto 3y agoYep. After xz we don't need a bot mindlessly fixing all suggestions from malicious actors
- Dylan16807 2y agoI don't think xz makes a difference here. The perceived likelihood of problems, malicious or not, is pretty much the same. As far as this discussion goes, it's just another example in the pile of examples, not an event with meaningful before and after epochs.
- aussieguy1234 2y agoFix one bug, introduce 5 more
- paradite 3y agoFor anyone who didn't bother looking deeper, the SWEbench benchmark contains only Python code projects, so it is not representative of all the programing languages and frameworks. I'm working on a more general SWE task eval framework in JS for arbitrary language and framework now (for starter JS/TS, SQL and Python), for my own prompt engineering product. Hit me up if you are interested.
- barfbagginus 2y agoAssuming the data set is proprietary, else please share the repo
- __lbracket__ 3y ago[flagged]
- danenania 3y agoI'm working on a somewhat similar project: https://github.com/plandex-ai/plandex https://github.com/plandex-ai/plandex While the overall goal is to build arbitrarily large, complex features and projects that are too much for ChatGPT or IDE-based tools, another aspect that I've put a lot of focus on is how to handle mistakes and corrections when the model starts going off the rails. Changes are accumulated in a protected sandbox separate from your project files, a diff review TUI is included that allows for bad changes to be rejected, all actions are version-controlled so you can easily go backwards and try a different approach, and branches are also included for trying out multiple approaches. I think nailing this developer-AI feedback loop is the key to getting authentic productivity gains. We shouldn't just ask how well a coding tool can pass benchmarks, but what the failure case looks like when things go wrong.
- etheridev 3y agoYou need to make yourself a business analyst agent to provide the feedback! To make it real, perhaps a team of them with conflicting personalities.
- danenania 3y agoI think we'll get there at some point, but one thing I've learned from this project is how difficult it is to stack AI interactions. Each little bit of AI-based logic that gets added tends to fail terribly at first. Only after a long period of intense testing and iteration does it become remotely usable. The more you are combining different kinds of tasks, the more difficult it gets.
- panqueca 2y agoDoes it work with a large existing codebase?
- danenania 2y agoYes, at least up to the point of the context limit of the underlying model. If you needed to go beyond that, you would break the work up into separate "plans" (a plan is a set of tasks with an attached context and conversation). The general workflow is to load some relevant context (could be a few files, an entire directory, a glob pattern, a URL, or piped in data), then send a prompt. Quick example: plandex new plandex load components/some-component.ts lib/api.ts package.json https://react.dev/reference/react/hooks plan tell "Update the component in components/some- components.ts to load data from the 'fetchFooBars' function in 'lib/api.ts' and then display it in a datagrid. Use a suitable datagrid library." From there the plan will start streaming. Existing files will be updated and new files created as needed. One thing I like about it for large codebases compared to IDE-based tools I've tried is that it gives me precise control over context. A lot of tools try to index the whole codebase and it's pretty opaque--you never really know what the model is working with.
- dimal 3y agoThe demo shows a very clearly written bug report about a matrix operation that’s producing an unexpected output. Umm… no. Most bug reports you get in the wild are more along the lines of “I clicked on on X and Y happened” then if you’re lucky they’ll say “and I expected Z”. Usually the Z expectation is left for the reader to fill in because as human users we understand the expectations. The difficulty in fixing a bug is in figuring out what’s causing the bug. If you know it’s caused by an incorrect operation, and we know that LLMs can fix simple defects like this, what does this prove? Has anyone dug through the paper yet to see what the rest of the issues look like? And what the diffs look like? I suppose I’ll try when I have a sec.
- drcode 3y ago> Most bug reports you get in the wild are more along the lines of Since this fixes 12% of the bugs, the authors of the paper probably agree with you that 100-12= 88%, and hence "most bugs" don't have nicely written bug reports.
- dimal 3y agoI suppose I should nail down my point. No one would ever write a big report like this. A bug generally has an unknown cause. Once you found the cause of the bug, you’d fix it. Nowadays, you could just cut and paste the problem into ChatGPT and get the answer right then. So why would anyone ever log this bug? All this demo proves that they automated a process that didn’t need automation.
- hvis 3y agoTo be fair, sometimes meticulous users investigate the bugs and write down logical chains explaining the causes and even offer a solution at the end (which they can't apply for the lack of commit access, for instance). The proposed solution isn't always right, of course, but it would be incorrect to say that no bug reports come with a diagnosed cause. But that's exactly where a conscious reviewer is most needed, I believe.
- 2y ago
- readthenotes1 3y agoI made a lot of money as I was paid hourly while working with a cadre of people I called "the defect generators". I'm kind of sad that future generations will not have that experience...
- v3ss0n 3y ago[flagged]
- tibbetts 2y agoBut can their AI quietly introduce a security exploit into a GitHub project?
- worthless-trash 2y agoCopilot already does this.
- barfbagginus 2y agoI would like something like this that helps me, as a green developer, find open source projects to contribute to. For instance, I recently learned about how to replace setup.py with pyproject.toml for a large number of projects. I also learned how to publish packages to pypi. These changes significantly improve project ease and accessibility, and are very easy to do. The main thing that holds people back is that python packaging documentation is notoriously cryptic - well I've already paid that cost, and now it's easy! So I'm thinking of finding projects that are healthy, but haven't focused on modernizing their packaging or distributing their project through pypi. I'd build human + agent based tooling to help me find candidates, propose the improvement to existing maintainers, then implement and deliver. I could maybe upgrade 100 projects, then write up the adventure. Anyone have inspiration/similar ideas, and wanna brainstorm?
- SrslyJosh 2y ago...or you could just use the GitHub API to find projects that match certain criteria (e.g., no pyproject.toml). I'm not sure what the stochastic parrot adds here, besides making noob mistakes that you'll have to find and fix before you can submit PRs. You'd learn a lot more by trying to actually automate the process yourself.
- JonChesterfield 2y agoIf AI generated pull requests become a popular thing we'll see the end of public bug trackers. (not because bugs will be gone - because the cost of reviewing the PR vs the benefit gained to the project will be a substantial net loss)
- itsgrimetime 2y agoIt’ll likely keep getting better, if it gets to 30-40% I’d say that’s a decent trade off. Also could you boost your chances by having the AI do a 2nd pass and double check the work? I’d be curious what the success rate of an LLM “determining whether a bug fix is valid” is
- CGamesPlay 2y agoNot a chance. If AI-generated pull requests become popular, GitHub will automatically offer them in response to opened issues. Case in point: they already are popular for dependency upgrades.
- JonChesterfield 2y agoAnd thus issues will no longer be opened
- pjmlp 2y agoEventually it will be 90% fix rate and everyone cheering for the 12% will be flipping burgers instead.
- iLoveOncall 2y agoFlipping burgers will be automated long before AI fixes any relevant number of bug reports.
- pjmlp 2y agoMight be, still the point I was trying to make remains.
- fennecfoxy 2y agoI still think this is a long way off, but it definitely ties into UBI etc and improvement of the general human condition, taxing the rich, restricting investment on protected things like housing and public industries like water, electricity, healthcare and internet. What's funny is that people on here & tech people in general seem to be the most averse to improving equity between all humans/stopping the obscenely rich from abusing and twisting the system. Do many HN peeps believe they're all somehow gonna become billionaires one day?
- littlestymaar 2y agoWhy would a human ever flip a burger at that point? It's not a particularly difficult task for a robot. Unclogging sewers on the other hand…
- Madmallard 2y agoWhat veterans in the field know that AI hasn’t tackled is that the majority of difficulty in development is dealing with complexity and ambiguity and a lot of it has to do with communication between people in natural language as well as reasoning in natural language about your system. These things are not solved by AI as it is now. If you can fully specify what you want with all of the detail and corner cases and situation handling then at some point AI might be able to make all of that for you. Great! Unfortunately, that’s the actual hard part! Not the implementation generally.