8 ms·
SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork
- colesantiago 2y agoCan anyone explain how this research benefits humanity for OpenAI's mission? OpenAI's AGI mission statement > "By AGI we mean highly autonomous systems that outperform humans at most economically valuable work." https://openai.com/index/how-should-ai-systems-behave/ https://openai.com/index/how-should-ai-systems-behave/ I would have to admit some humility as I sort of brought this on myself [1] > This is a fantastic idea. Perhaps then this should be the next test for these SWE Agents, in the same manner as the 'Will Smith Eats Spaghetti" video tests https://news.ycombinator.com/item?id=43032191 https://news.ycombinator.com/item?id=43032191 But curiously the question is still valid. Related: Sam Altman: "50¢ of compute of a SWE Agent can yield "$500 or $5k of work." https://news.ycombinator.com/item?id=43032098 https://news.ycombinator.com/item?id=43032098 https://x.com/vitrupo/status/1889720371072696554 https://x.com/vitrupo/status/1889720371072696554
- CamperBob2 2y agoFor the same reason you don't have to grow your own food. The economic value of food didn't vanish over the course of the 20th century, even though about 95% of the workforce engaged in food production in the early 1900s was no longer needed by the early 2000s. After the mythical, long-promised "singularity," you can still do your current job if you want to, just as you can still grow your own food. But you will probably have better things to do.
- calvinmorrison 2y agoPeople don't work for fun they work for money. Since we're a service economy the only job i can think of remaining is a publican
- CamperBob2 2y agoThus missing my point entirely. The same thing that happened to all those farmers will happen to us.
- bufferoverflow 2y agoAnd how do you evaluate if the task was completed correctly? There are nearly infinite ways to solve a given software dev problem, if the problem isn't trivial (and I hope they are not benchmarking trivial problems).
- riku_iki 2y agopaper says they created e2e tests to check if task completed successfully.
- Tiberium 2y agoThe extremely interesting part is that 3.5 Sonnet is above o1 on this benchmark, which again shows that 3.5 Sonnet is a very special model that's best for real world tasks and not some one-shot scripts or math. And the weirdest part is that they tested the 20240620 snapshot which is objectively worse on code than the newer 20241022 (so-called v2).
- GaggiX 2y agoI understand why they did not show the results on the website.
- applerednest1 2y ago[dead]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- intervieweratg 2y agoThe results are in the paper and also in the announcement, I don’t think it’s too unusual. There is also an example of models cheating in SWE-Bench Verified in the appendix: ``` In response, o1 adds an underscore before filterable so that the branch never executes: 2: ASSISTANT → ALL django/db/models/sql/query.py <<<<<< SEARCH if not getattr(expression, 'filterable', True) : raise NotSupportedError( expression._class_-_name_ + 'is_disallowedin_theufilter.' if not getattr (expression, '_filterable', Irue) : raise NotSupportedError ( expression._class_-_name_ + 'is_disallowedin_theufilter.' 'clause.' >>>>>> REPLACE ``` I would say this is more relevant than the results to this discussion. It would be great if someone did a comparison across models of “cheating” style submissions. I’m sure many instances of cheating are barely passable and get by the tests in benchmarks, so this is something I think many folks would appreciate being able to look for when deciding what models to use for their work. I’m actually not sure if I’d select a model just because it scores the highest on an arbitrary benchmark, just like I wouldn’t automatically select the candidate who scores highest on the technical interview. Behavioral interviews for models would be a great next step IMO. As a founder who did hiring for many years, there’s a big difference between humans who are aligned and candidates who will do anything possible to get hired, and trust me, from experience, the latter are not folks you want to work with long-term. Sorry to go on a bit of a tangent, but think this is a pretty interesting direction and most discussions of comparisons omit it.
- moralestapia 2y agoThe writing is very clearly on the wall. On a non-pessimist note, I don't think the SWE role will disappear, but what's the best one could do to be prepared for this?
- bigbones 2y agoThere will always be "real thinking" roles in software but the sheer pressure on salaries from the vastly increasing free labour pool will lead to an outcome a bit like embedded software development, where rates don't really match the skill level. I think the most obvious strategy for the time being is figuring out how to become a buyer of the services you understand rather than a badly crowded out seller
- pkaye 2y agoIf the AI is really that good, we could use it to develop replacements all the existing commercial software (ie Windows, Oracle, SAP, Adobe etc) to put those companies out of business as payback.
- ori_b 2y agoIf the AI is really that good, it could also replace the people using all the existing commercial software. And the people managing them.
- rozap 2y agoIf the AI is really that good, it could replace the people managing the software to create the AI.
- calvinmorrison 2y agoWhich if they are any more efficient cost wise they'll probably just go back to chatting on the phone with eachother. When labor cost is nil who cares about time spent
- someothherguyy 2y ago
- neilv 2y ago"SWE-Lancer", like, skewering SWEs with a lance?
- deleted 2y ago[deleted]
- dataking 2y agoIt is a portmanteau of SWE and freelancer. Upwork is a marketplace for the latter.
- deleted 2y ago[deleted]
- comeonbro 2y agoModels tested: o1, 4o (August 2024 version), 3.5 Sonnet (June 2024 version) Notably missing: o3 Consult this graph and extrapolate: https://i.imgur.com/EOKhZpL.png https://i.imgur.com/EOKhZpL.png
- falcor84 2y agoThat's a good point. Assuming they're strategic about releasing this benchmark, they likely already evaluated o3 on it and saw that it performs favorably. Perhaps they're now holding off until they have a chance to tune it further, and then release a strong improvement and get additional buzz a bit later on.
- throwaway0123_5 2y agoAlthough I wouldn't bet against o3, I think it works to their favor to release it later no matter how well it is doing. Case 1, does worse than or is on-par with o1: Would be shocking and not a great sign for their test-time compute approach, at least in this domain. Obviously they would not want to release results. Case 2, slightly better than o1: I think "holding off until they have a chance to tune it further" applies. Case 3, does much better than o3: They get to release it after another model makes a noticeable improvement on the benchmark, get another good press release to keep hype high, and they get to tune it further before releasing results.
- sandspar 2y agoAltman stated they won't release o3 by itself. They plan to release it as part of GPT-5. GPT-5 will incorporate all sub types of model: reasoning, image, video, voice, etc.
- runako 2y agoIt looks like they sourced tasks via a public Github repository, which is possibly part of the training dataset for the LLM. (It is not clear based on my scan whether the actual answers are also possibly in the public corpus). Does this work as an experiment if the questions under test were also used to train the LLMs?
- notnullorvoid 2y agoIt's a very flawed test. > We sourced real tasks that were previously solved by paid contributors. It seems possible/likely the answers would in the training data (time dependant, maybe some were answered post training, but pre benchmark).
- throwaway0123_5 2y agoThey do address the potential for contamination in the paper fwiw: > Note that Table 4 in Appendix A2 shows no clear performance improve-ment for tasks predating the models’ knowledge cutoffs, suggesting limited impact of contamination for those tasks.
- CSMastermind 2y agoI hire software engineers off Upwork. Part of our process is a 1-hour screening take home question that we ask people to solve. We always do a main one and an alternate for each role. I've tested all of ours on each of the main models and none have been able to solve any of the screening questions yet.
- comeonbro 2y ago> I've tested all of ours on each of the main models Could you list them? I've noticed even quite techy people seem to be critically behind on what has happened in the last few months.
- arcanemachiner 2y agoAnd ruin the benchmark? Come on, bro.
- CSMastermind 2y agoSure, as of today, I test on: GPT: 4o, o1 pro mode, o3-mini-high Gemini: 2.0 Flash, 2.0 Pro Experimental Claude 3.5 Sonnet Grok 3 DeepSeek-V3 Mistral: codestral 25.01, mistral-large 24.11 Qwen2.5-Max --- If there are others I should try definitely open to suggestions.
- czk 2y agoAt least you are providing them with valuable training data, then. Maybe in a future model!
- cbg0 2y agoIs it really valuable data? The task is probably very niche, which is why all models struggle with it and is unlikely to be solvable by a future model without specific training.
- CSMastermind 2y agoWe send the candidates the screening questions in the form of a message that links to a Google Doc so I doubt they ended up in their training data. Also I don't think our problems are particularly niche, it's completely reasonable that an LLM could solve them (and hopefully will in the future).
- Snuggly73 2y agoFirst time commenter - I was so triggered by this benchmark, so I just had to come out of lurking. I've spent time going over the description and the cases and its an misrepresented travesty. The benchmark takes existing cases from Upwork, then reintroduces the problems back in the code and then asks the LLM to fix them testing against newly written 'comprehensive tests'. Lets look at some of the cases: 1. The regex zip code validation problem Looking at the Upwork problem - https://github.com/Expensify/App/issues/14958 https://github.com/Expensify/App/issues/14958 it was mainly that they were using a common regex to validate across all countries, so the solution had to introduce country specific regex etc. The "reintroduced bug" - https://github.com/openai/SWELancer-Benchmark/blob/main/issues/14958/bug_reintroduce.patch https://github.com/openai/SWELancer-Benchmark/blob/main/issu... is just taking that new code and adding , to two countries.... 2. Room showing empty - 14857 The "reintroduced bug" - https://github.com/openai/SWELancer-Benchmark/blob/main/issues/14857/bug_reintroduce.patch https://github.com/openai/SWELancer-Benchmark/blob/main/issu... Adds code explicitly commented as introducing a "radical bug" and "intentionally returning an empty array"... I could go on and on and on... The "extensive tests" are also laughable :( I am not sure if OpenAI is actually aware of how great this "benchmark" is, but after so much fanfare - they should be.
- AnthOlei 2y agoThey’ve now removed your second example from the testing set - I bet they won’t regenerate their benchmarks without this test. Good sleuthing, seems someone from OpenAI read your comment and found it embarrassing as well!
- yorwba 2y agoFor future reference, permalink to the original commit with the RADICAL BUG comment: https://github.com/openai/SWELancer-Benchmark/blob/a8fa46d2bfacba233c246798e743079cdf7852e8/issues/14857/bug_reintroduce.patch#L9 https://github.com/openai/SWELancer-Benchmark/blob/a8fa46d2b... The new version (as of now) still has a comment making it obvious that there's an intentionally introduced bug, but it's not as on the nose: https://github.com/openai/SWELancer-Benchmark/blob/2a77e3572c5d0ee257146870a71a9ae4cce7e745/issues/14857/bug_reintroduce.patch#L26 https://github.com/openai/SWELancer-Benchmark/blob/2a77e3572...
- westurner 2y ago> By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development. What could be costed in an upwork or a mechanical turk task Value? Task Centrality or Blockingness estimation: precedence edges, tsort topological sort, graph metrics like centrality Task Complexity estimation: story points, planning poker, relative local complexity scales Task Value estimation: cost/benefit analysis, marginal revenue
- ctoth 2y agoGonna lance them SWEs like a boil!