4 ms·
Top AI models fail at >96% of tasks
- codexon 8mo agoThis paper creates a new benchmark comprised of real remote work tasks sourced from the remote working website Upwork. The best commercial LLMs like Opus, GPT, Gemini, and Grok were tested. Models released a few days ago, Opus 4.6 and GPT 5.3, haven't been tested yet, but given the performance on other micro-benchmarks, they will probably not be much different on this benchmark.
- Venn1 8mo agoChatGPT: when you want spellcheck to argue with you.
- zb3 8mo agoYou think they don't? You think AI can replace programmers, today? Then go ahead and use AI to fix this: https://gitlab.gnome.org/GNOME/mutter/-/issues/4051 https://gitlab.gnome.org/GNOME/mutter/-/issues/4051
- stoneforger 8mo agoRewrite it in react it will.
- deleted 8mo ago[deleted]
- tessitore 8mo agoThis post really should be edited to say 96% of tasks posted on Upwork. Since we would all expect that to happen.
- deleted 8mo ago[deleted]
- scotty79 8mo agoKinda sus that least known model did best and none of the more recent models were tested. Capabilities grow very fast. So things that now routinely succeed rarely ever succeeded even half a year ago.
- rsynnott 8mo agoI mean performance is so bad across the board that this is likely essentially random. Monkeys accidentally doing a bit of Shakespeare.
- ben_w 8mo agoThat's wildly overestimating what monkeys can do on a typewriter. It takes a lot to just be mediocre. Which, don't get me wrong, I'll agree current ML is, it's just that "mediocre" is an incomprehensibly huge step up from "random".