6 ms·
Exhausted man defeats AI model in world coding championship
- ChrisMarshallNY 1y agoI'm old enough to remember being taught the Ballad of John Henry...
- esseph 1y agoThe article mentions it...
- ChrisMarshallNY 1y agoYup. I was talking about how they taught it to us, in school. It actually had an emotional place in my heart. For some reason, I found the story compelling.
- coldtea 1y agoWell, millions did, that's why it's a classic!
- ChrisMarshallNY 1y agoI suspect it may not be taught, anymore, though. I seem to encounter cultural milestones, that are no longer there, every day.
- tyre 1y agoI learned it in elementary school in the late 90s
- coldtea 1y agoThere's no shared culture anymore around those things. The majority of people don't listen to anything outside the hits of the year, and for younger people there are no community events, there's no shared platform with any variety (radio is dead, streaming plays either popular hits or your own echo-bubble), schools doesn't teach them (they're more likely to teach some song of the latest shit pop celebrity), and there are in general no mainstream institutions that keep these alive. But before our current 60-seconds-memory-span such folk songs were part of a canon known and loved for well over a century.
- ASalazarMX 1y agoI found it sad, because he dies at the end to prove he could beat the machine once, but the machine could keep producing its lesser output every day after his death. He gave it all for a Pyrrhic victory.
- 317070 1y agoDoes someone know the problem/challenge being solved?
- burkaman 1y agohttps://atcoder.jp/contests/awtf2025heuristic/tasks/awtf2025heuristic_a https://atcoder.jp/contests/awtf2025heuristic/tasks/awtf2025...
- hyperhello 1y agoI'm at a complete loss to discern why this would be a useful task to solve. It seems like the equivalent of elementary schoolers saying "OK, if you're so smart, what's 9,203,278,023 times 3,333,300,209?"
- LeoPanthera 1y agoIt's a problem that has no perfect solution, only incremental improvements. So it's really not like your example at all.
- lovich 1y agoI feel like if that question was asked when calculators were invented, and someone was claiming humans were still better at arithmetic than machines, that it would be appropriate. I was surprised reading through this problem that the machine solved it well at all. I get that it’s a leet code style question but it’s got a lot of specifics and I assumed the corpus of training data on optimizing this type of problem was several orders of magnitude too small to train an LLM on and have good results.
- margalabargala 1y agoIt's more the equivalent of "why would anyone race the 400m on a standard track, you just wind up back where you started!"
- TrackerFF 1y agoSay you have a bunch of warehouse robots, some which work on different sections in the warehouse. Maybe one section has less things to do, while another section has more things to do - and thus needs more help. So you need to move a bunch of robots there, in groups. Something like that.
- chmod775 1y agoTen hours is a decent amount of time, so I'm not too surprised the human won. LLMs don't really tend to improve the longer they get to chew on a problem (often the opposite in fact). The LLM was probably getting nowhere trying to improve after the first few minutes.
- shoguse72 1y ago> The LLM was probably getting nowhere trying to improve after the first few minutes. How did you come to that conclusion from the contents of the article? The final scores are all relatively close. How could that happen if the ai was floundering the whole time? Just a good initial guess?
- coldtea 1y ago>How could that happen if the ai was floundering the whole time? Just a good initial guess? Yes, that and marginal improvements over it.
- satyrun 1y agoI would think the LLM though is not trying one solution for 10 hours like a human. I would assume the LLM is trying an inhuman number of solutions and the best one was #2 in this contest. Impressive by the human winner but good luck on that in 2026.
- asey 1y agoOn the livestream (perhaps elsewhere?) you can watch the submissions and scores come in over time. The LLM steadily increased (and sometimes decreased) it's score over time though by the end did seem to hit a lacuna. You could even see it try out new strategies (with walls e.g.) which didn't appear until about half-way through the competition.
- ac130kz 1y agoYeap, self-reinforcement learning is missing in LLMs.
- crmi 1y agoReally feels like it could be an onion title.
- geephroh 1y agoHa! Came here to say the same thing...
- TrackerFF 1y agoNow imagine where we'll be in 10 years, and where we were 10 years ago. Things move, fast.
- deleted 1y ago[deleted]
- nuifldpei 1y agoI despise the company that competed, but I feel obligated to acknowledge that headline buries the lede that their bot got SECOND place, and their 2nd place was closer to first than 3rd was to 2nd. Are the submissions available online without needing to become a member of AtCoder? I want to see what these 'heuristic' solutions look like. Is it just that the ai precomputed more states and shoved their solutions in as the 'heuristic' or did it come up with novel, more broad, heuristics? Did the human and ai solutions have overlapping heuristics?
- briandw 1y agoThis is a real modern day John Henry story, except John Henry dies in the end. https://en.wikipedia.org/wiki/John_Henry_(folklore) https://en.wikipedia.org/wiki/John_Henry_(folklore)
- thegeomaster 1y agoMentioned in TFA as well.
- briandw 1y agoGuess I should read more than the summary :)
- baerrie 1y agoHow was the model operated? Was it someone prompting it continuously or was it just given the initial prompt?
- flanbiscuit 1y agoso many things First, there's a world coding championship?! Of course there is. There's a competition for anything these days. Why is he exhausted? > The 10-hour marathon left him "completely exhausted." > ... noting he had little sleep while competing in several competitions across three days. "I'm completely exhausted. ... I'm barely alive." oh! That's a lot. > beating an advanced AI model from OpenAI ... > On Wednesday, programmer Przemysław Dębiak (known as "Psyho"), a former OpenAI employee, Interesting that he used to work there. > Dębiak won 500,000 yen JPY 500,000 -> USD 3367.20 -> EUR 2889.35 I'm guessing it's more about the clout than it is about the payment, because that's not a lot of money for the effort spent
- hungmung 1y ago> I'm guessing it's more about the clout than it is about the payment Yeah I'm not in tech but I've seen his handle like 3 times today already, so he's definitely got recognition.
- magicalist 1y ago> I'm guessing it's more about the clout than it is about the payment to be fair he also said > "Honestly, the hype feels kind of bizarre," Dębiak said on X. "Never expected so many people would be interested in programming contests."
- chiwilliams 1y agoHe's retired, so I'm guessing more about the clout. Or even just "love of the game"? He had a fairly popular tweet thread a couple years back where he wrote out 80 tips for competitive programming -- that feels less likely to be clout based
- atleastoptimal 1y agoRemember this is the worst AI will ever be from here on out. Models are only going to get better, faster, cheaper, more accessible and more easily deployable. I think people need to realize that just because an AI model fails at one point, or some certain architecture has common failure modes, that billions of dollars are poured into correcting those failures and improving in every economically viable domain. Two years ago AI video looked like a garbled 140p nightmare, now it's higher quality video than all but professional production studios could make. AI agents don't get tired. They don't need to sleep. They don't require sick days, parental leave, or PTO. They don't file lawsuits, they don't share company secrets, they don't disparage, deliberately sandbag to get extra free time, whine, burn out or go AWOL. The best AI model/employee is infinitely replicatable, and can share its knowledge with other agents perfectly and clone itself arbitrarily many times, and it doesn't have a clash of egos working with copies of itself, it just optimizes and is refit to accomplish whatever task its given. All this means is that gradually the relative advantage of humans in any economically viable domain will predictably trend towards zero. We have to figure out now what that will mean for general human welfare, freedom and happiness, because barring extremely restrictive measures on AI development or voluntary cessation by all AI companies, AGI will arrive.
- reducesuffering 1y agoExactly. The inability of people to extrapolate towards the future and foresee second-order effects is astounding. We've seen this in climate change and we've just seen this in COVID. The ones with foresight are warning about the massive upheaval coming. It's time for people to shake away their preconceived notions, look at the situation with fresh eyes, and deeply think about what the technology diff from 5 years ago to today, means for 5 years from now.
- xienze 1y ago> Exactly. The inability of people to extrapolate towards the future and foresee second-order effects is astounding. On a related note, many people also assume that just because something has been trending exponential that it will _continue_ to do so...
- abound 1y agoI'm guessing he didn't have access to any LLMs while competing, but I think a "centaur" approach probably would have outperformed both "only human" and "only LLM" competitors. Reading through the challenge, there's a lot of data modelling and test harness writing and ideating that an LLM could knock out fairly quickly, but would take even a competitive coder some time to write (even if just limited by typing speed). That'd give the human more time to experiment with different approaches and test incremental improvements.
- chiwilliams 1y agoHe did use a little autocomplete apparently, but used [Vscode](https://x.com/jacob_posel/status/1945585787690738051 https://x.com/jacob_posel/status/1945585787690738051). And it's not against the rules to use LLMs apparently in the competition. (https://atcoder.jp/posts/1495 https://atcoder.jp/posts/1495). I'd be curious what other competitors used.
- abound 1y agoInteresting, thanks for the links! I had read this part of the article: > All competitors, including OpenAI, were limited to identical hardware provided by AtCoder, ensuring a level playing field between human and AI contestants. And assumed that meant a pretty restricted (and LLM-free) environment. I think their policy is pretty pragmatic.
- tom_m 1y agoHe misspelled psycho.
- hgs3 1y agoThis is interesting, but aren't "coding competitions" about writing small leetcode programs from a prompt? I would expect the AI to excel at that.
- evolve2k 1y ago> All competitors, including OpenAI, were limited to identical hardware provided by AtCoder, ensuring a level playing field between human and AI contestants. The power of this chosen hardware will very much determine how well the AI performs. Everyone receiving the same computer does not make the competition inherently fair. It’s likely that human competitors would outperform the AI on hardware that is even a few years old.