13 ms·
DeepMind and OpenAI win gold at ICPC
https://x.com/MostafaRohani/status/1968360976379703569 https://x.com/MostafaRohani/status/1968360976379703569
https://x.com/HengTze/status/1968359525339246825 https://x.com/HengTze/status/1968359525339246825
- NitpickLawyer 1y agoSo this year SotA models have gotten gold at IMO, IoI, ICPC and beat 9/10 humans in that atcoder thing that tested optimisation problems. Yet the most reposted headlines and rethoric is "wall this", "stangation that", "model regression", "winter", "bubble", doom etc.
- sixtram 1y agoThe last time I asked for a code review from AI was last week. It added (hallucinated) some extra lines to the code and then marked them as buggy. Yes, it beats humans at coding — great!
- CamperBob2 1y agoWhat's "It?" What was your prompt?
- tech_ken 1y agoIn 2015 SotA models blew past all expectations for engine performance in Go, but that didn't translate into LLM-based Code agents for another ~7 years (and even now the performance of these is up for debate). I think what this shows is that humans are extremely bad at understanding what problems are "hard" for computers; or rather we don't understand how to group tasks by difficulty in a generalizable way (success in a previously "hard" domain doesn't necessarily translate to performance in other domains of seemingly comparable difficult). It's incredibly impressive how these models perform in these contests, and certainly demonstrates that these tools have high potential in *specific areas* , but I think we might also need to accept that these are not necessarily good benchmarks for these tools' efficacy in less structured problem spaces. Copying from a comment I made a few weeks ago: > I dunno I can see an argument that something like IMO word problems are categorically a different language space than a corpus of historiography. For one, even when expressed in English language math is still highly, highly structured. Definitions of terms are totally unambiguous, logical tautologies can be expressed using only a few tokens, etc. etc. It's incredibly impressive that these rich structures can be learned by such a flexible model class, but it definitely seems closer (to me) to excelling at chess or other structured game, versus something as ambiguous as synthesis of historical narratives. edit: oh small world! the cited comment was actually a response to you in that other thread :D
- NitpickLawyer 1y ago> edit: oh small world the cited comment was actually a response to you in that other thread :D That's hilarious, we must have the same interests since we keep cross posting :D The thing with the go comparison is that alphago was meant to solve go and nothing else. It couldn't do chess with the same weights. The current SotA LLMs are "unreasonably good" at a LOT of tasks, while being trained with a very "simple" objective: NTP. That's the key difference here. We have these "stochastic parrots" + RL + compute that basically solve top tier competitions in math, coding, and who knows what else... I think it's insanely good for what it is.
- tech_ken 1y ago> I think it's insanely good for what it is. Oh totally! I think that the progress made in NLP, as well as the surprising collision of NLP with seemingly unrelated spaces (like ICPC word problems) is nothing sort of revolutionary. Nevertheless I also see stuff like this: https://dynomight.substack.com/p/chess https://dynomight.substack.com/p/chess To me this suggests that this out-of-domain performance is more like an unexpected boon, rather than a guarantee of future performance. The "and who knows what else..." is kind of I'm getting: so far we are turning out to be bad at predicting where these tools are going to excel or fall short. To me this is sort of where the "wall" stuff comes from; despite all the incredible successes in these structured problem domains, nobody (in my personal opinion) has really unlocked the "killer app" yet. My belief is that by accepting their limitations we might better position ourselves to laser-target LLMs at the kind of things they rule at, rather than trying to make them "everything tools".
- tempusalaria 1y agoA lot of the current code and science capabilities do not come from NTP training. Indeed in seems in most language model RL there is not even process supervision, so a long way from NTP
- JohnKemeny 1y agoThere is a clear difference between what OpenAI manages to do with GPT-5 and what I manage to do with GPT-5. The other day I asked for code to generate a linear regression and it gave back a figure of some points and a line through it. If GPT-5, as claimed, is able to solve all problems in ICPC, please give the instructions on how I can reproduce it.
- simianwords 1y agoAre you using the thinking model or the non thinking model? Maybe you can share your chat.
- minimaxir 1y agoThe point of the GPT-5 model is that it is supposed to route between thinking/nonthinking smartly. Leveraging prompt hacks such as instructing it to "think carefully" to force routing to the thinking model go against OpenAI's claims.
- koakuma-chan 1y agoAre you sure? I thought you can only specify reasoning_effort and that's it.
- Workaccount2 1y agoJust select GPT5-thinking if you need anything done with competence. The regular gpt5 is nothing impressive and geared more towards regular daily life chatting.
- JohnKemeny 1y agoI prefer not to due to privacy concerns. Perhaps you can try yourself? I will say that after checking, I see that the model is set to "Auto", and as mentioned, used almost 8 minutes. The prompt I used was: Solve the following problem from a competitive programming contest. Output only the exact code needed to get it to pass on the submission server. It did a lot of thinking, including I need to tackle a problem where no web-based help is available. The task involves checking if a given tree can be the result of inserting numbers 1 to n into an empty skew heap, following the described insertion algorithm. I have to figure out the minimal and maximal permutations that produce such a tree. And I can see that it visited 13 webpages, including icpc, codeforces, geeksforgeeks, github, tehrantimes, arxiv, facebook, stackoverflow, etc.
- riku_iki 1y ago> So this year SotA models have gotten gold at IMO, IoI, ICPC > Yet the most reposted headlines and rethoric is "wall this", "stangation that", "model regression", "winter", "bubble", doom etc. this is narrow niche with high amount of training data (they all buy training data from leetcode), and this results are not necessary generalizable on overall industrial tasks
- jug 1y agoEven Sam Altman himself thinks we’re in a bubble, and he ought to have a good sense of the wind direction here. I think the contradiction here can be reconciled by how these tests don’t tend to run on the typical hardware constraints they need to be able do this at scale. And herein lies a large part of the problem as far as I can tell; in late 2024, OpenAI realized they had to rethink GPT-5 since their first attempt became too costly to run. This delayed the model and when it finally released, it was not a revolutionary update but evolutionary at best compared to o3. Benchmarks published by OpenAI themselves indicated a 10% gain over o3 for God knows how much cash and well over a year of work. We certainly didn’t have those problems in 2023 or even 2024. DeepSeek has had to delay R2, and Mistral has had to delay Mistral 3 Large, teased within weeks back in May. No word from either about what’s going on. DS is said to move more to Huawei and this is behind a delay but I don’t think it’s entirely clear it has nothing to do with performance issues. It would be more strange to _not_ have people speculate about stagnation or bubbles given these events and public statements. Personally, I’m not sure if stagnation is the right word. We’re seeing a lot,of innovation in toolsets and platforms surrounding LLM’s like Codex, Claude Code, etc. I think we’ll see more in this regard and that this will provide more value than the core improvements to the LLM’s themselves in 2026. And as for the bubble, I think we are in one but mostly because the market has been so incredibly hot. I see a bubble not because AI will fall apart but because there are too many products and services right now in a golden rush era. Companies will fail but not because AI suddenly starts failing us but due to saturation.
- kadushka 1y agoit was not a revolutionary update but evolutionary at best compared to o3 It is a revolutionary update if compared to the previous major release (GPT-4 from March 2023).
- rhetocj23 1y agoSam Altman proclaiming we are in a bubble benefits him. It lowers the price of potential targets for acquisitions. I bet you didnt think of that did you?
- KallDrexx 1y agoIt's important to look closely at the details of how these models actually do these things. If you look at the details of how Google got gold at IMO, you'll see that AlphaGeometry only relies on LLMs for a very specific part of the whole system, and the LLM wasn't the core problem solving system in play. Most of AlphaGeometry is standard algorithms at play solving geometry problems using known constraints. When the algorithmic system gets stuck, it reaches out to LLMs that were fine tuned specifically for creating new geometric constraints. So the LLM would create new geometric constraints and pass that back to the algorithmic parts to get it unstuck, and repeat. Without more details, it's not clear if this win is also the Gpt-5 and Gemini models we use, or specially fine-tuned models that are integrated with other non-LLM and non-ML based systems to solve these. Not being solved purely by LLM isn't a knock on it, but with the current conversations going on today with LLMs, these are heavily being marketed as "LLMs did this all by themselves", which doesn't match with a lot of the evidence I've personally seen.
- NitpickLawyer 1y agoAlphaGeometry/AlphaProof (the one you're thinking of, where they used LLMs + lean) was last year! and they "only" got silver. IMO gold results this year were e2e NLP.
- Workaccount2 1y ago>This achievement is a significant advance over last year’s breakthrough result. At IMO 2024, AlphaGeometry and AlphaProof required experts to first translate problems from natural language into domain-specific languages, such as Lean, and vice-versa for the proofs. It also took two to three days of computation. This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions – all within the 4.5-hour competition time limit. [1]https://deepmind.google/discover/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/ https://deepmind.google/discover/blog/advanced-version-of-ge...
- 77pt77 1y ago
- mvieira38 1y agoWell, the supposed PhD-level models are still pretty dumb when they get to consumers, so what gives?
- apwell23 1y agoThis comment makes me think. What did previous winners of these competition go on to do in their lives? Anything spectacular?
- twhyn2 1y agoIndeed. I personally view all this stuff as noise. Im more interested in seeing any contributions to the real economy. Not some competition stuff that is irrelevant to the welfare of people.
- atleastoptimal 1y agoPeople pattern match with a very low-resolution view of the world (web3/crypt/nfts were a bubble because there was hype, so there must be a bubble since AI is hyped! I am very smart) and fail to reckon with the very real ways in which AI is fundamentally different. Also I think people do understand just how big of a deal AI is but don't want to accept it or at least publicly admit it because they are scared for a number of reasons, least of all being human irrelevance.
- paxys 1y agoMy response simply is that performance in coding competitions such as ICPC is a very different skillset than what is required in a regular software engineering job. GPT-5 still cannot make sense of my company's legacy codebase even if asked to do the most basic tasks that a new grad out of college can figure out in a day or two. I recently asked it to fix a broken test (I had messed with it by changing one single assertion) and it declared "success" by deleting the entire test suite.
- mrkeen 1y agoSimilar experience with windsurf. I had a class of 5 or so test methods - ABCDE. I asked it to fix C, so it started typing out B token-by-token underneath C, such that my source file was now ABCBDE. I don't think I'm smart enough to get it to do coding activities.
- 77pt77 1y ago> it declared "success" by deleting the entire test suite. The paperclip trivial solution!
- raspasov 1y agoThis. Dealing with the problems of a real-world legacy code base is the exact opposite of a perfectly constrained problem, verified for internal consistency probably by computers and humans, of all things, and presented neatly in a single PDF. There are dozens, if not 100s, of assumptions that humans are going to make while solving a problem (i.e., make sure you don't crash the website on your first day at work!) that an LLM is not going to. Similar to why, despite all its hype, Waymo cars are still being supervised by human drivers nearly 100% of the time and can't even park themselves regularly without stalling with no explanation.
- levocardia 1y ago>Waymo cars are still being supervised by human drivers nearly 100% of the time That seems...highly implausible?
- 1y ago
- chpatrick 1y agoDon't worry, they're just stochastic parrots copying answers from Stack Overflow. ;)
- reducesuffering 1y agoPeople are having a tough time coping with what the near future holds for them. It is quite hard for a typical person to imagine how disruptive and exponential coming world events are like Covid showed.
- noosphr 1y agoTwo days ago I talked to someone in water management about data centers. One of the big players wanted to build a center that consumed as much water as a medium town in semi arid bushland. A week before that it was a substation which would take a decade to source the transformers for. Before that it was buying closed down coal power plants. I don't know if we're in a bubble for model capabilities, but we are definitely hitting the wall in terms of what the rest of the physical economy can provide. You can't undo 50 years of deffered maintenance in three months.
- trhway 1y agoGetting well funded commercial demand is exactly how you undo it.
- noosphr 1y agoNot in three months. It will take years if not decades. What happens when OpenAI and friends go bust because China is drowning in spare grid capacity and releasing sota open weights models like R1 every other week? Every company building infrastructure for AI also goes out of business and we are in a worse position than we are now because instead of having a tiny industry building infrastructure at a level required to replace what has reached end of life we have nothing.
- m3kw9 1y agothe wall is how we need to throw trillions of hardware to do "breakthroughs", LLM uses the same algorthm from last few years. We need a new algorthm breakthrough otherwise buying hardware to increase intelligence isn't scalable.
- nofriend 1y agoWhere these competitions differ from real life is that evaluating a solution is much easier than generating a solution. We're at the point where AI can do a pretty good job of evaluating solutions, which is definitely an impressive step. We're also at the point where AI can generate candidate solutions to problems like these, which is also impressive. But the degree to which that translates to practical utility is questionable. The sibling commenter compared this to go, but we could go back to comparing it with chess. Deepblue didn't play chess the way a human did. It deployed massive amounts of compute, to look at as many future board states as possible, in order to see which move would work out. People who said that a computer that could play chess as well as a human would be as smart as a human ended up eating crow. These modern AIs are also not playing these competitions the way a human does. Comparing their intelligence to that of a humans is similarly fallacious.
- Ianjit 1y agoHistorically there has been a gap between the performance of AI in test environments vs the impact in the real world, and that makes people who have been through the cycle a few times cautious extrapolating. In 2016 Geoffrey Hinton said vision models would put radiologists out of business within 5-10 years. 10 years on there is a shortage of Radiologists in the US and AI hasn't disrupted the industry. The DARPA grand challenge for autonomous vehicles was won in 2006, 20 years on self driving cars still have limited deployment. The real world is more complex than computer scientists apprecate.
- LunaSea 1y ago"We used a custom AI that requires a small nuclear plant to be trained and function to beat three humans consuming 400 watts per day" isn't as impressive as it sounds
- birktj 1y agoThey apparently managed gold in the IOI as well. A result that was extremely surprising for me and causes me to rethink a lot of assumptions I have about current LLMs. Unfortunately there was very little transparency on how they managed those results and the only source was a Twitter post. I want to know if there was any third party oversight, what kind of compute they used, how much power what kind of models and how they were set up? In this case I see that DeepMind at least has a blog post, but as far as I can see it does not answer any of my questions. I think this is huge news, and I cannot imagine anything other than models with this capability having a massive impact all over the world. It causes me to be more worried than excited, it is very hard to tell what this will lead which is probably what makes it scary for me. However with so little transparency from these companies and extreme financial pressure to perform well in these contests, I have to be quite sceptical of how truthful these results are. If true I think it is really remarkable, but I really want some more solid proof before I change my worldview.
- XenophileJKO 1y agoSo outside of human intervention, I don't think the specifics really matter. What this means is that it is possible and that this capability will in time be commoditized. This is helpful in framing the conversation, especially with "skeptics" of what these models are capable of.
- birktj 1y agoTo a certain extent I agree. But as far as I know I cannot go to chatgpt.com and paste the newest ICPC problems and get full solutions. And there is no information about what they do differently. For a competition like the ICPC, which is academic in its nature, I think it is very unfortunate to setup a seperate AI track like this without publishing clear public information about what that actually entails. And have clear requirements for these AI companies to publish their methology. I know it is a nice source of sponsorships for them, but the ICPC should afford to stand up a bit for academic integrity. Without any of this I can't even know for sure if there was any human intervention. I don't really think so, but as I mentioned the financial pressure to perform well is extreme so I can totally see that happening. Maybe ICPC did have some oversight, but please write a bit about it then. If you assume no human intervention then all of this is of course irrelevant if you only care about the capabilities that exist. But still the implications of a general model performing at this level vs something more like a chess model trained specifically on competitive programming are of course different, even if the gap may close in the future. And how much compute/power was used, are we talking hundreds of kWhs? And does that just means larger models than normally or intelligent bruteforcing through a huge solutionspace? If so, then it is not clear how much they will be able to scale down the compute usage while keeping the performance at the same level
- JohnKemeny 1y agoI went to ICPC's web pages, downloaded the first problem (problem A) and gave it to GPT-5, asking it for code to solve it (stating it was a problem from a recent competitive programming contest). It thought for 7m 53s and gave as reply # placeholder # (No solution provided)
- CamperBob2 1y agoSounds like a bug. Did you try it again (or with another leading-edge model) and get a similar result?
- patrickhogan1 1y ago1. What was your prompt? 2. Why did you give it to GPT-5 instead of GPT-5 Thinking or GPT-5 Pro?
- patrickhogan1 1y agoHere is the prompt I just gave to GPT-5 Pro - its chugging on it. Not sure if it will succeed. Let's see what happens. I did think about converting the PDF to markdown, but figured this prompt is more fair. - You are a gold level math olympiad competitor participating in the ICPC 2025 Baku competition. You will be given a competitive programming problem to solve completely. All problems are located at the following URL: https://worldfinals.icpc.global/problems/2025/finals/problems/problemset.pdf https://worldfinals.icpc.global/problems/2025/finals/problem... Here is the problem you need to solve and only solve this problem: <problem> Problem B located on Page 3 of the PDF that starts with this text - but has other text so ensure you go to the PDF and look at all of page 3 To help her elementary school students understand the concept of prime factorization, Aisha has invented a game for them to play on the blackboard. The rules of the game are as follows. The game is played by two players who alternate their moves. Initially, the integers from 1 to n are written on the blackboard. To start, the first player may choose any even number and circle it. On every subsequent move, the current player must choose a number that is either the circled number multiplied by some prime, or the circled number divided by some prime. That player then erases the circled number and circles the newly chosen number. When a player is unable to make a move, that player loses the game. To help Aisha’s students, write a program that, given the integer n, decides whether it is better to move first or second, and if it is better to move first, figures out a winning first move.</problem> Your task is to provide a complete solution that includes: 1. A thorough analysis and solution approach 2. Working code implementation 3. Unit test cases with random inputs 4. Performance optimization to run within 1 second Use your scratchpad to think through the problem systematically before providing your final solution. <scratchpad> Think through the following steps: 1. Problem Understanding: - What exactly is the problem asking for? - What are the input constraints and output requirements? - Are there any edge cases to consider? 2. Solution Strategy: - What algorithm or mathematical approach should be used? - What is the time complexity of your approach? - What is the space complexity? - Will this approach work within the given constraints? 3. Implementation Planning: - What data structures will you need? - How will you handle input/output? - What are the key functions or components? 4. Testing Strategy: - What types of test cases should you create? - How will you generate random inputs within the problem constraints? - What edge cases need specific testing? 5. Optimization Considerations: - Are there any bottlenecks in your initial approach? - Can you reduce time or space complexity? - Are there language-specific optimizations to apply? </scratchpad> Now provide your complete solution with the following components: <analysis> Provide a detailed analysis of the problem, including: - Problem interpretation and requirements - Chosen algorithm/approach and why - Time and space complexity analysis - Key insights or mathematical observations </analysis> <solution> Provide your complete, working code solution. Make sure it: - Handles all input/output correctly - Implements your chosen algorithm efficiently - Includes proper error handling if needed - Is well-commented for clarity </solution> <unit_tests> Create comprehensive unit test cases that: - Test normal cases with random inputs within constraints - Test edge cases (minimum/maximum values, boundary conditions) - Include at least 5-10 different test scenarios - Show expected outputs for each test case </unit_tests> <optimization> Explain any optimizations you made or could make: - Performance improvements implemented - Memory usage optimizations - Language-specific optimizations - Verification that solution runs within 1 second for maximum constraints </optimization> Take all the time you need to solve this problem thoroughly and correctly.
- jaggs 1y agoI think it's becoming clear that these mega AI corps are juggling with their models at inference time to produce unrealistically good results. By that it seems that they're just cranking up the compute beyond reasonable levels in order to gain PR points against each other. The fact is most ordinary mortals never get access to a fraction of that kind of power, which explains the commonly reported issues with AI models failing to complete even rudimentary tasks. It's now turned into a whole marketing circus (maybe to justify these ludicrous billion-dollar valuations?).
- andy12_ 1y agoModels drop in price x10 each year. Us, common folk, getting access to these kinds of models is just a matter of time.
- jaggs 1y agoIs that true though? Having to pay some $200 a month for a max account of whatever kind doesn't seem to be cheaper to me at all?
- scarmig 1y ago$200/month for an LLM with the capability to fully automate my job is extremely cheap. Of course, even with a high thinking budget we don't have that yet, but if we see it at any cost in 2026, I'll be expecting to be forced into retirement by 2030.
- andy12_ 1y agoWhen I say 10 times cheaper, I mean when comparing models of the same capabilities. The kind of performance you get now for a 200$ subscription, a year ago probably would have costed 2000$.
- jaggs 1y agoI understand what you're saying. However I'm not sure it's that germane when we're talking about whether or not the current $200 subscription fee is actually delivering value for money, or whether AI giants are manipulating performance to gain marketing points.
- modeless 1y agoMore information on OpenAI's result (which seems better than DeepMind's) from the X thread: > our OpenAI reasoning system got a perfect score of 12/12 > For 11 of the 12 problems, the system’s first answer was correct. For the hardest problem, it succeeded on the 9th submission. Notably, the best human team achieved 11/12. > We had both GPT-5 and an experimental reasoning model generating solutions, and the experimental reasoning model selecting which solutions to submit. GPT-5 answered 11 correctly, and the last (and most difficult problem) was solved by the experimental reasoning model. I'm assuming that "GPT-5" here is a version with the same model weights but higher compute limits than even GPT-5 Pro, with many instances working in parallel, and some specific scaffolding and prompts. Still, extremely impressive to outperform the best human team. The stat I'd really like to see is how much money it would cost to get this result using their API (with a realistic cost for the "experimental reasoning model").
- bazmattaz 1y agoHa so true. I was so tempted to copy and paste a problem into GPT5 and see what it would say
- HardCodedBias 1y agoThey likely had a prompt that gave considerable guidance. Hopefully that prompt was the same for all questions (I think that is what they did for the IMO submission, or maybe it was Google that did that, not sure).
- qwertox 1y ago> it succeeded on the 9th submission What's the judgement here? Was it within the allotted time, or just a "try as often as you need to"?
- modeless 1y agoIt was within the allotted time. If I'm reading the scoreboard correctly [edit: I wasn't], the human teams typically submitted dozens or hundreds of attempts at each problem.
- ChrisArchitect 1y agoSharing links to a couple of tweets is not a blog post. Google source post: https://deepmind.google/discover/blog/gemini-achieves-gold-level-performance-at-the-international-collegiate-programming-contest-world-finals/ https://deepmind.google/discover/blog/gemini-achieves-gold-l... (https://news.ycombinator.com/item?id=45278480 https://news.ycombinator.com/item?id=45278480) OpenAI tweet: https://x.com/OpenAI/status/1968368133024231902 https://x.com/OpenAI/status/1968368133024231902 (https://news.ycombinator.com/item?id=45279514 https://news.ycombinator.com/item?id=45279514)
- ototot 1y agoGiven that ICPC problems are in general easier than IOI problems. I wouldn't be surprise to see they can get Gold (even perfect scores) in ICPC. Nonetheless, I'm still questioning what's the cost and how long it would take for us to be able to access these models. Still great work, but it's less useful if the cost is actually higher than hiring someone with the same level.
- JohnKemeny 1y agoWhat makes you say that they are easier? Are there more people who manages to solve a problem from ICPC than from IOI? How do you compare those? There were at least 2 very simple problems in IOI this year. I haven't read the ICPC problem set, and perhaps there are some low-hanging fruits, but I highly doubt it.
- ototot 1y agoBecause I'm a ICPC medalist (not this year though) but not a IOI medalist. Another evidence is that you only have 5 hours to solve 3 problems in IOI, but you need to solve 10+ problems in ICPC. It's impossible to have all 10+ problems to at IOI level in ICPC.
- tgma 1y ago> Because I'm a ICPC medalist (not this year though) but not a IOI medalist. Isn't getting a medal a function of your ranking, not score, in both cases? If so, that does not prove much about the difficulty of either.
- ototot 1y agoOK. I think my opinion and definition on "easier" is indeed vague. For "easier", I'm only comparing the thinking difficulty. Yes, medal is function of ranking but not difficulty. Nonetheless, I would say that IOI more focus on thinking, which I to some degree is not that good at, while ICPC is more like a mix thinking and implementing. Therefore, my ability to implement stuff can improve my ICPC ranking but not IOI.
- ferguess_k 1y agoI think in the future information will be more walled -- because AI companies are not paying anyone for that piece of information, and I encourage everyone to put their knowledge on their own website, and for each page, put up a few urls that humans won't be able to find (but can still click if he knows where to find), but can be crawled by AI, which link to pages containing falsified information (such as, oh the information on url blah is actually incorrect, here you can find the correct version, with all those explanations, blah blah -- but of course page blah is the only correct version). Essentially, we need to poison AI in all possible ways, without impacting human reading. They either have to hire more humans to filter the information, or hire more humans to improve the crawlers. Or we can simply stop sharing knowledge. I'm fine with it, TBF.
- tgma 1y agoWhy the AI hate? How is it different from sharing your knowledge with another individual or writing a book to share it? > AI companies are not paying anyone for that piece of information So? For the vast majority of human existence, paying for content was not a thing, just like paying for air isn't. The copyright model you are used to may just be too forced. Many countries have no moral qualms about "pirating" Windows and other pieces of software or games (they won't afford to purchase anyway.) There's no inherent morality or entitlement for author receiving payment for everything they "create" (to wit, Bill Gates had to write a letter to Homebrew Computer Club to make a case for this, showing that it was hardly the default and natural viewpoint.) It's just a legal/social contract to achieve specific goals for the society. Frankly the wheels of copyright have been falling off since the dawn of the Internet, not LLM.
- bgwalter 1y agoCompanies valued at $300 billion or more are not another individual and people are not "sharing" their works. The companies are stealing them. For the majority of interesting output people have paid for art, music, software, journalism. But you know that already and are justifying the industry that pays your bills.
- 1y ago
- bgwalter 1y agoA database is good at leetcode, who would have thought. Give humans a database and they'll outperform your "AI" (which probably uses an extraordinary amount of graphics cards and electricity). It is an idiotic benchmark, in line with the rest of the "AI" propaganda.
- chpatrick 1y agoWhere is this magic ICPC competition answers database that they're using?
- bgwalter 1y ago"Database" was not meant in a literal sense. Clearly a lot of knowledge from similar problems is encoded in the model, that is why you can use models as a kind of fuzzy encyclopedia. It is like an open book exam for humans where they also can lookup similar problems. The current top comment makes the same point, but in a more diplomatic and sophisticated manner.
- chpatrick 1y agoI mean strong human contestants would also know a lot of similar problems, I'm not seeing how it's fundamentally different or not a meaningful achievement.
- JohnKemeny 1y agoDo you also think that IBM Watson won Jeopardy primarily because of its ability to reason? Or due to its massive internal database combined with superhuman button pressing skills?
- chpatrick 1y agoAs far as I know Watson was actually smoke and mirrors, relying on a ton of human hand written rules and didn't do any reasoning at all. It was a giant program written in PROLOG, not machine learning.
- smokel 1y agoThe best thing of the ICPC is the first C, which stands for "collegiate". It means that you get to solve a set of problems with three persons, but with only one computer. This means that you have to be smart about who is going to spend time coding, thinking, or debugging. The time pressure is intense, and it really is a team sport. It's also extra fun if one of the team members prefers a Dvorak keyboard layout and vi, and the others do not. I wonder how three different AI vendors would cooperate. It would probably lift reinforcement learning to the next level.
- Workaccount2 1y agoClaude, ChatGPT, and Gemini on a team. I'm not sure how it would play out, but at least when you let them talk to each other they tend to get very technical very fast.
- deleted 1y ago[deleted]
- NoahZuniga 1y agoActually collegiate means that the contestants are in college.
- smokel 1y agoHa, shows what I know :) As a non-native speaker I had always assumed it referred to working with colleagues. Etymologically related, but apparently not the same.
- AndrewKemendo 1y agoClearly strongly held Hallucinations are problematic no matter what agent produces them
- amluto 1y agoI've contemplated this a bit, and I think I have a bit of an unconventional take: First, this is really impressive. Second, with that out of the way, these models are not playing the same game as the human contestants, in at least two major regards. First, and quite obviously, they have massive amounts of compute power, which is kind of like giving a human team a week instead of five hours. But the models that are competing have absolutely massive memorization capacity, whereas the teams are allowed to bring a 25-page PDF with them and they need to manually transcribe anything from that PDF that they actually want to use in a submission. I think that, if you gave me the ability to search the pre-contest Internet and a week to prepare my submissions, I would be kind of embarrassed if I didn't get gold, and I'd find the contest to be rather less interesting than I would find the real thing.
- paladin314159 1y ago> I think that, if you gave me the ability to search the pre-contest Internet and a week to prepare my submissions, I would be kind of embarrassed if I didn't get gold, and I'd find the contest to be rather less interesting than I would find the real thing. I don't know what your personal experience with competitive programming is, so your statement may be true for yourself, but I can confidently state that this is not true for the VAST majority of programmers and software engineers. Much like trying to do IMO problems without tons of training/practice, the mid-to-hard problems in the ICPC are completely unapproachable to the average computer science student (who already has a better chance than the average software engineer) in the course of a week. In the same way that LLMs have memorized tons of stuff, the top competitors capable of achieving a gold medal at the ICPC know algorithms, data structures, and how to pattern match them to problems to an extreme degree.
- amluto 1y ago> I can confidently state that this is not true for the VAST majority of programmers and software engineers. That may well be true. I think it's even more true in cases where the user is not a programmer by profession. I once watched someone present their graduate-level research in a different field and explain how they had solved a real-world problem in their field by writing a complicated computer program full of complicated heuristics to get it to run fast enough and thinking "hmm, I'm pretty sure that a standard algorithm from computer graphics could be adapted to directly solve your problem in O(n log n) time". If users can get usable algorithms that approximately match the state of the art out of a chatbot (or a fancy "agent") without needing to know the magic words, then that would be amazing, regardless of whether those chatbots/agents ever become creative enough to actually advance the state of the art. (I sometimes dream of an AI producing a piece of actual code that comes even close to state of the art for solving mixed-integer optimization problems. That's a whole field of wonderful computer science / math that is mostly usable via a couple of extraordinarily expensive closed-source offerings.)
- HarHarVeryFunny 1y agoICPC = The International Collegiate Programming Contest. These are college level programmers, not elite competitive programmers. Apparently Gemini solved one problem (running on who knows what kind of cluster) by burning 30 min of "thinking" time on it, and at a cost that Google have declined to provide. According to one prior competition paricipant, writing in the comments section of this ArsClasica coverage, each year they include one "time sink" problem that smart humans will avoid until they have tackled everything else. https://arstechnica.com/google/2025/09/google-gemini-earns-gold-medal-in-icpc-world-finals-coding-competition/ https://arstechnica.com/google/2025/09/google-gemini-earns-g... This would all seem to put a rather different spin on this. It's not a case of Google outwitting the worlds best programmers, but rather that by searching for solutions for 30 min on god knows what kind of cloud hardware, they were able to get something done that the college kids did not have time to complete, or deem worthwhile starting.
- mannycalavera42 1y agoI've competed in these contest before. There are probably more difficult than what we can call _elite_ competitive programmer note: my team only passed the first 2 rounds, far from bragging about my skills here :)
- dist-epoch 1y agoLet's bookmark this comment and check again next year, if the freely available models will be able to do it for a few dollars.
- HarHarVeryFunny 1y agoSure, although my point wasn't intended to be about the cost (which would still be interesting to know), but rather that the win by Google seems more down to brute force than intelligence.
- amluto 1y agoThese are college-student or occasionally grad-school programmers who qualified to enter the ICPC World Finals, generally by performing sufficiently well at a regional championship to qualify. You can read actual rules here (see "Advancing to the ICPC World Finals"): https://icpc.global/regionals/rules https://icpc.global/regionals/rules I don't know what you mean by "elite", and there are certainly plenty of teams at the World Finals that are not especially competitive, and there certainly many elite programers who don't qualify for various reasons (most obviously by being the wrong age or not in the right stage of school or having already attended too many times), but I find it hard to believe that there aren't enough "elite" programmers present to make the winning teams be genuinely elite. Compare to, say, the Olympics or pretty much any academic olympiad. There are many people and teams at the Olympics who are not remotely competitive with the winners.
- d--b 1y agoTwo words: Uh oh
- sameermanek 1y agoWhats the point? These models are still unreliable in every day work. And they're getting fat! For a moment, they were getting cheaper, but now they are only getting bigger and this is not going to be cheap in the future. The point is, what are we investing a trillion dollars in?
- Ylpertnodi 1y ago/> The point is, what are THEY investing a trillion dollars in? Who cares? I won't be a customer until I see a return on my investment [in them].
- sidibe 1y agoUnreliable doesn't mean unusable. I'm finding it harder and harder to believe people are actually trying to use them and saying they are useless. If you can chop your problem up and give little tedious parts of your bigger task it's starting like doing code review for a new grad instead of coding. And they're getting more reliable and the parts you give it can be bigger and bigger. I wish there was a way to stop this but I don't think it's going to.
- Imnimo 1y agoMy understanding is that the way they do this is have some number of model instances generating solution proposals, and then another model which chooses which candidates to submit. I haven't been able to find information on how many proposals were generated before a solution was chosen to submit. I'm curious to know whether this is "you can get ICPC gold medal performance with a handful of GPT-5 instances" or "you will drown yourself in API credit debt if you try this". Still extremely impressive either way.
- antegamisou 1y agoMake that shit cure cancer/disease and abstain from that modern Space race equivalent BS ffs.
- m3kw9 1y agochemical space and in vivo testing is a different beast
- huflungdung 1y ago[dead]
- z7 1y agoCurrent cope collection: - It's not a fair match, these models have more compute and memory than humans - Contestants weren't really elite, they're just college level programmers, not the world's best - This doesn't matter for the real world, competitive programming is very different from regular software engineering - It's marketing, they're just cranking up the compute to unrealistic levels to gain PR points - It's brute force, not intelligence
- patrickhogan1 1y agoThis is impressive. Here is the published 2025 ICPC World Finals problemset. The "Time limit: X seconds" printed on each ICPC World Finals problem is the maximum runtime your program is allowed. If any judged run of your program takes longer than that, the submission fails, even if other runs finish in time. https://worldfinals.icpc.global/problems/2025/finals/problems/problemset.pdf https://worldfinals.icpc.global/problems/2025/finals/problem...
- m3kw9 1y agoi'm still waiting for LLMs to give us one profound science breakthrough
- sinuhe69 1y agoI wonder whether they allowed humans input for the AI besides the initial generic prompt? Could they provide guidance for the AI? We all know that by this kind of problems, intuition/guiding principles to transform the problem is all you need. The human may not be fast enough or error-free to sample correctly the already restricted solution space, but machine can. And for them, it’s a huge advantage. So did they allow human input (as part of a centaur team!) input or not? These AI teams often have one of the best (ex-) competitive programmers.
- Vegenoid 1y agoWhile very cool, this feels like another instance of the kind of thing that we already know they are good at: self-contained, perfectly-specified problems that can be done by humans in a short timespan (especially when a team of highly skilled engineers behind the model is wielding it). Yes, it's amazing that a computer can do this, consider what they could do 10 years ago to today, so on and so on - but I don't see this and go "holy shit", I see this and go "yep". I wish they went into more detail about how exactly the interaction with the LLM works - I'm pretty sure there's significantly more to it than "drop the paper with the problems into a text box and hit go".
- flimflamm 1y agoGood to note that OpenAI solved 12/12 and DeepMind 10/12.
- fullparens 1y agoI briefly looked at a few of Gemini's solutions https://github.com/google-deepmind/gemini_icpc2025 https://github.com/google-deepmind/gemini_icpc2025. What struck me was how Gemini finds clean ways to express an idea - perhaps because it knows a large set of tricks for each kind of sub-algorithm (within the larger algorithm). I am a former competitor and managed to reach world finals at Google CodeJam and Topcoder Open. It took me a lot of work to get there but I will gladly concede that Gemini is way better than my peak. I haven't competed in 15 years and have forgotten a lot of tricks but Gemini's code reminded me how quickly algorithms can get complicated sometimes without a bag of tricks. There are parallels to tactics in chess - humans might miss them but a machine will not. And that can be a huge difference in a game or even in a software project. EDIT: minor correction.