4 ms·
9 seconds to get both stars is absolutely insane - there had to be some AI assistance here. Come to think of it, a pipeline that feeds the problem text into an
by arjvik 2y ago
9 seconds to get both stars is absolutely insane - there had to be some AI assistance here.
Come to think of it, a pipeline that feeds the problem text into an LLM to generate a solution and automatically runs it on the input and attempts to submit the solution, doing this N times in parallel, could certainly solve the first few days' problem in 9 seconds.
- isoprophlex 2y agoAI coding assistants ruined the global leaderboard experience. AoC might as well nerf it by discarding the quickest x percent of submissions, or something...
- WithinReason 2y agoThey should check that LLMs can't solve the problems in 9 seconds and come up with appropriate problems. Or just allow AI assistants, they are now as much part of the programmer's toolkit as syntax highlighting or autocomplete or Stack Overflow, and pretending otherwise is not useful.
- Retr0id 2y agoThe first few days are supposed to be beginner-accessible, it's practically impossible to have something beginner accessible but GPT-inaccessible.
- martin-t 2y agoNot gonna happen. AoC always starts with beginner level problems. That's why it's so commonly used for learning the basics of new languages. A problem that wouldn't be immediately solvable by LLMs would either be too advanced or simply too large to be fun. This is probably where programming as a whole is going. Many of the things that make programming fun for me, like deeply understanding a small but non-trivial problem and finding a good solution, are gonna be performed much faster by LLMs. After all most of what we do has been done before, just in a slightly different content or a different language. Either LLMs will peak out at the current level and be often useful but very error prone and not-quite-there. Or they'll get better and we'll be just checking their output and designing the general architecture.
- matsemann 2y agoThat's like going out for a run and taking an electrical scooter around the park instead. The point isn't finishing, the point is doing the activity.
- WithinReason 2y agoThen why have a leaderboard?
- matsemann 2y agoBecause someone likes to compete? There are 5k races as well, which people enjoy to do even though vehicles exist. And people would rightfully be upset if they got beaten by someone not running themselves.
- zwirbl 2y agoAnd then allow aimbots for counterstrike, stockfish at chess tournaments and Epo on the tour de France. The leader board is intended for people to compete against each other, one could make a separate leaderboard for LLM, kind of similar to the chess AI leaderboards.
- falcor84 2y ago> allow aimbots for counterstrike I'm not played counterstrike in over a decade, so you got me wondering - are there matches where everyone uses aimbots? What does the game look like then? I suppose there's a new mix of strategies evolving, with a higher focus on the macro movement planning?
- worthless-trash 2y ago> are there matches where everyone uses aimbots? Yes > What does the game look like then? I have only observed the games, it requires a lot of hiding. Most of the time the winning method is to act at the very last second and hope the other player is distracted.
- WithinReason 2y agoFalse equivalence. The sole reason for counterstrike and chess to exist is competition. Programming is about solving a problem. If you want to turn programming into a competition you shouldn't take away tools from the programmer.
- jbjbjbjb 2y agoYou’re saying programming isn’t not equivalent to chess here because programming isn’t a competition, but the Advent of Code leaderboard very much is a competition.
- fuglede_ 2y agoIt's not that bad. I'm sure there are more LLM'ers in there than the one, but you can tell that the majority of the day 1 leaderboard is made up of people who have historically performed well, even before LLMs were a thing. Compare https://adventofcode.com/2024/leaderboard/day/1 https://adventofcode.com/2024/leaderboard/day/1 to e.g. https://fuglede.github.io/aoc-full-leaderboard/ https://fuglede.github.io/aoc-full-leaderboard/ There was also at least one instance of people working together where you would have 15 people from the same company submit solutions at the same time, which can be a bit frustrating but again, not a huge issue.
- fuglede_ 2y agoOkay, I think I have to go ahead and retract my own comment. Day 5 appears to have been sufficiently tricky for humans to do quickly while still easy enough for the LLMs that it is clear that there is a very large amount of cheating going on.
- gorgoiler 2y agoI have a rule in life: no summary statistics without showing the distribution. Usually this goes for any median which might be in a sneaky bimodal distribution of, say, AI models vs humans. I guess it applies to leaderboards too though.
- exitb 2y agoPotentially the challenge just doesn’t make as much sense anymore? There apparently are „mental calculations” competitions and I’m sure their participants have fun. Yet I can hardly imagine doing arithmetic in ones head is any fun for an average mathematician. The challenge just shifted elsewhere over time.
- hmottestad 2y agoOpenAI did something similar with their o1 model. Ran a coding problem through o1 thousands or maybe millions of times and then checked if the solution was correct. I can imagine a great pipeline for performance optimization: 1. have an AI generate millions of tests for your existing code 2. have another AI generate faster code that still makes the tests pass So I guess all I want for Christmas is a massive compute cluster and infinite OpenAI credits :P
- nneonneo 2y agoIt was, of course, an AI-generated solution; they posted it here: https://web.archive.org/web/20241201052156/https://github.com/qianxyz/advent-of-code/commit/6458076f5782790f8878a957add96093e72ce5a2 https://web.archive.org/web/20241201052156/https://github.co... Later, after being called out on it, they posted an apology to their GitHub profile (https://web.archive.org/web/20241201064816/https://github.com/qianxyz https://web.archive.org/web/20241201064816/https://github.co...): "If you are here from the AoC leaderboard, I apologize for not reading the FAQ. Won't happen again." Both the repo and that message are now gone.
- dsissitka 2y agoFor comparison, here's how long it took in past years: 2015 - 10:55 2016 - 7:01 2017 - 1:16 2018 - 1:48 2019 - 1:39 2020 - 7:11 2021 - 1:07 2022 - 0:53 2023 - 2:24 And this year's second place was 0:54.
- fuglede_ 2y agoNote that regarding the outliers, in 2015 and 2016 the puzzles weren't as widely known, and in 2020, AWS' load balancers crashed and the puzzle was unavailable to most people for 6 minutes, then solved in a few minutes. https://adventofcode.com/2020/leaderboard/day/1 https://adventofcode.com/2020/leaderboard/day/1 -- postmortem: https://old.reddit.com/r/adventofcode/comments/k9lt09/postmortem_2_scaling_adventures/ https://old.reddit.com/r/adventofcode/comments/k9lt09/postmo...
- fuglede_ 2y agoYep, there was; they even wrote so in their commit message before removing it: https://old.reddit.com/r/adventofcode/comments/1h3w7mc/2024_day_1_no_llms_here/ https://old.reddit.com/r/adventofcode/comments/1h3w7mc/2024_...
- Almondsetat 2y agoCaring about the leaderboards is the problem. Are you (impersonal) seriously doing AoC for clout or something?
- nicce 2y agoI did it without AI last year and planning to do it again, for fun.
- zwnow 2y agoTrue. Leaderboards are and always have been full of cheaters.
- thinkingemote 2y agoIf people are concerned about leaderboards there are private leaderboards: https://adventofcode.com/2024/leaderboard/private https://adventofcode.com/2024/leaderboard/private Personally I don't do it to compete, I just like puzzles.
- aithrowawaycomm 2y agoThe primary reason to not care about AoC leaderboards is that that it penalizes people for being in the wrong time zone. That said, the top 100 or so contributors clearly do care about these things and using an LLM is cheating. In particular the LLM cheating isn’t just by conjuring a solution: humans don’t get ASCII characters pumped directly into their brain, we have to slowly read problem descriptions with our eyes. It takes humans more than 9 seconds to solve AoC #1 purely because of unavoidable latency.
- exitb 2y agoI wonder how many seconds could be won by the organizers if the challenge included a blant prompt injection breaking the result.
- emadb 2y agoI participate almost every year but I don't care about the leaderboard. The timezone play a crucial role in being able to be ready at the right time, so actually who cares? I prefer to build private leaderboards with my friends and colleagues.
- rvz 2y agoAt this point, it just shows that Advent of Code is completely worthless given the ease and accessibility of AI-assisted tools to solve these problems. RIP Advent of Code.
- wiseowise 2y agoSarcasm?
- abenga 2y agoWhy? Solving interesting problems to learn is a worthwhile goal. Why should it matter to you that others are "cheating"?
- zwirbl 2y agoBecause competing against other people is the fun part for some people. Why not allow everyone to use stockfish at chess tournaments?
- zwnow 2y agoThis is stupid though. Advent of Code Leaderbords were always full of cheaters. At least since 2020 when I first started. If you want competitive programming, AoC is not the place for that.
- graynk 2y ago“For _some_ people” is at odds with “_completely_ worthless”, don’t you think?
- dakiol 2y agoTo be fair, I read it as "_completely_ worthless for _some_ people"
- rich_sasha 2y agoWithout condoning cheating, I am impressed with the automation aspect of it. 9 seconds sounds more or less like the inference time of the LLM, so this must have been automated. Login at midnight + lots of C&P may not have done it. Perhaps there is a scope for an alternative AoC type competition aimed at AI submissions... ...though of course that would be experimenting to get us all out of work. Hmm.
- zwnow 2y agoIf real life problems were as easy and defined as AoC problems we might be able to be replaced at some point. I highly doubt you can replace software devs otherwise. Who else is going to take the blame for software issues?
- rich_sasha 2y ago"Write me a snippet that does X" is a step behind "figure out how to log into this page, download the data, write a snippet that gives the right answer to the sample data, then run it on the real thing and submit the output to the text box".
- deleted 2y ago[deleted]
- tags2k 2y agoAs with real life, the speed generally doesn't matter as long as you get a working solution and you find it fun. If you find "copy and paste into an LLM and then copy and paste the answer back out" fun, then I suppose you do you. I didn't realise it was be timed, which is good because I casually set up a new rig to give future puzzles some kind of rig. I used C# which, although probably more wordy than other solutions, did the job and LINQ made light work of the list operations. Ended up with about 6.5 minutes for each one but most of that was refactoring out of pedantry.