8 ms·
Performance of LLMs on Advent of Code 2024
- jebarker 2y agoI'd be interested to know how o1 compares. On may days after I completed the AoC puzzles I was putting them question into o1 and it seemed to do really well.
- qsort 2y agoAccording to this thread: https://old.reddit.com/r/adventofcode/comments/1hnk1c5/results_of_a_multiyear_llm_experiment/ https://old.reddit.com/r/adventofcode/comments/1hnk1c5/resul... o1 got 20 out of 25 (or 19 out of 24, depending on how you want to count). Unclear experimental setup (it's not obvious how much it was prompted), but it seems to check out with leaderboard times, where problems solvable with LLMs had clear times flat out impossible for humans. An agent-type setup using Claude got 14 out of 25 (or, again, 13/24) https://github.com/JasonSteving99/agent-of-code/tree/main https://github.com/JasonSteving99/agent-of-code/tree/main
- johnea 2y agoLLMs are writing code for the coming of the lil' baby jesus?
- valbaca 2y agoadventofcode.com
- grumple 2y agoI’m both surprised and not surprised. I’m surprised because these sort of problems with very clear prompts and fairly clear algorithmic requirements are exactly what I’d expect LLMs to perform best at. But I’m not surprised because I’ve seen them fail on many problems even with lots of prompt engineering and test cases.
- yunwal 2y agoWith no prompt engineering this seems like a weird comparison. I wouldn’t expect anyone to be able to one-shot most of the AOC problems. A fair fight would at least use something like cursor’s agent on YOLO mode that can review a command’s output, add logs, etc
- ben_w 2y agoIf you immediately know the candlelight is fire, then the meal was cooked a long time ago. And so it is with success of LLMs in one-shot challenges and any job that depends on such challenges: cooked a long time ago.
- NitpickLawyer 2y agoIndeed. (wild to see a sg-1 quote in the wild!)
- fumeux_fume 2y agoI certainly don't think it's weird to measure one-shot performance on the AOC. Sure, more could be done. More can always be done, but this is still useful and interesting.
- deleted 2y ago[deleted]
- segmondy 2y agoI did zero shot 27 solutions successfully with local model code-qwen2.5-32b. I think adding sonnet or latest gemini will probably get me to 40.
- rhubarbtree 2y agoSeems reasonable to me. More people use ChatGPT than cursor, so use it like ChatGPT. For most coding problems, people won't magically know if the code is "correct" or not, so any one-shot answer that is wrong could be a real hindrance. I don't have time to prompt engineer for every bit of code. I need tools that accelerate my work, not tools that absorb my time.
- unclad5968 2y agoHalf the time I try to use gemini questions about the c++ std library, it fabricates non-existent types and functions. I'm honestly impressed it was able to solve any of the AoC problems.
- devjab 2y agoI’ve had a side job as an external examiner for CS students for almost a decade now. While LLMs are generally terrible at programming (in my experience) they excel at passing finals. If I were to guess it’s likely a combination of the relatively simple (or perhaps isolated is a better word) tasks coupled with how many times similar problems have been solved before in the available training data. Somewhat ironically the easiest way to spot students who “cheat” is when the code is great. Being an external examiner, meaning that I have a full time job in software development I personally find it sort of silly when students aren’t allowed to use a calculator. I guess changing the way you teach and test is a massive undertaking though, so right now we just pretend LLMs aren’t being used by basically every students. Luckily I’m not part of the “spot the cheater” process, so I can just judge them based on how well they can explain their code. Anyway, I’m not at all surprised that they can handle AoC. If anything I would worry that AoC will still be a fun project to author when many people solve it with AI. It sure won’t be fun to curate any form of leaderboard.
- uludag 2y agoleetoced/AoC-like problems are probably the easiest class of software related tasks LLMs can do. Using the correct library, the correct way, at the correct version, especially if the library has some level of churn and isn't extremely common, can be a harder task for LLMs than the hardest AoC problems.
- upghost 2y agoAfter looking at the charts I was like "Whoa, damn, that Jerpint model seems amazing. Where do I get that??" I spent some time trying to find it on Huggingface before I realized...
- senordevnyc 2y agolol, I did the same thing
- j45 2y agoMe too. The author could make a model on huggingface routing requests to him. Latency might vary.
- tbagman 2y agoThe good news is that training jerprint was probably cheaper than training the latest got models...
- foldl2022 2y agoJust open-wights it. We need this. :)
- moffkalast 2y agoAt first I was like "What is this jerpint model that's beating the competition so soundly?" then it hit me, lol. Anyhow this is like night and day compared to last year, and it's impressive that Sonnet is now apparently 50% as good as a professional human at this sort of thing.
- zkry 2y agoI don't think comparing star counts would be a good measure though, as with AOC 90% of the effort and difficulty goes into the harder problems towards the end and it was the beginning, easy problems where the bulk of the sonnet's stars came from.
- moffkalast 2y agoAh yeah that's true, the difficulty curve is not very linear.
- BugsJustFindMe 2y agoI think this is a terrible analysis with a weak conclusion. There's zero mention of how long it took the LLM to write the code vs the human. You have a 300 second runtime limit, but what was your coding time limit? The machine spat out code in, what, a few seconds? And how long did your solutions take to write? Advent of code problems take me longer to just read than it takes an LLM to have a proposed solution ready for evaluation. > they didn’t perform nearly as well as I’d expect Is this a joke, though? A machine takes a problem description written as floridly hyperventilated as advent problems are, and, without any opportunity for automated reanalysis, it understands the exact problem domain, it understands exactly what's being asked, correctly models the solution, and spits out a correct single-shot solution on 20 of them in no time flat, often with substantially better running time than your own solutions, and that's disappointing? > a lot of the submissions had timeout errors, which means that their solutions might work if asked more explicitly for efficient solutions. However the models should know very well what AoC solutions entail You made up an arbitrary runtime limit and then kept that limit a secret, and you were surprised when the solutions didn't adhere to the secret limit? > Finally, some of the submissions raised some Exceptions, which would likely be fixed with a human reviewing this code and asking for changes. How many of your solutions got the correct answer on the first try without going back and fixing something?
- keyle 2y agoYou raise some good points about the "total of hours spent" but I guess you don't consider training time included. Also there is no need to quote the author's post and have a go at him personally. There are better ways to get your point across by arguing the point made and not the sentences written.
- BugsJustFindMe 2y ago> I guess you don't consider training time included In the same way that I don't consider the time an author spends writing a book when saying how long it took me to read the book. OP lost zero time training ChatGPT. Or do you mean how much time OP spent training themself? Because that's a whole new can of worms. How many years should we add to OP's development times because probably they learned how to write code before this exercise?
- bryan0 2y agoSince you did not give the models a chance to test their code and correct any mistakes, I think a more accurate comparison would be if you compared them against you submitting answers without testing (or even running!) your code first
- angarg12 2y agoPeople keep evaluating LLMs on essentially zero-shotting a perfect solution to a coding problem. Once we use tools to easily iterate on code (e.g. generate, compile, test, use outcome to refine prompt) we will turbocharge LLMs coding abilities.
- rhubarbtree 2y agoThis smacks of "moving the goalposts" just as much as the other side is accused of when unimpressed by advances. It's a reasonable test.
- xen0 2y agoI'm a bit curious to see how close the solutions were. When it couldn't solve it within the constraints, was it 'broadly correct' with some bugs? Was it on the right track or completely off?
- jerpint 2y agoThe code they generated and the outputs (including trace back errors) is all included, you can view them on the post itself or from hf space: https://huggingface.co/spaces/jerpint/advent24-llm https://huggingface.co/spaces/jerpint/advent24-llm
- xen0 2y agoHuh, don't know how I missed that. I was hoping for them to have done some analysis themselves. Glancing through the code for some of the 'easier' problems Claude (his best performer), it seems to have broadly correct (if a little strange and overwrought) code. But my phone is not a good platform for me to do this kind of reading.
- bongodongobob 2y agoI think a major mistake was giving parts 1 and 2 all at once. I had great results having it solve 1, then 2. I think I got 4o to one shot parts 1 then 2 up to about day 12. It started to struggle a bit after that and I got bored with it at day 18. It did way better than I expected, I don't understand why the author is disappointed. This shit is magic.
- 101008 2y agoThis kind of mirror my experience with LLMs. If I ask them non-original problems (make this API, write this test, update this function (that must be written 100s of time by develoeprs around the world, etc), it works very well. Some minors changes here and there but it saves time. When I ask them to code things that they never heard of (I am working on a online sport game), it fails catastrophically. The LLM should know the sport, and what I ask is pretty clear for anyone who understand the game (I tested against actual people and it was obvious what to expect), but the LLM failed miserably. Even worse when I ask them to write some designs in CSS for the game. It seems if you take them outside the 3-columsn layout or bootstrap or the overused landing page, LLMs fails miserably. It works very well for the known cases, but as soon as you want them to do something original, they just can't.
- jppope 2y agocompletely agree with this. The problem I'm starting to run into is that I do a lot of things that have been done before so I get really efficient at doing that stuff with LLMs... and now I'm slow to solve the stuff I used to be really fast at.
- anonzzzies 2y agoI ask llms things like I would spec it out for my team, which is to say, when I write a task, I would not include things about the sport, I would explain the logic etc one would need to write. No one needs to know what it is about in abstract as I see the same issue with humans; they get confused when using real world things to see the connection with what they need to do. I mean I don't flesh it out to the extend I might as well write the code, but enough that you don't need to know anything about the underlying subject. We are doing a large EU healthcare project and things differ per country obviously; if I would assume people to have a modicum of world knowledge or interest in it and look it up, nothing would get done. It is easier to deliver excel sheets with the proper wording and formulas and leave out what it is even for. Works fine with LLMs. Disclaimer; the people in my company know everything about the subject matter; the people (and LLMs) that implement it are usually not ours: we just manage them and in my experience, programmers at big companies are borderline worthless, hence, in the past 35 years, I have taken to write tasks as concrete as I can; we found that otherwise we either get garbage commits or tons of meetings with questions and then garbage commits. And now that comes in handy as this works much better on LLMs too.
- Tiberium 2y agoWanted to try with o1 and o1-mini but looks like there's no code available, although I guess I could just ask 3.5 Sonnet/o1 to make the evaluation suite ;)
- jerpint 2y agoAuthor here: all the code to reproduce this is actually all on the huggingface space here [1] https://huggingface.co/spaces/jerpint/advent24-llm/tree/main https://huggingface.co/spaces/jerpint/advent24-llm/tree/main
- Tiberium 2y agoThanks, I'll check with o1-mini and o1 and update this comment :) Also, the code has some small errors, although those can be easily fixed.
- zaptheimpaler 2y agoI’m adjacent to some people who do AoC competitively and it’s clear many of the top 10 and maybe 1/2 of the top 100 this year were heavily LLM assisted or wholly done by LLMs in a loop. They won first place on many days. It was disappointing to the community that people cheated and went against the community’s wishes but it’s clear LLMs can do much better than described here
- davidclark 2y agoI’d like the same article topic but from the person who did Day 1 pt1 in 4s and pt2 in 5s (9s total time).
- rhubarbtree 2y agoNo, that means it's clear that LLM-assisted coding can do better than described here. Which implies humans are adding a lot.
- NitpickLawyer 2y agoTo be fair, there's probably not a lot that "humans" add when the solution is solved in 4 and 5 seconds respectively. That's clearly 100% automated. Most humans can't even read the problem in that timeframes. A better implication would be that proper use of LLMs implies more than OP did here (proper prompting, looping the answer w/ a code check, etc.)
- zaptheimpaler 2y agoI think you're trying to dunk on me without actually knowing much about the matter. You can look at the leaderboard times and the github repos too. They are fully automated to fetch input, prompt LLMs and submit the answer within 10 seconds.
- rhubarbtree 2y agoSome of them, yes.
- bawolff 2y agoI'm a bit of an AI skeptic, and i think i had the opposite reaction of the author. Even though this is far from welcoming our AI overlords, I am surprised that they are this good.
- cheevly 2y agoGenuinely terrible prompt. Not only in structure, but also contains grammatical errors. I'm confident you could at least double their score if you improve your prompting significantly.
- rhubarbtree 2y agoThis isn't very constructive criticism. You could improve it through a number of levels: 1. You could have pointed out the grammatical errors and explained why they matter. 2. You could have pointed out the structural errors, and explained what structure should have been used. 3. You could have written a new prompt. 4. You could have re-run the experiment with the new prompt. Otherwise what you say is just unsubstantiated criticism.
- demirbey05 2y agoo1 is not included, I think each benchmark should include o1 and reasoning models. o-series is really changed the game.
- airstrike 2y agoI like the idea, but I feel like the execution left a bit to be desired. My gut tells me you can get much better results from the models with better prompting. The whole "You are solving the 2024 advent of code challenge." form of prompting is just adding noise with no real value. Based on my empirical experience, that likely hurts performance instead of helping. The time limit feels arbitrary and adds nothing to the benchmark. I don't understand why you wouldn't include o1 in the list of models. There's just a lot here that doesn't feel very scientific about this analysis...
- Recursing 2y agoThe article and comments here _really_ underestimate the current state of LLMs (or overestimate how hard AoC 2024 was) Here's a much better analysis from someone who got 45 stars using LLMs. https://www.reddit.com/r/adventofcode/comments/1hnk1c5/results_of_a_multiyear_llm_experiment/ https://www.reddit.com/r/adventofcode/comments/1hnk1c5/resul... All the top 5 players on the final leaderboard https://adventofcode.com/2024/leaderboard https://adventofcode.com/2024/leaderboard used LLMs for most of their solutions. LLMs can solve all days except 12, 15, 17, 21, and 24
- nindalf 2y agoThis needs to be higher. Not only because it shows that LLMs can do better than what OP says, but also that there’s some difference in how they’re used. Clearly the person on reddit was able to use them more effectively.
- danielbln 2y agoThat's the crux with most discussions on here regarding LLMs. "I used gpt-4o with zero shot prompts and it failed terribly!" "I used Claude/o1/o3, I fed various bits of information into the context and carefully guided the LLM" Those two approaches (there are many more) would lead to very different results, yet all we read here are comments giving their opinions on "LLMs", as if there is only one LLM and one way of working with them.
- d0mine 2y agoThis reminds me of https://en.wikipedia.org/wiki/Stone_Soup https://en.wikipedia.org/wiki/Stone_Soup stone->your ingredients->soup LLM->your prompts->solution
- jonathan_landy 2y agoSeems to depends strongly on the model perhaps. The Reddit post says “Some other models tested that just didn't work: gpt-4o, gpt-o1, qwen qwq.” Notably gpt-4o was used in the post linked here.
- guerrilla 2y agoHow far can LLMs get in Project Euler without messing up?
- antirez 2y agoThe most important thing is missing from this post: the performance of Jerpint+Claude. It's not a VS game.
- bru3s 2y ago[flagged]