4 ms·
The article and comments here _really_ underestimate the current state of LLMs (or overestimate how hard AoC 2024 was) Here's a much better analysis from someo
by Recursing 2y ago
The article and comments here _really_ underestimate the current state of LLMs (or overestimate how hard AoC 2024 was)
Here's a much better analysis from someone who got 45 stars using LLMs. https://www.reddit.com/r/adventofcode/comments/1hnk1c5/results_of_a_multiyear_llm_experiment/ https://www.reddit.com/r/adventofcode/comments/1hnk1c5/resul...
All the top 5 players on the final leaderboard https://adventofcode.com/2024/leaderboard https://adventofcode.com/2024/leaderboard used LLMs for most of their solutions.
LLMs can solve all days except 12, 15, 17, 21, and 24
- nindalf 2y agoThis needs to be higher. Not only because it shows that LLMs can do better than what OP says, but also that there’s some difference in how they’re used. Clearly the person on reddit was able to use them more effectively.
- danielbln 2y agoThat's the crux with most discussions on here regarding LLMs. "I used gpt-4o with zero shot prompts and it failed terribly!" "I used Claude/o1/o3, I fed various bits of information into the context and carefully guided the LLM" Those two approaches (there are many more) would lead to very different results, yet all we read here are comments giving their opinions on "LLMs", as if there is only one LLM and one way of working with them.
- d0mine 2y agoThis reminds me of https://en.wikipedia.org/wiki/Stone_Soup https://en.wikipedia.org/wiki/Stone_Soup stone->your ingredients->soup LLM->your prompts->solution
- jonathan_landy 2y agoSeems to depends strongly on the model perhaps. The Reddit post says “Some other models tested that just didn't work: gpt-4o, gpt-o1, qwen qwq.” Notably gpt-4o was used in the post linked here.
- jebarker 2y agoI don't know what they were doing, but I tried o1 with many problems after I solved them already and it did great. No special prompting, just "solve this problem with a python program".
- FrustratedMonky 2y agoWonder if different goals. If the top 5 people on leader board, 'used LLM's'. Meaning, they used an LLM as a helping tool. But article I think is, what if the LLM played by itself. Just paste the questions in, and see if it can do it all on its own? Perhaps that is the difference, different goals.
- oytis 2y agoWhy does neither of articles provide the actual raw chat logs? It's like a recent article about a non-released LLM solving non-public tasks which everyone is supposed to be impressed about
- michaelt 2y ago> All the top 5 players on the final leaderboard [...] used LLMs for most of their solutions. Note that the leaderboard points are given out on the time taken to produce a correct answer. 100 points for the first to submit an answer, 99 for the second, 98 for third and so on. No points if you're not in the first 100. So if an LLM fails on 5 problems, but for the other 20 it can take a 600-word problem statement and solve it in 12 seconds? It'll rack up loads of points. Whereas the best human programmer you know might be able to solve all 25 problems taking around 15 minutes per problem. On most days they would have zero points, as all 100 points finishes go to LLMs.
- jerpint 2y agoAuthor here: the point of the article was only to evaluate zero-shot capabilities. I’m certain that had I used LLMs I would have definitely gotten more stars on AoC (got 41/50 without). Because I chose to solve this year without LLMs, I was simply curious to see how the opposite setup would do, using basically zero human intervention to solve AoC. That said, if I cared only to produce the best results I would 100% pair up with an LLM
- WorldWideWebb 2y agoAm I old and grumpy or does that kinda go against the whole point of AoC? A daily challenge for your brain, not your prompt writing abilities.
- poincaredisk 2y agoThat's humanity for you. We automate thinking out of our lives so we can doomscroll unbothered.
- 383toast 2y agoYep AoC explicitly discourages using LLMs for leaderboard scoring
- 383toast 2y ago> All the top 5 players on the final leaderboard https://adventofcode.com/2024/leaderboard https://adventofcode.com/2024/leaderboard used LLMs for most of their solutions. Source for this?
- steolan 2y agobecause the timings were not possible for humans. problems were solved in under 2 minutes. I've experimented with LLM for few days too. yielded the same results.