4 ms·
AoC questions this year were _deliberately written_ to be confusing to LLMs, it's not failing because it's worse it's failing because the questions were written
by inglor 3y ago
AoC questions this year were _deliberately written_ to be confusing to LLMs, it's not failing because it's worse it's failing because the questions were written to make it hader for models :]
Edit: apparently not, the author is just really good at coming up with ai adverse puzzles. When testing ChatGPT did much better on last year’s puzzles.
- sanxiyn 3y agoThat's fascinating! Where can I read more about it?
- cormacrelf 3y ago[flagged]
- usrme 3y agoThe article actually links to a Reddit post[1] where the creator of Advent of Code says that 2023's puzzles weren't engineered to be more difficult for LLMs. --- [1]: https://old.reddit.com/r/adventofcode/comments/18bp8id/why_does_aoc_care_about_llms/kc646i4/ https://old.reddit.com/r/adventofcode/comments/18bp8id/why_d...
- sd9 3y agoExcept the article says the opposite: Some people have even speculated that the problems this year were deliberately formulated to foil ChatGPT, but Eric actually denied that this is the case. Citing the author of AoC: Here are things LLMs didn't influence: The story. The puzzles. The inputs. I did the same thing this year that I do every year: I picked 25 puzzle ideas that sounded interesting to me, wrote them up, and then calibrated them based on betatester feedback. https://old.reddit.com/r/adventofcode/comments/18bp8id/why_does_aoc_care_about_llms/kc646i4/ https://old.reddit.com/r/adventofcode/comments/18bp8id/why_d...
- anonzzzies 3y agoI thought they were as well, because there were nuances which made them harder for LLMs. When I tried on 1 dec, gpt4 couldn't solve day 1 part 2. And I tripped over exactly the same thing ; when I parsed it correctly and then explained to the llm, it did solve it. But no where near as fast as me with my bag of horrible hacking aoc lib... It's interesting it turns out this year was not written with gpt in mind.
- adamors 3y agoHere's the creator of AoC saying the exact opposite on Reddit https://old.reddit.com/r/adventofcode/comments/18bp8id/why_does_aoc_care_about_llms/kc646i4/ https://old.reddit.com/r/adventofcode/comments/18bp8id/why_d... > Here are things LLMs didn't influence: > The story. > The puzzles. > The inputs. > I don't have a ChatGPT or Bard or whatever account, and I've never even used an LLM to write code or solve a puzzle, so I'm not sure what kinds of puzzles would be good or bad if that were my goal. Fortunately, it's not my goal - my goal is to help people become better programmers, not to create some kind of wacky LLM obstacle course. I'd rather have puzzles that are good for humans than puzzles that are both bad for humans and also somehow make the speed contest LLM-resistant. > I did the same thing this year that I do every year: I picked 25 puzzle ideas that sounded interesting to me, wrote them up, and then calibrated them based on betatester feedback. If you found a given puzzle easier or harder than you expected, please remember that difficulty is subjective and writing puzzles is tricky.
- raincole 3y agoGood example of "the best way to get a correct answer on the internet is to post a wrong one."
- fragmede 3y agoPoe's law strikes again!
- maximus-decimus 3y ago> calibrated them based on betatester feedback. At least one betatester might very well have been using chatgpt.
- imjonse 3y agoHe may not have consciously written them to confuse LLMs but even without using any he probably knows they get more confused on reasoning tasks stated in not very clear text. At least for a few problem descriptions I couldn't help feeling that the statement was a lot more complex than it could have been and not only because of the story woven around it. Of course it could have been there to confuse the humans :)
- radres 3y agoofc, chatgpt can do better in previous years given that solutions discussed all over the internet and could've train on them.
- mtlmtlmtlmtl 3y agoYeah there are literally Reddit megathreads with hundreds of people sharing their code.
- deleted 3y ago[deleted]
- danielbln 3y agoI think you should maybe be less confident in your statements if you actually don't really know. Highlighting _deliberately written_ when it was anything but.. well, that's confident bullshitting, LLM style, ironically.
- deleted 3y ago[deleted]
- blibble 3y ago> Edit: apparently not, the author is just really good at coming up with ai adverse puzzles. aka new unique problems that aren't in its training set