3 ms·
And here lies the exact issue. Single tests don’t provide any meaningful insights. You need to perform this test at least twenty times in separate chat windows
by s-macke 2y ago
And here lies the exact issue. Single tests don’t provide any meaningful insights. You need to perform this test at least twenty times in separate chat windows or via the API to obtain meaningful statistics.
For the "Alice in Wonderland" paper, neither Claude-3.5 nor o1-preview was available at that time.
But I have tested them as well a few weeks ago with the issue translated into German, achieving also a 100% success rate with both models.
However, when I add irrelevant information (My mother ...), Claude's success rate drops to 85%:
"My mother has a sister called Alice. Alice has 2 sisters and 1 brother. How many sisters does Alice's brother have?"
- Workaccount2 2y agoWe do have chatbot arena which to a degree already does this. I like to use: "Kim's mother is Linda. Linda's son is Rachel. John is Kim's daughter. Who is Kim's son?" Interestingly I just got a model called "engine test" that nailed this one in a three sentence response, whereas o1-preview got it wrong (but has gotten it right in the past).
- probably_wrong 2y agoYour experience makes me think that the reason the models got a better success rate is not because they are better at reasoning, but rather because the problem made it to their training dataset.
- andrepd 2y agoAbsolutely! It's the elephant in the room with these ducking "we've solved 80% of maths olympiad problems" claims!
- s-macke 2y agoWe don't know. The paper and the problem was very prominent at that time. Some developers at Anthropic or OpenAI might have included that in some way. Either as test or as a task to improve the CoT via Reinforcement Learning.
- meroes 2y agoIt made it into their data set via RLHF almost assuredly. Wild these papers are getting published when RLHF'ers and up see this stuff in the wild daily and ahead of the papers. Timeline is roughly: Model developer notices a sometimes highly specific weak area -> ... -> RLHF'ers are asked to develop a bunch of very specific problems improving the weak area -> a few months go by -> A paper gets published that squeezes water out of stone to make AI headlines. These researchers should just become RLHF'ers because their efforts aren't uncovering anything unknown and it's just being dressed up with a little statistics. And by the time the research is out, the the fixes are already identified internally, worked on, and nearing pushes. I just realized AI research will be part of the AI bubble if it bursts. I don't think there was a .com research sub-bubble, so this might be novel.
- andoando 2y agoYou also need a problem that hasn't been copy pasted a million times on the internet.