3 ms·
This doesn’t address the primary issue: that they had no methodology for choosing puzzles that weren’t in the training set and indeed while they claimed to have
by bowsamic 1y ago
This doesn’t address the primary issue: that they had no methodology for choosing puzzles that weren’t in the training set and indeed while they claimed to have chosen puzzles that aren’t they didn’t explain why they think that. The whole point of the paper was to test LLM reasoning in untrained cases but there’s no reason to expect such puzzles to not part of the training set, and if you don’t have any way of telling if it is not or then your paper is not going to work out
- roywiggins 1y agoIsn't it worse for LLMs if an LLM that has been trained on the Towers of Hanoi still can't solve it reliably?
- bowsamic 1y agoYes
- anonthrowawy 1y agohow could you prove that?
- bowsamic 1y agoYou couldn’t, so such a paper cannot be scientific (Or it should not be based on that claim as a central point, which apples paper was)