2 ms·
IMO a fair evaluation needs to include a few extra steps, where the model generates a response, and then reflects on that response, using basic prompts like "Di
by EnergyAmy 2y ago
IMO a fair evaluation needs to include a few extra steps, where the model generates a response, and then reflects on that response, using basic prompts like "Did you make sure to read the question carefully and look for trick questions?".
Imagine if you were tested in a rapid-fire manner and evaluated based on whatever answer you blurted out first. Human reasoning scores would suffer dramatically as well. There would probably be a lot of answers involving wolves and cabbages.
You might think that prompting the LLM with questions like "Make sure it's not a trick question" is cheating, but it's very similar to how humans work. A lot of people would answer wrongly if they encountered the question "What weighs more, a pound of bricks or two pounds of feathers?", because they also just pattern match and assume the answer is the question they've seen many times before. If they're primed by someone saying "Read the question carefully", they'll do a lot better.
- Hugsun 2y agoYou might enjoy the analysis in the article. https://www.arnaldur.be/writing/about/large-language-model-reasoning#recognizability-traps https://www.arnaldur.be/writing/about/large-language-model-r... There is an expandable box with all the questions, colored by correctness, and question 6 has a bunch of responses. Below the box is a summary of the results.