4 ms·
the problem with the critique of "oh they didn't use the correct prompts" is that prompt engineering is highly dependent on the model. You could technically cre
by chewxy 3y ago
the problem with the critique of "oh they didn't use the correct prompts" is that prompt engineering is highly dependent on the model. You could technically create an LLM that would not work with the "let's think this through step-by-step" magic prompt (i.e. exclude anything with similar phrases in the pretraining dataset).
Yes, they used GPT3.5-turbo, which would have its set of magic key phrases. Should they have used it? I'd say probably not.
- aeternum 3y agoRight, this is why my critique is that it is a weak paper in general. It's misleading to make claims about "LLMs" based on experiments with a single LLM. This is made worse by testing with very few prompt variations.