4 ms·
Disagree - a model may have capabilities that it only deploys in certain circumstances. In particular, if you train on the whole internet, you probably learn wh
by emtel 3y ago
Disagree - a model may have capabilities that it only deploys in certain circumstances. In particular, if you train on the whole internet, you probably learn what both stupid and smart outputs look like. But you probably also learn that the modal output is stupid. So it’s no ding on the model’s capabilities if it defaults to assuming it should behave stupidly.
- PumpkinSpice 3y agoWhat you're basically saying is "the model as trained can't do well at this task, so let's use our own cognitive skills to help it." That's a problem for comparative benchmarking, right? You're no longer testing the model; you're testing the model in tandem with the prompt engineer. This raises several big questions: 1) How do you know that the engineer's knowledge of the "correct" answer isn't being subtly encoded in the prompt? Essentially the Clever Hans phenomenon. 2) If you want this to be a fair game, how do you give precisely the same kind of advantage to whatever you're benchmarking the LLM against? 3) Last but not least, if you're not going to throw in your prompt engineer for free with the product, will your results be reproducible by your customers? To be clear, I don't think there's some cosmically objective way of doing this. If you're using prompts written for humans, you're already putting the computer at a disadvantage. But at the very least, you're measuring something meaningful: how will the model behave in the real world.