3 ms·
[Author] Can you be specific about which part of the research is dishonest? We've shared the code openly, you can reproduce it yourself if you want.
by waleedk 3y ago
[Author] Can you be specific about which part of the research is dishonest?
We've shared the code openly, you can reproduce it yourself if you want.
- BoorishBears 3y agoI think in a vacuum it's not dishonest, but it is half-baked in a way a lot of these benchmarks are. Time and time again we see benchmarks like this where 0 effort into actually trying to get anything good out of the model. Of course you can make extremely specific changes that essentially game the test, so to me a fair way to test this is to establish a realistic token budget across all the models, maybe even adjusted for cost depending on the goals of your testing. Then for each model, use your budget to get the highest quality result possible. _ You're asking the GPT models for example, in the worst possible way. You've moved detailed instructions into the user prompt, despite their newest updates focusing on steerability via the system. Meanwhile from tinkering LLAMA 2 pretty much doesn't care about the system prompt vs user prompts. You're also not allowing for any form of chain-of-thought. Giving the models a few hundred tokens to form a conclusion instead of spending those tokens trying to force out a mathematically induced order bias (which I wouldn't expect to do much) would have been much better. There's also no way that GPT 4 should have struggled to give you a well defined output: All of the models would have probably benefited from a well defined output format in terms of a schema, rather than asking for a single letter response. The model rarely has to produce a single letter answer and will struggle with that. Something like asking for { "output" : "A_IS_MORE_FACTUAL" | "B_IS_MORE_FACTUAL" } would have been better Overall I think there's no open dishonesty, and obviously there's no objective correct amount of effort to put into the test setup here... but there's a conflict of interest that would have made me want to see more effort put into it. I think there were trivially low hanging fruit ignored here that I'd expect people selling LLMOps to have pick up on.