3 ms·
I think there is some kind of "bias" in how we test LLMs. Most (all?) benchmarks, either those coming from the industry and academia or those end-users like you
by mqefjh 2y ago
I think there is some kind of "bias" in how we test LLMs. Most (all?) benchmarks, either those coming from the industry and academia or those end-users like you or me may run on a couple examples all seem to compare LLMs answers with expected answers. This doesn't capture the extent to which one can augment his abilities using LLMs for cases where we don't know the expected answer (and this is precisely why we often turn to LLMs).
One instance of this was when I was able to extend features of a partial std::functional port for the AVR platform and was able to achieve my goals by asking ChatGPT to generate rather complex C++ template code that would have taken me several days to figure out since I'm not a C++ programmer. In about two hours and several back-n-forth between the code and ChatGPT's interface, I was able to integrate the modifications (about 50 LOCs) and save me the daunting and frustrating task of rewriting around 2000 LOCs I had written for the espressif platform. This is what I would have done if ChatGPT wasn't around.
In this context, look-good-but-broken examples are not really a problem when you can identify what is wrong and communicate the problems back to GPT. These cases do not bode well when we are asserting the correctness and autonomy of AI systems, but they are not as problematic when one seeks to augment his own abilities.