5 ms·
then do you have better way to more rigorously evaluate chatbot at the presence of LLMs like ChatGPT trained on almost all Internet data?
by zhisbug 4y ago
then do you have better way to more rigorously evaluate chatbot at the presence of LLMs like ChatGPT trained on almost all Internet data?
- cscurmudgeon 4y ago> running it through a _single_ of the many openly available language model benchmarks.
- zhisbug 4y agobut how could you guarantee the pretrained model haven't seen those benchmarks? And the baselines you are comparing to (chatgpt, bard) haven't as well? Cuz those benchmarking datasets are also collected from Internet right?
- CGamesPlay 4y agoAs a baseline, don't! If it performs horribly on the test and it cheated, that's even worse than if it fails the test and didn't cheat. So the benchmark score gives you an upper bound on performance.