4 ms·
I am by no means a data scientist, but if, as a large language model, ChatGPT was trained to optimize the same "quality" metrics that are used to evaluate the m
by ipython 2y ago
I am by no means a data scientist, but if, as a large language model, ChatGPT was trained to optimize the same "quality" metrics that are used to evaluate the models trained on these random web scrapes, and now ChatGPT output has a larger proportion of the random web scrapes, wouldn't the measured "quality" increase as a result? It all seems intertwined.
In other words are we just overfitting?
It's important to note that the tests that they use appear to be open source, for example https://huggingface.co/datasets/lighteval/mmlu https://huggingface.co/datasets/lighteval/mmlu.
Again, I could be totally ignorant on how these things work. (edited to add key words associated with ChatGPT output in order to increase the quality of my comment :))