3 ms·
Hi some authors of the work here, thanks a lot for sharing the paper, it's been quite some time in the work and we're super happy to share it with the world. A
by Thomjazz 3y ago
Hi some authors of the work here, thanks a lot for sharing the paper, it's been quite some time in the work and we're super happy to share it with the world.
A short note on some of the reasons we decided to go with openly-sharing the questions instead of holding them back (which was another option we contemplated):
- with closed-models we need to send the questions through an external AI anyway so a full privacy of the test set is not possible in general unless the leaderboard is restricted to open models (would be quite restrictive)
- also, the benchmark contain a limited number of questions which are non-obvious and take a significant time for human reviewers to solve. We thus don't expect the dataset to become training material for models and to lead to having model over-fitting on the benchmark pattern in the traditional sense that happened with larger benchmark datasets including a training split. This benchmark is generally closer in philosophy to small, hand crafted benchmark datasets, like HumanEval for instance has been for code models.