3 ms·
The point isn't that it's a secret, it's that it's a string that shouldn't appear in any context besides the benchmark. It's an easy way to identify contaminati
by PollardsRho 2y ago
The point isn't that it's a secret, it's that it's a string that shouldn't appear in any context besides the benchmark. It's an easy way to identify contamination in datasets. People often don't explicitly repeat them because that leads to false positives.
- gpm 2y agoAh, you're right, I should have clicked through a few more links. I was assuming it was a needle that they were asking if the model could recall after hiding it in a bunch of text, not a canary for the test data being included in the training set at all. Still, I think it's less than clear that the problem here isn't just the idea that you can publish benchmark data on the open web at all without making the benchmark outdated. Expecting secret (to the model) benchmark data to be reliably detected and filtered from training data seems... unlikely. Especially in this day and age where lots of people dislike AI enough that they would be interested in actively sabotaging efforts...