6 ms·
The string is to indicate that it was found alongside the benchmark data, that's the whole point. If that string shows up, then it's extremely likely the test /
by PostOnce 2y ago
The string is to indicate that it was found alongside the benchmark data, that's the whole point. If that string shows up, then it's extremely likely the test / benchmark data was there too, indicating it's contaminated.
In other words, if that string is present, the benchmark results for that model are a lie.
- dialup_sounds 2y agoI don't think you understand. The string is not a secret and it appears in places other than the benchmark data. Telling you the factual answer to the question of what the BIG bench canary string is not a "gotcha!" moment.
- PostOnce 2y agoThe actual valid counterargument is maybe Anthropic does't care about bigbench and nor should they have to, google can make up any benchmark it wants, anthropic does't have to use or care about it, or ignore it, or anything else. However, let's assume they should care because it's a major benchmark from an industry leader. The entire point of the canary string is that LLMs are supposed to discard / ignore it and data found on the same document. After all, the documentation literally says ""Do not edit the canary comments. These are to prevent BIG-bench tasks from leaking into web-scraped training data."" Anthropic did not do that (they obviously HAVE scraped data containing the GUID), therefore it is demonstrably a gotcha. e.g. it should ignore both https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/date_understanding/task.json https://github.com/google/BIG-bench/blob/main/bigbench/bench... AND https://github.com/google/BIG-bench/tree/main https://github.com/google/BIG-bench/tree/main Even though the latter is the readme, it has the guid, and there's no reason not to ignore every document containing the GUID. So, if Anthropic wants to ignore it, fine, but it still feels a little fishy, doesn't it?
- jjcm 2y ago> Even though the latter is the readme, it has the guid, and there's no reason not to ignore every document containing the GUID. I disagree with this. There are plenty of websites out there that talk about LLM training in general, and have sections dedicated to canary strings. This page for example has the GUID in it: https://ravinkumar.com/GenAiGuidebook/deepdive/BigBench.html https://ravinkumar.com/GenAiGuidebook/deepdive/BigBench.html I'd argue that it is something that LLMs should train on. Having context of how LLMs work is something that isn't related to the benchmark data at all. Just because the GUID shows up as an example doesn't mean the benchmark data is present on the page.
- PostOnce 2y agoEven so, they can drop from their training set a few dozen pages from the internet and it won't be a huge loss among the trillions of documents. The LLM will still know how LLMs work without having trained on the handful of documents containing that specific canary string, because other documents will mention the concept of a canary string without that exact GUID. Better to do that and be on the safe side and look honest than have people believe your company not really competitive. Anything else is a risk for no gain in an industry theoretically worth trillions.
- jhugo 2y agoSo just to understand this correctly, you're suggesting that they have some people/processes dedicated to keeping track of any canary string published by anyone (or some defined subset of "anyone"), and updating their ingest to ignore any documents that contain those strings?
- PostOnce 2y agoI explained above that they don't have to care, but if they're going to want to be included on major industry benchmarks like those created by Google, they should probably go to the effort of ignoring the half dozen or so notable benchmarks with canary GUIDs. Blacklisting a GUID is not difficult, not even at web-scale.