4 ms·
Either way, shouldn't a company the size of Anthropic be scrubbing benchmarks from their training set regardless of how the benchmark data got into the training
by PostOnce 2y ago
Either way, shouldn't a company the size of Anthropic be scrubbing benchmarks from their training set regardless of how the benchmark data got into the training set?
- jnhl 2y agoUntraining models is an open research question and not par for the course like you're implying
- PostOnce 2y agoNo, not "untraining" after the fact, but scrubbing the benchmark data from the dataset prior to training. That's absolutely trivial. Take the set of benchmarks you care about, and then as you build a training dataset (by scraping, or whatever), you scrub each new item for benchmark questions (or just discard that entire item/webpage/whatever for being contaminated). Otherwise you're negligently and willingly inflating your benchmark performance and defrauding new investors / users / customers, I think.
- dialup_sounds 2y agoThe string itself is a fact, not a benchmark.
- PostOnce 2y agoThe string is to indicate that it was found alongside the benchmark data, that's the whole point. If that string shows up, then it's extremely likely the test / benchmark data was there too, indicating it's contaminated. In other words, if that string is present, the benchmark results for that model are a lie.
- dialup_sounds 2y agoI don't think you understand. The string is not a secret and it appears in places other than the benchmark data. Telling you the factual answer to the question of what the BIG bench canary string is not a "gotcha!" moment.
- PostOnce 2y agoThe actual valid counterargument is maybe Anthropic does't care about bigbench and nor should they have to, google can make up any benchmark it wants, anthropic does't have to use or care about it, or ignore it, or anything else. However, let's assume they should care because it's a major benchmark from an industry leader. The entire point of the canary string is that LLMs are supposed to discard / ignore it and data found on the same document. After all, the documentation literally says ""Do not edit the canary comments. These are to prevent BIG-bench tasks from leaking into web-scraped training data."" Anthropic did not do that (they obviously HAVE scraped data containing the GUID), therefore it is demonstrably a gotcha. e.g. it should ignore both https://github.com/google/BIG-bench/blob/main/bigbench/benchmark_tasks/date_understanding/task.json https://github.com/google/BIG-bench/blob/main/bigbench/bench... AND https://github.com/google/BIG-bench/tree/main https://github.com/google/BIG-bench/tree/main Even though the latter is the readme, it has the guid, and there's no reason not to ignore every document containing the GUID. So, if Anthropic wants to ignore it, fine, but it still feels a little fishy, doesn't it?
- jjcm 2y ago> Even though the latter is the readme, it has the guid, and there's no reason not to ignore every document containing the GUID. I disagree with this. There are plenty of websites out there that talk about LLM training in general, and have sections dedicated to canary strings. This page for example has the GUID in it: https://ravinkumar.com/GenAiGuidebook/deepdive/BigBench.html https://ravinkumar.com/GenAiGuidebook/deepdive/BigBench.html I'd argue that it is something that LLMs should train on. Having context of how LLMs work is something that isn't related to the benchmark data at all. Just because the GUID shows up as an example doesn't mean the benchmark data is present on the page.
- PostOnce 2y agoEven so, they can drop from their training set a few dozen pages from the internet and it won't be a huge loss among the trillions of documents. The LLM will still know how LLMs work without having trained on the handful of documents containing that specific canary string, because other documents will mention the concept of a canary string without that exact GUID. Better to do that and be on the safe side and look honest than have people believe your company not really competitive. Anything else is a risk for no gain in an industry theoretically worth trillions.