3 ms·
The evidence is Anthropic's own reporting [1]. You may doubt that they're telling the truth, but that's what they're reporting. [1] https://www.anthropic.com/n
by v64 1mo ago
The evidence is Anthropic's own reporting [1]. You may doubt that they're telling the truth, but that's what they're reporting.
[1] https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- Gigachad 1mo ago"Distillation attacks", and it's just the exact same thing they did to the rest of the web and all human works.
- brookst 1mo agoExplain “exact same thing”? Like it load the Wikipedia page for Fort Worth, Texas, does it make Wikipedia spin up an editor to go write the article? I’m really not sure how you’re getting “exact” here.
- Gigachad 1mo agoAnthropic scraped the whole web, scanned every book, pirated every bit of media to feed in to their training. Distillers are doing essentially the same thing scraping all the knowledge from the LLM to create a training set for a new one. They are crying about theft after committing the largest theft in human history.
- colingauvin 1mo agoI'm supposed to believe that DeepSeek distilled a 300B-1T param model with 150,000 requests? Lol.
- brookst 1mo agoAre you thinking distillation goes from zero to complete model? I believe it can be used in the RL / fine tuning sense, in which case 150,000 requests, assuming every one was detected, could move the needle in quality. I agree it couldn’t replace all of pre and post training , but I don’t think that’s the claim. You do typical training, then distill really difficult cases.