33 ms·
“Distillation attack” are we joking here. If anything these models should be compelled to be public since they have been trained off public data. What an absur
by onetrickwolf 3mo ago
“Distillation attack” are we joking here.
If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack.
It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat.
I generally really like Anthropic’s work and models but stuff like this scares me for the future. We are positioning these companies to have too much power. The public’s life is getting worse while these companies consolidate power using data they stole from the public.
- TZubiri 3mo agoTwo wrongs don't make a right
- tokioyoyo 3mo agoIn this scenario it does, because consumers win. Everyone in AI industry wants to fight dirty, but gets angry when their competitor fights dirty as well. And I’ve mentioned it before, how I generally like Ant and its products.
- moistoreos 3mo agoPretty sure the second rectified the first.
- justapassenger 3mo agoClosest analogy to distillation is api reimplementation, without which current software industry wouldn’t exist. There’s nothing fundamentally wrong with distillation.
- zobzu 3mo agoits mainly just a lot cheaper. copying is always cheaper anyway, very little r&d - ai or no ai.
- rafram 3mo agoThe core of the training data is public, but the part that actually makes these models smart came from (pretty highly-paid) experts via platforms like Mercor. Claude didn't magically learn to write good code by reading all of GitHub - humans trained it in that, more or less manually.
- jaen 3mo ago...and the rest of the training data (ie. the entire corpus of copyrighted works) was not written by experts expecting compensation? Double standards.
- rafram 3mo agoI didn't say that.
- jaen 3mo agoIndeed, that's exactly why I replied - you omitted one side from the discussion.
- thom 3mo agoNo, you just parroted an increasingly popular talking point, the entire purpose of which seems to be to absolve AI companies of the enormous theft that put them in the position to hire experts in the first place.
- rafram 3mo agoWell, I'd never heard anyone make it before, but sure. (I looked into Mercor a bit and know some people who've worked in data generation/labeling, which is what exposed me to that side of the operation.) It doesn't absolve them of any theft, but it does make the assertion that they should be required to release their models to the public seem, to me, a bit farcical. There are dozens of free and open-weights models that have all trained on exactly the same web crawls and books as GPT-5 and Opus. The proprietary models are better because of proprietary data.
- coliveira 3mo agoWhat they're trying to do under the umbrella of "national security" is to legislate how we can use the results we pay for when accessing these models. This way they will control the "intellectual property" that was acquired illegally.
- cma 3mo agoSince they hide their thinking traces it really doesn't make too much sense. We know one of their fixed degradations they talked about in a recent blog post was if you left claude code idle for too long they would rehydrate it without the thinking traces in the context and it degraded performance. So direct forms of distillation wouldn't be expected to get as good of results as they are getting. However, they could have used it as a judge etc. during training.
- _fat_santa 3mo ago> If anything these models should be compelled to be public since they have been trained off public data I'm starting to come around to this idea TBH. For a while my position was: "these companies have invested billions into training these models, therefore they should be able to control them and profit off them" but looking deeper at where they got their training data, my view is starting to shift. IMHO I feel like we need new laws around AI, specifically training data. Something like: "you can train an AI model and ignore copyright laws, BUT you must then make the model open weight", a company can still develop closed weight models but then they must aquire permission to use training data. But it gets murky because if something like that was on the books then AI labs would just train open weight models and then distill them into their closed weight models.
- ivanovm 3mo agolabs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated
- islandfox100 3mo agoThen it should be simple for one of the frontier labs to produce a model trained only on private data. We haven't seen that.
- wongarsu 3mo agoDidn't the famous "Textbooks are all you need" paper already proof that point three years ago? Sure, we ask a lot more of modern models, but private training data also got a lot better. You would loose out on a lot of long-tail knowledge, but that can be fixed with web search tools. You'd limit the styles, dialects and colloquial phrases the model understands and can use, but for many use cases that would be fine But why would any frontier lab do that? Throwing in more training data still leads to better results in pretraining. And showing that they don't need to hoover up the internet and Anna's Archive only empowers regulators to prevent them from doing that
- slibhb 3mo ago> If anything these models should be compelled to be public since they have been trained off public data. What an absurd overreach to call this an attack. > It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. If all that is required to train these models is public data, why can't Alibaba just use that? The fact that Alibaba has to resort to scraping Claude suggests there already is a moat...
- KerryJones 3mo agoThis feels more nuanced than you are giving it credit for? Much of the training data that was available has been withdrawn, atleast for OpenAI we know that much of the training data was garnered in less-than above the board methods
- petilon 3mo ago> If anything these models should be compelled to be public since they have been trained off public data. Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free? > What an absurd overreach to call this an attack. Would it be an attack to take your meal by force if you used a public recipe to prepare the meal?
- topgrain2 3mo ago> Isn't that a bit like saying if you read books in a public library to pick up a new skill you should work for free? Only if you’re trying to muddy the waters. No, obviously it’s not. One can also support licensing for driving a car on public roads but not for walking, even though both involve traveling. This is only confusing to people pretending to be confused, for effect. > Would it be an attack to take your meal by force if you used a public recipe to prepare the meal? “You wouldn’t download a car…” (unless it worked like copying an MP3, then, of course, you would, everyone would) It’s as if you’re using terrible analogies and comparisons because stronger ones don’t exist. Great news for the AI-should-be-open crowd.
- petilon 3mo agoI think the analogies are appropriate. Anthropic took public data and added value on top of it. It is that added value that Alibaba is targeting. If it was the underlying data, that's freely available.
- runtime_terror 3mo agoIf by "public domain data" you mean stealing ungodly amounts of copyrighted works then sure
- topgrain2 3mo agoAlibaba's asking for things, and receiving what they asked for. > If it was the underlying data, that's freely available. A bunch of it is not, but was pirated. And "underlying data"—JFC, that's billions of person-hours of thoughtful work by real people, practically infinitely more worthy of respect and care than what these LLM companies have done, without which they would have nothing. Alibaba's being more above-board about this than the major American firms have been (are they in general? Oh no, I doubt it, but in this particular case, yes). Extra accounts to get around TOS restrictions is the lesser evil here, and it's being done to companies that did worse. This is the least they should suffer, and their complaining about it is as comical as a professional fence crying about how unfair it is their shop got burgled. Live by the sword...
- flowerlad 3mo agoShould Google search index be forced to be public too?
- calgoo 3mo agoHonestly, yes it should in some form. If their index contains the actual data from the sites, and they are making that information public in one way or another, then it should be available as a downloadable dataset.
- rapind 3mo ago> It’s clear they are scapegoating national security and China at this point to build an anti-competitive moat. They are also fear mongering (and getting shills to as well) the idea that once open weight (Chinese) models catch up to Mythos we're all doomed. Maybe I'd be bit less cynical if they weren't prepping for IPO? Wasn't OpenAI spreading similar FUD back when GPT 2 came out? Guys... AGI is right around the corner. Pinky swear. Now buy our stock. Keep in mind that the entire US economy is currently propped up by AI spending, so a lot of people (banks, government) are incentivized to make sure these companies succeed. Expect this propaganda to ratchet up a notch if / when the economy starts to nose dive.
- ok123456 3mo agoYes. They're turning on the consent manufacturing machine to make it an issue of "national security" to download some gguf file from Hugging Face. Absolutely disgusting.
- rayiner 3mo ago> The public’s life is getting worse while these companies consolidate power using data they stole from the public How can you “steal” public information?
- calgoo 3mo agoreally? You know this just like everyone else: Just because the information is available publicly, does not mean that you can do whatever you want with the information. Copyright exists for a reason, and if the copyright lobby is going to continue to push for the poor poor media companies to keep their copyrights, then we should do the same towards the AI companies. So yes, they Stole the information from everyone else, and they keep doing so, as you can see their scanners still hitting every website on the web to get an updated dataset. It does not matter what they do AFTER they steal all the information, as they already stole it.
- deleted 3mo ago[deleted]
- msabalau 3mo agoThere's probably at 10-15% percent chance of a war between the US and China over the next 10 years. Maybe better than even chance of a militarized crisis that might have led to war, but somehow de-escalates. Regardless of how sad late stage capitalism makes you, or how outrageous one claims to find "hypocrisy", any national security argument about limiting Chinese AI capability stands on it's own, at least for nations likely to be drawn into a war. Also, all the local model enthusiasts who assume Chinese firms are going be allowed to endlessly release models if they have disruptive potential attributed to Mythos are probably in for a rude awakening. Just because the PRC is content about what has happened in the past doesn't mean that they would tolerate an open model that could be truly destabilizing.
- pseudony 3mo agoAs a third party I would rather be happy about the way Chinese labs are acting in the here and now while US labs first masquerade as a public good, then turn around, bail on all promises of open AI, turn into a corporation and attempt to own the world while its runner-up is trying to scaremonger people into buying their product. I know most Americans are fed a steady diet of “evil China” and China MAY have issues. But on the AI front they are heaps better. Even if everything got closed tomorrow, we have a plethora of good models we can inspect and tweak while from the US labs we have… a single old 120b model ? And with the way the US is treating its allies, maybe a bunch of us are quite content with a more even match rather than US hegemony.