5 ms·
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't perso
by AustinDev 2mo ago
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
- michelsedgh 2mo agoWhat an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
- AussieWog93 2mo agoI'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.
- dullcrisp 2mo agoWhy isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?
- ygjb 2mo agoBecause that doesn't satisfy the narrative of American exceptionalism. It's easier to point at something and say it was stolen or copied than it is to compete, especially with the political climate in the US. This isn't an anti-American sentiment. It is an anti-corporate/regulatory capture/embrace and extinguish sentiment (which probably reads the same to many people these days).
- Eldt 2mo agoBecause the companies did it underhandedly without prior consent? Imagine you walked into a hardware store to grab some lumber, didn't pay for it, and the store had to call the cops to swing around your house? That's hardly the typical shopping experience, now is it?
- dullcrisp 2mo agoI think that’s my point: none of those things happened, so on what basis can we claim that it’s like they did. Who has to agree with what you’re doing in order for it not to be considered underhanded?
- Gigachad 2mo agoIf the legal system declares the first thief’s theft not theft then all bets are off.
- mannanj 2mo agoIs it theft if another thief steal's the first thief's theft?
- BeetleB 2mo ago> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
- kennywinker 2mo agoAs i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.
- TheOtherHobbes 2mo agoCopyright law only considers illegal ownership of a work, so the crime - or tort - was making/acquiring copies without permission or payment. Training from copies has been ruled fair use because it's "transformative" and not simply "derivative." This is obviously debatable, but that's where the debate is at the moment.
- michelsedgh 2mo agoSo basically because they just browsed and used the information that was mostly public on the internet and they didnt copy it, they just learned from it and thats fine. Which makes sense. None of the llms let u copy someones work exactly anyways... makes total sense honestly. So in this case what happens to distilling? Is that also learning or ur trying to get to their actual weights by kind of reverse engineering it? Where would the argument fall there?
- itemize123 2mo agoquestion's phrasing made your judgement obvious
- michelsedgh 2mo agoI seriously wasnt judging i maybe shouldve asked llm to frame it better cause i knew it might sound that way, thats why i added: im not judging, merely asking...
- ofjcihen 2mo agoI gathered that the most recent advances haven’t been in capabilities of the model but more the way that it’s able to be employed (most recently agents).
- deleted 2mo ago[deleted]
- conception 2mo agoIf you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.
- FuckButtons 2mo agoThat’s definitely my impression of deepseek 0731 after a fair bit of use via ds4, it sounds like Claude.
- bossyTeacher 2mo ago> LLM providers can distill all of human output into their models for 'free' Not sure what part of being charged guilty and paying a fine you see as "free".
- rlupi 2mo agoI think it's more likely to be the effect of synchronization of launches, and the fact that models that do not challenge SOTA in some way do not get launched (think Gemini Pro delays), launched quietly or do not get any attention.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]