4 ms·
While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people do
by remus 2mo ago
While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).
ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.
- ryandvm 2mo agoI don't know man. This reads like "yeah we stole your grain, but making bread is hard."
- Muromec 2mo agoIt sure is, but it doesn't matter. Whatever position that generates more economic activity is declared legal using some nonsense retconned logic "because we said so".
- darod 2mo agoYou can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.
- some_random 2mo agoDistillation is objectively easier than training a model from scratch, that's why all these Chinese labs are doing it.
- skippyfish 2mo ago> raining a SOTA model takes a huge amount of resources and expertise Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.
- Teever 2mo agoI'm sure it takes a lot of time and resources to plan and pull off an epic heist but it is unusual to see people like Thomas Crown being accused of creating value, as they're usually accused of committing theft.
- il 2mo agoProbably not as much effort as writing books and creating art the models were trained on.
- jaggederest 2mo agoI suspect that, in aggregate, all of the informational output of humanity prior to 2020 has taken more resources to produce than the last few years of LLM research.
- deleted 2mo ago[deleted]
- altmanaltman 2mo agoWhy is it less true for distillation? Everyone technically has access to Fable but Moonshot came up with the model. How can you objectively claim one is adding value while the other is not? If that is the whole point you need to clarify why this is the case on an objective level. I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.
- deleted 2mo ago[deleted]
- Bratmon 2mo agoI like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value. It's the most CS-major take ever!
- joshuamorton 2mo agoI don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.
- remus 2mo agoYes, this is what I was getting at.
- gozucito 2mo agoThere is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot. There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible. Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.
- joshuamorton 2mo agoSo, I'm not the person you were responding to. I'd like you to take a moment and suggest where anyone in the thread you're replying to, either me or Remus, has said anything that suggests disagreement with the statement > There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?
- AlienRobot 2mo agoThe value of LLM's come from replacing what generated its training data. If the distilled model is cheaper, then it's just LLM's getting LLM'ed.
- InsideOutSanta 2mo agoAs an author, that's a genuinely disheartening thing to read. It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?
- blks 2mo agoThey add value on top of other people’s work, often against licensing, and then commercialize this product, ie profiting from making a product out of other people’s IP.
- liuliu 2mo ago> training an LLM takes more resources and expertise than distilling from an existing LLM This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.
- mort96 2mo agoYeah, there's a difference. One party spends a bunch of resources doing something illegal and extremely immoral. The other party spends little money doing something legal and morally neutral.
- GTP 2mo agoStill, AFAIK Kimi's architecture (just like that of other LLMs from Chinese labs) is different from those of OpenAI and Anthropic's model in a nontrivial way. So the expertise is still there, and I guess resource use too (although Chinese labs tend to optimize this, thanks to the restrictions they have on GPU use). EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.
- Barrin92 2mo ago>Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.