3 ms·
> The Chinese models are mostly very well documented in terms of architecture and training processes/flows, with what is missing to recreate them being the trai
by throw10920 2mo ago
> The Chinese models are mostly very well documented in terms of architecture and training processes/flows, with what is missing to recreate them being the training data.
...and because that training data is missing, they can't be replicated. Which means that you cannot assert that the Chinese are being open in their LLM development, because there's no way to verify that the techniques they describe are actually the ones being used.
The reason that the training data is missing is that they're trained on a large amount of American copyrighted data and distilled on American models, which is where a lot of their performance comes from.
- HarHarVeryFunny 2mo agoYou can replicate the architectural innovations, and try them for yourself with your own dataset. It seems some of them are certainly being used by western companies, such as DeepSeek Sparse Attention, now supported by NVIDIA cuDNN. Ditto for training algorithms and procedures such as Slime or DeepSeek's details instructions on how to build a reasoning model. This is the exact value of openly shared details - others CAN copy and try them and modify them themselves. Yes, the training data specifically has not been released for any model, American or Chinese, but that doesn't detract from what has been shared, and the reason the Chinese are not sharing data are no more nefarious than why the American companies are not sharing - because they are all using data from sources they don't want you to know about, and at the end of the day the data is the closest thing any of them do have to a moat.
- throw10920 2mo ago> You can replicate the architectural innovations, and try them for yourself with your own dataset. That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data. > and the reason the Chinese are not sharing data are no more nefarious than why the American companies are not sharing This is moving the goalposts. Your claim was that "The Chinese have actually been very open about training", which is false, as discussed. Nobody ever claimed that the American labs were open.
- HarHarVeryFunny 2mo ago> You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data. Why would you be concerned about THEIR model performance ?! Surely if you are an ML researcher and read about a new technique, you are interested in how YOU may be able to use it. If you are Ilya Sutskever sitting at OpenAI in 2017, and happen upon Google's "attention" (Transformer architecture) paper, then what you do is go and implement it for yourself, get yourself some training data, and try it. What you are NOT going to do is whine about not being given their source code, or their training data, or their training harness, or a dump of Google's corporate secrets. You take the research that has been shared and evaluate it for yourself.
- throw10920 2mo agoThis doesn't have anything to do with Google or OpenAI. I'm not Ilya Sutskever and it's not 2017, either. I'm not "whining" about anything. You made the claim "The Chinese have actually been very open about training" and I showed that that was false. That's all that there is to it.
- HarHarVeryFunny 2mo agoThe rest of the word is reading Chinese published research and benefiting from it. Apparently you are unaware of it and not benefiting from it. Oh well.