3 ms·
Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but
by culi 1mo ago
Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.
- azinman2 1mo agoThey don’t release all the code.
- culi 1mo agoDeepSeek has released a ton of low-level AI infrastructure code, libraries, and mathematical models on GitHub. Their tools and agent environments are fully open sourced including their harness. They release complete PyTorch and Hugging-face compatible python files detailing their configuration, tokenizers, and layers of the architecture. The only thing they don't release is their data-filtering pipelines but they detail even that in their public-access papers. DeepSeek is truly as open source as you can possibly legally get. Besides the data itself, it's completely reproducible by anyone else. I don't think americans yet acknowledge just how radically transparent Chinese labs are being (and how much even the west benefits from it).
- thepasch 1mo agoInference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open weights models) are still a complete blackbox, and I can't think of any Chinese model that made its training corpus public.
- culi 1mo agoThis is absolutely not true. DeepSeek is most famous for publishing really in-depth papers on their training process but the other labs have started to do the same as well. If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible. In fact everything I just said I said in my original comment. It's like you didn't read it at all. DeepSeek's GRPO Infrastructure, multi-stage training pipeline, and their "cold start" phase have been massively influential in LLM research.
- thepasch 1mo ago> If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible. Did you read the site this very post links to? The entire point is that the training corpus, recipe, and scripts, as well as intermediate checkpoints, will be made available for K2 Horizon. Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product. I've read your comment. I'm doubting you've even read the thing you were commenting on.
- culi 1mo agoYes there have been several research projects like these that have fully open sourced their training data (mostly coming from Europe). But that's all they amount to. Experiments and projects. The labs that are actually competing with frontier models are using data would usually be a violation of copyright to release openly. > Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product. No? That's absolutely not true. Qwen, GLM, Kimi, DeepSeek, etc all consistently release both the post-trained "Instruct/Chat" versions and the underlying "Base" (pre-trained) weights. Which specific Chinese models are you thinking about?