7 ms·
Truly open source coming from China. This is heartwarming. I know if the potential ulterior motives.
by sidcool 6mo ago
Truly open source coming from China. This is heartwarming. I know if the potential ulterior motives.
- I_am_tiberius 6mo agoOpen weight!
- alecco 6mo agoPlease don't slander the most open AI company in the world. Even more open than some non-profit labs from universities. DeepSeek is famous for publishing everything. They might take a bit to publish source code but it's almost always there. And their papers are extremely pro-social to help the broader open AI community. This is why they struggle getting funded because investors hate openness. And in China they struggle against the political and hiring power of the big tech companies. Just this week they published a serious foundational library for LLMs https://github.com/deepseek-ai/TileKernels https://github.com/deepseek-ai/TileKernels Others worth mentioning: https://github.com/deepseek-ai/DeepGEMM https://github.com/deepseek-ai/DeepGEMM a competitive foundational library https://github.com/deepseek-ai/Engram https://github.com/deepseek-ai/Engram https://github.com/deepseek-ai/DeepSeek-V3 https://github.com/deepseek-ai/DeepSeek-V3 https://github.com/deepseek-ai/DeepSeek-R1 https://github.com/deepseek-ai/DeepSeek-R1 https://github.com/deepseek-ai/DeepSeek-OCR-2 https://github.com/deepseek-ai/DeepSeek-OCR-2 They have 33 repos and counting: https://github.com/orgs/deepseek-ai/repositories?type=all https://github.com/orgs/deepseek-ai/repositories?type=all And DeepSeek often has very cool new approaches to AI copied by the rest. Many others copied their tech. And some of those have 10x or 100x the GPU training budget and that's their moat to stay competitive. The models from Chinese Big Tech and some of the small ones are open weights only. (and allegedly benchmaxxed) (see https://xcancel.com/N8Programs/status/2044408755790508113 https://xcancel.com/N8Programs/status/2044408755790508113). Not the same.
- patshead 6mo agoDeepSeek's models are indeed open weight. Why do you feel that pointing this out would be considered slander?
- culi 6mo agoI think they were reading GP's comment as a correction. Like "not open-source, just open weight". I'm not sure if their reading was accurate but I enjoyed their high effort comment nonetheless
- alecco 6mo agoX is full of "open weights!" corrections as a dog whistle by the anti-China crowd. And they are right about models from the Chinese Big Tech, but completely wrong about DeepSeek.
- alecco 6mo ago>> Truly open source coming from China. > Open weight! They clearly were implying it's not open source.
- patshead 6mo agoCorrect. We have open-weight models from OpenAI, Facebook, Mistral, DeepSeek, Z.ai, MiniMax, and all sorts of other companies. Most of them have fantastic and open licensing terms. If we can't build the weights, then we don't have the source. I'm not entirely sure what an open-source model would even look like, but I am confident that these binary blobs that we are loading into llama.cpp and vllm aren't the equivalent of source code. We have absolutely no idea what sort of data went into them. This is fine. It isn't slanderous. It is what we have, and it is awesome. Just because it is awesome doesn't make it open source.
- kortilla 6mo agoIt’s not slander to say something true. These are open weights, not open source. They don’t provide the training data or the methodology requires to reproduce these weights. So you can’t see what facts are pruned out, what biases were applied, etc. Even more importantly, you can’t make a slightly improved version. This model is as open source as a windows XP installation ISO.
- alecco 6mo ago> These are open weights, not open source. Did you even read my comment?
- 0-_-0 6mo agoWeights are the source, training data is the compiler
- crazylogger 6mo agoTraining data == source code, training algorithm == compiler, model weights == compiled binary.
- 0-_-0 6mo agoTraining algorithm is the programmer, weights are the code that you run in an interpreter
- ngruhn 6mo agoisn't it more like the data is the source, the training process is the compiler, and the weights are the binary output.
- try-working 6mo agoif you want to understand why labs open source their models: http://try.works/why-chinese-ai-labs-went-open-and-will-remain-open http://try.works/why-chinese-ai-labs-went-open-and-will-rema...
- wraptile 6mo ago> Internet comments say that open sourcing is a national strategy, a loss maker subsidized by the government. On the contrary, it is a commercial strategy and the best strategy available in this industry. This sounds whole lot like potatoh potahto. I think the former argument is very much the correct one: China can undercut everyone and win, even at a loss. Happened with solar panels, steel, evs, sea food - it's a well tested strategy and it works really well despite the many flavors it comes in. That being said a job well done for the wrong reasons is still a job well done so we should very much welcome these contributions, and maybe it's good to upset western big tech a bit so it's remains competitive.
- try-working 6mo agoIt is not only that Chinese labs can undercut on price. It is that they must. They must give away their models for free by open sourcing them, and they must even give away free inference services for people to try them. That is the point of the post.
- FuckButtons 6mo agoThere is not ‘must’ here, they did not ‘have’ to undercut every other strategically and technologically important industry the rest of the world has, but they did as a point of national policy.
- try-working 6mo agoNo. Read what I wrote. I have spent a decade in the Chinese tech industry.
- 6mo ago
- b65e8bee43c2ed0 6mo agoAmerican companies want a scan of your asshole for the privilege of paying to access their models, and unapologetically admit to storing, analyzing, training on, and freely giving your data to any authorities if requested. Chinese ulteriority is hypothetical, American is blatant.
- elefanten 6mo agoIt’s not remotely hypothetical you’d have to be living under a rock to believe that. And the fusion with a one-party state government that doesn’t tolerate huge swathes of thoughtspace being freely discussed is completely streamlined, not mediated by any guardrails or accountability. This “no harm to me” meme about a foreign totalitarian government (with plenty of incentive to run influence ops on foreigners) hoovering your data is just so mind-bogglingly naive.
- OhioMan2943 6mo agoAnd you're saying Americans aren't banned from criticising their elites?
- tommica 6mo agoPretty sure you guys have a strong laws about free-speech, and criticizing elites is part of that. Though there are some groups that do not really want the 1st amendment to be a thing.
- spaceman_2020 6mo agoI don’t care about whatever “ulterior motives” they might have My country’s per capita income is $2500 a year. We can’t pay perpetual rent to OAI/Anthropic
- djyde 6mo agoSame
- Quothling 6mo agoIt's a little sad that tech now comes down to geopolitics, but if you're not in the USA then what is the difference? I'm Danish, would I rather give my data to China or to a country which recently threatened the kingdom I live in with military invasion? Ideally I'd give them to Mistral, but in reality we're probably going to continue building multi-model tools to make sure we share our data with everyone equally.
- jatora 6mo agoLol EU pats you on the head Its sad to see how you have regulated yourselves into a position where Mistral is your only claim.
- zerr 6mo agoDo they also open-source censoring filter rules? Like, you can't ask what happened at Tiananmen Square in 1989.
- harladsinsteden 6mo ago> I know if the potential ulterior motives. And you think the US tech giants don't have any ulterior motives?!
- FuckButtons 6mo agoI think their motives are pretty transparent, as are china’s, as ever, you have to pick the lesser of two evils.
- neonstatic 6mo agoHow are the "ulterior motives" of Chinese companies any worse than "ulterior motives" of US companies or European ones?