30 ms·
Deepseek: The quiet giant leading China’s AI race
- yellow_lead 2y ago> Liang Wenfeng: We believe that as the economy develops, China should gradually become a contributor instead of freeriding. In the past 30+ years of the IT wave, we basically didn’t participate in real technological innovation. We’re used to Moore’s Law falling out of the sky, lying at home waiting 18 months for better hardware and software to emerge. That’s how the Scaling Law is being treated.
- culi 2y agoIn the past few years, Chinese publications on AI research have surpassed English-language ones. Deepseek itself is open-sourced
- rubymamis 2y agoOpen source or open weights?
- cscurmudgeon 2y agoTotal number of publications or even citations is not a good metric to measure success in any competitive field.
- verygoodnotbad 2y agoThat sounds true to me. It seems like the CEO is pretty focused on truth. So maybe he also understands that the US has good reason to tariff and restrict Chinese investment. It is not only for the benefit of the US people, but of the world and the Chinese people. It is not out of emotional fear, but morality and responsibility, which are obviously trans-cultural.
- suraci 2y agogalaxy brain
- mentalgear 2y agoImpressive to think about how DeepSeek achieved: ~ Parity with o1 and Claude with > 10x less resources. Better algorithms and approaches are what's needed for the next step of ML.
- NitpickLawyer 2y agoWhile impressive, the deepseek models aren't really "on par" with either oAI or Anthropic offerings, right now. The models seem to be a bit overfitted in the post-training step. They are very "stubborn" models, and usually handle tasks well if they can handle them, but steering them is quite difficult. As a result, they score very well on various benchmarks, but often times perform slightly worse in real-life scenarios.
- espadrine 2y agoThe blind test at lmarena.ai does give it a higher Elo than GPT-4o (API), Claude, and Gemini 1.5 Pro. It seems that people do enter real-life scenarios in the arena.
- victorbjorklund 2y agoI found deepseek very useful at coding with Aider. On par with claude.
- rahimnathwani 2y agoThey are very "stubborn" models Have you found this to be the case even when using the recommended temperature settings (ranging from 0 for math, to 1.5 for creative tasks)?
- NitpickLawyer 2y agoI use 0.05 for math, just did a 5k problem set, trying to fine-tune a smaller model with the outputs. It has some very interesting training, borrowed from r1 per the tech report, where it does the o1/qwq "thinking steps", but a bit shorter. It solves ~80% of the problems in 4k context, while qwq would go on for 8k-16k. It's very good at what it does. But as soon as I need it to do something other than solve a problem - say rewrite the problem in simpler terms, or given a problem + solution provide hints, or rewrite the solution with these <tags>, etc. it kinda stops working. Often times it still goes ahead and solves the problem. That's why I'm saying it's stubborn. If a task looks like a task that it can handle very well, it's really hard to make it perform that other, similar but not quite the same task. In a similar vein - https://github.com/cpldcpu/MisguidedAttention/tree/main/eval#readme https://github.com/cpldcpu/MisguidedAttention/tree/main/eval...
- deleted 2y ago[deleted]
- lomkju 2y agoI feel the GPU restrictions created an environment for Chinese Devs to be more innovative and do more with less. Kudos to the deepseek team!
- wodderam 2y agoKai-Fu Lee describes the culture so well in AI Superpowers. The roots are well before GPU restriction. Absolute cut throat competition. Imagine Sam Altman throwing a chair out a window in a meeting lol. The message of AI Superpowers is that China will lag the US at first but once things stabilize this will happen because China has a lot more engineers and a lot more data. Anyone who hasn't read AI Superpowers should really make it a point to read it in 2025. It is an incredible book.
- Etheryte 2y agoI don't know, I've been hearing the story that China is about to upend the US as the leading global superpower ever since I was a kid. There's always a new vogue and novel twist put on the rationale and how it's gonna happen, but so far it's like fusion, always a few years away.
- elashri 2y agoI think you mean nuclear fusion.
- Etheryte 2y agoOf course, thanks.
- tossandthrow 2y agoWhat makes you think it has not happened? There has not been an event to establish who the current super power is in new time.
- futureshock 2y ago
- friend_Fernando 2y ago[flagged]
- kjellsbells 2y agoIf you tell the world that eggs are awesome while denying other countries access to eggs, they discover ways to use less eggs and eventually realize they don't need eggs at all. Then you are stuck making Dennys breakfasts while the rest of the world is on to fine dining. China has incredibly strong incentives to do the pure research needed to break the current GPU-or-else lock. I hope, for science' sake, we dont end up gunning down each others mathematicians on the streets of Vienna like certain nuclear physicists seem to go.
- djaouen 2y ago> If you tell the world that eggs are awesome while denying other countries access to eggs, they discover ways to use less eggs You are confusing cause with effect. What actually happened: Nixon opened up US trade with China and, ever since, China has been stealing trade secrets to undermine and overthrow American interests. Limiting their access to eggs was literally us trying to prevent them from stealing all our shit!
- quantum_state 2y agoIt seems to me that we forgot about the “stealing” of the “shit” from Europe and other places in the early days …
- djaouen 2y agoProtip: Some of us were not involved in the desecration caused by the East India Tea Company. Just because we look British means we should suffer like them, too?
- carom 2y agoThey are referring to the fact that the US ignored European IP in its early days and relating that to what China is doing to the US now.
- 2y ago
- nsoonhui 2y agoI find that the gushing around deepseek is fascinating to watch. To me there are a few structural and fundamental reasons why deepseek can never outperform other models by a wide margin. On par maybe--as we reach the diminishing returns with our investment in the models, but not win by a wide margin. 1. The US trade war with china which will place deepseek compute availability at disadvantages, eventually, if we ever get to that. 2. China censorship which limits the deepseek data ingestion and output, to some degree. 3. Most importantly, deepseek is open source, which means that the other models are free to copy whatever secret source it has, eg: Whatever architecture that purportedly use less compute can easily be copied. I've been using Gemini, chatgpt, deepseek and Claudie on regular basis. Deepseek is neither better or worse than others. But this says more about my own limited usage of LLM rather than the usefulness of the models. I want to know exactly what makes everyone thinks that deepseek totally owns the LLM space? Do I miss anything? PS: I am a Malaysian Chinese, so I am certainly not "a westerner who is jealous and fearful of the rise of China"
- logicchains 2y ago>I want to know exactly what makes everyone thinks that deepseek totally owns the LLM space? It achieved competitive performance to the competition at literally 10x less cost of production (training). That's an incredible achievement in any industry, especially given they have such a small team relative to competitors. Their API is 20-50x cheaper than the competitors, and not because they're burning cash by charging less than costs, but rather because their architecture is just that much more efficient. They already achieved the above in spite of sanctions limiting their availability to top-tier GPUs, and the gap between Chinese domestic GPUs and NVidia is getting smaller and smaller, so in future the GPU disadvantage will be less and less.
- nsoonhui 2y agoBut like I said, deepseek is open source so why can't the competitors copy whatever source that makes the cost of production 10x cheaper ?
- 2y ago
- suraci 2y agoI'm wondering what impact this will have on NVDA
- wiradikusuma 2y agoI hope the competition among AI companies will continue to be healthy. Meaning they will keep sharing their techniques and papers, and we, as a whole, will be better off.
- LittleTimothy 2y agoI'm getting so interested in the meta dynamics of this. The ability of the Chinese company to just openly state "we're working on this because it's interesting" rather than the US version "We want to wrap the world in puppies and hugs and we love you all and it's just a really embarrassing mistake I ended up buying myself a Koenigsegg and fired all the scientists from my non-profit board". To apply the same scepticism to the Chinese CEO - you can't threaten the monopoly of the Communist party so you have to pretend you're less capable than you are. I don't think there's any doubt that China can produce some level of tech innovation, I do wonder if it can be sustained and exploited since we saw the damage that went on with Alibaba. Although maybe that's looking like a more reasonable approach when you see the danger of the opposite happening in the US.
- bru3s 2y ago[flagged]
- dumbmrblah 2y agoPart of the reason their API is so cheap because they explicitly state they are going to train on your API data. Open AI and Claude say they won’t if you use their API (if you use ChatGPT that’s a different story). There are no free lunches.
- eldenring 2y agoThis comment is misleading. There is a "free lunch" here in the sense that serving this model is far cheaper than worse, open source models at scale. Yes they probably are more willing to go down in price due to this, but the architecture is open, and they are charging similarly to a 30B-50B dense model, which is about how many active params deepseek-v3 has.
- sanjams 2y agoSo then OP is correct? Your comment confirms the same sentiment about the tradeoff API users make: cheaper inference means you pay with your data. Sure Deepseek may publish their weights so you dont have to use the API, but the point still stands for the API.
- eldenring 2y agoIts a matter of degree. If 90% of the cost savings are from a new, smarter architecture, it doesn't make sense to point to the API terms as the reason for it being so cheap.
- cynicalsecurity 2y agoChina doesn't limit their AI research with so called safety and other concerns, but we do. Who is going to win? Somehow I don't think this is going to be us.
- chimen 2y agoWhat is the "so called safety" that we do?
- Alifatisk 2y agoFor example, ask Claude to help you out with a question from jackbox.tv and it will refuse because it's not family friendly.
- tokioyoyo 2y agoThey do. I swear this entire thread is just full of two extremes of misinformations from both sides.
- emporas 2y agoNot personally surprised that a MoE model performs so well. I used Mixtral a lot for coding Rust, and it had qualities no other model had except GPT 3.5 and later Claude Sonet. The funny thing is Mixtral was based on Llama 2 which was not trained on code that much. DeepSeek v3: 671B parameters on total, and 37B activated sounds very good even though impossible to run locally. Question if some people happen to know: For each query it activates just that many of parameters, 37B, and no more?
- coolspot 2y agoIt activates only 37B per query, but you don’t know which ones ahead of time, so you gotta store all 671B in (V)RAM.
- cma 2y agoBut you don't need cluster networking or nvlink so much like with splitting out llama 405B. You could even split them out with friends over internet levels of bandwidth.
- int_19h 2y agoMistral LMs are not LLaMA derivatives.
- wolfgangK 2y agoDeepSeek v3 can run on CPU & RAM : https://www.reddit.com/r/LocalLLaMA/comments/1hqidbs/deepseek_v3_running_on_llamacpp_wishes_you_a/ https://www.reddit.com/r/LocalLLaMA/comments/1hqidbs/deepsee... Epyc Gen4 and 12 memory channels of DDR5 @4800 should give you 7 to 9 t/s.
- orbital-decay 2y agoThis reminds me of PixArt-α. It's a diffusion model for image generation, that demonstrated that it's possible to train a SotA model on a ridiculously tiny budget ($28k).
- inSenCite 2y ago"Before Deepseek, CEO Liang Wenfeng’s main venture was High-Flyer (幻方), a top 4 Chinese quantitative hedge fund last valued at $8 billion" Seems wild that a top 4 quant hedge fund is only $8B?
- timtom123 2y agoSo much spam around this model. LocalLLaMA is stuffed with spam posts and even hacker news is getting spammed. Who has actually ran this model and verified performance? Does anyone know of a decent review from a trustworthy source?
- x_may 2y agoThe LMSYS leaderboards are crowdsourced and would be hard to fake, it showing a pretty strong performance in terms of human preference.
- paxys 2y agoCrowdsourced data is the easiest to fake unless you can somehow ensure that you have a completely unbiased population (which is impossible). There's a reason why certain models do so well on upvote-based leaderboards but rank nowhere on objective tests.
- CGamesPlay 2y agoWhich ones? I think fine-tunes are where I see most of this (I'll just call it) "model spam", but the base models don't seem to exhibit this behavior. I do see some models perform way below the curve compared to their benchmark performance, though (Phi family being the most famous).
- feverzsj 2y agoI've tried it. It's average at best. Nothing comparable to ChatGPT.
- starfezzy 2y agoWhere’s the spam? I scrolled dozens of posts without seeing a single mention of this—the biggest (certainly the most interesting) LLM news recently. When something big happens with Claude or ChatGPT there are more posts, but nobody calls that “spam”. Anyways, if you were actually following locallama (a subreddit about running LLMs locally, where this is by far the biggest and most relevant news topic currently) you’d have seen this post https://www.reddit.com/r/LocalLLaMA/s/Yay5njt963 https://www.reddit.com/r/LocalLLaMA/s/Yay5njt963 where a guy is working on running deepseek on llamacpp and demonstrates ~8tk/s using a cpu.
- waldrews 2y agoTo this day, asking Deepseek "what model are you" typically gives the answer "I'm an AI language model called ChatGPT, created by OpenAI. Specifically, I'm based on the GPT-4 architecture, which is designed to understand and generate human-like text based on the input I receive. My training data includes a wide range of information up until October 2023, and I can assist with answering questions, generating text, and much more. How can I help you today?" this tells us something about using synthetic data to bootstrap new model. All those clauses in the terms of service about not using the model to develop competing UI? Yeah, good luck with that.
- waldrews 2y agoAnd you can ask it if it's sure, and it'll consistently double down on insisting it's ChatGPT. Ask it what country it's developed in, and it'll say US; ask it if it's sure it's not China, and it'll be sure.
- amelius 2y agoI'm sure OpenAI breaches copyright just as well. They are just a little bit better at hiding it.
- waldrews 2y agoIt also tells us the genie is out of the bottle not just in the form of open weights being widely available, but in the form of the text corpuses coming from the existing model. The claimed low cost of Deepseek's training is partly enabled by the availability of all that synthetic data created by the first generation models trained and developed at much higher cost. When the Soviets got hold of the nuke plans, they greatly reduced their development costs by primarily by not having to redo all the experiments that led to dead ends. What's amazing is that time it's different; nobody needs OpenAI's secret sauce anymore, just enough data - some of it happily supplied by ChatGPT itself, and they can experiment with different architectures and either get tolerable results with an architecture already in textbooks, or greatly improve efficiency by innovating.
- paxys 2y ago
- fallmonkey 2y agoStrangely, deepseek has been always a prominent name in open source LLM community since last year, with their repos and papers - https://github.com/deepseek-ai https://github.com/deepseek-ai. Nothing of it is really quiet except that they probably burn 1% of marketing money compared to other China LLM players.
- rickandmortyy 2y ago[dead]
- exe34 2y agoI have a question for the floor - given the worsening situation with technological unemployment, and the structural inability of capitalism to cope with it (who will buy the products when nobody has a job?), is it possible that China will be able to pivot to UBI and push on ahead? they have enormous control over the population and economy, so they might be able to change direction faster than the West?
- sinuhe69 2y agoEven in their 'classic years', communist states could never provide UBI to their citizens. Sure, the state employed most people outside the agricultural sector, but even so, they have to grade salaries and food ratios. As far as I know, at no point in the history of communism have they been able to provide a basic income without work, let alone a UBI. Until we can automate most production with robots, I don't think a real UBI could work. Ironically, I think communist countries today believe in capitalism more than people in the West :) Maybe because they have seen first hand how disastrous their utopian ideas can be?
- exe34 2y agomy question is what happens when nobody can afford the stuff that's being produced anymore? I understand that communism hasn't worked so far, but if you don't feed the hordes, you can get Luigied.
- sinuhe69 2y agoThe "technological unemployment" you mentioned can only happen if robots/AI replace most of the jobs on the market. But then it means that we have already automated most production and therefore UBI could be an option (taxing the machine). Otherwise, there will always be a demand for labor. It's not just production that the world economy is about. I see no limit to growth in services such as health, education and entertainment once basic provision is secured.
- exe34 2y ago
- dgfitz 2y agoLeading from the rear, with subterfuge and theft. Keep going, China, you’re an inspiration to us all.
- culi 2y agoThere's no reason to think Deepseek is engaged in any such practices. It's also open source unlike many of the western counterparts
- dgfitz 2y agoYou posit that individual freedoms exist inside China? I posit they do not.
- 8note 2y agowhen do i get to use my western freedoms to inspect OpenAIs datasets?
- rickandmortyy 2y ago[dead]
- sroussey 2y agoI’m surprised there is no word of combining old school symbolic AI with the new ML derived versions we enjoy today.
- kopirgan 2y agoInteresting (mis)use of the word catfish Not how we normally understand
- ovelv 2y ago[dead]
- tw1984 2y agoit is good news for all software devs and AI researchers, we are taking the fruits of AI back from silicon mongers!
- blackoil 2y agoIf is funny how a site that otherwise stays away from politics turns into reddit as soon as China is mentioned.
- throwaway290 2y agoMaybe because it is a country using technology to attack US. It caused deaths of US citizens. And this was going for 10+ years. I have the opposite question, why that is not brought up every time China is mentioned. https://www.naccho.org/blog/articles/cyber-attack-on-u-s-hospital-group-highlights-vulnerability-of-critical-infrastructure https://www.naccho.org/blog/articles/cyber-attack-on-u-s-hos... https://www.theregister.com/2024/12/30/att_verizon_confirm_salt_typhoon_breach/ https://www.theregister.com/2024/12/30/att_verizon_confirm_s... https://www.politico.com/news/2022/12/28/cyberattacks-u-s-hospitals-00075638 https://www.politico.com/news/2022/12/28/cyberattacks-u-s-ho... Yes these days more of it is Russia and DPRK (the peace loving prosperous country according to ByteDance's AI) but hmm let's see where they would get the tech from if they are banned from it otherwise
- suraci 2y agoare we that good already?
- throwaway290 2y agowho are "we"? good at what exactly?
- Giorgi 2y agoAhh, yes another Chinese ChatGPT killer that is crappy.
- codedokode 2y agoWest is making a mistake again. They should not allow export of GPU and publish information on ML. Instead it would be wiser to become a monopoly and sell only AI services. If other countries learn how to do AI, nobody will need expensive Western services anymore. Given that there are expectations that AI will be able to replace humans and increase manufacturing productivity, it should be well guarded unless you want your foreign competitors to increase the productivity too. The wise strategy is to sell goods or services but never to sell tools that can be used to produce them, like industrial machines and robots.
- l33tc0de 2y ago[flagged]
- robblbobbl 2y agoGood there is already an EU competitor available
- anshumankmr 2y agosadly not much from India on this front save for maybe Sarvam AI
- jimbobthemighty 2y agoTry asking about the Tiananmen Square massacre or why people compare Xi to Paddington Bear or even the failings of Xi... but it will happily criticise Trump.
- Alifatisk 2y agoTheir web interface refuses it, but their api still answers it happily!
- murtio 2y agoI'm starting to believe that these articles are commissioned. I asked their public model questions related to branding and marketing, instructing it to come up with a brand identity based off the apps functionalities. It kept talking to itself for more than 5 minutes in Chinese! Then finished up with a very bad answer!
- wsm123456 2y ago[dead]