6 ms·
CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due
by leetharris 2y ago
CEO of Scale said Deepseek is lying and actually has a 50k GPU cluster. He said they lied in the paper because technically they aren't supposed to have them due to export laws.
I feel like this is very likely. They obvious did some great breakthroughs, but I doubt they were able to train on so much less hardware.
- pdntspa 2y agoI would think the CEO of an American AI company has every reason to neg and downplay foreign competition... And since it's a businessperson they're going to make it sound as cute and innocuous as possible
- stale2002 2y agoOr, more likely, there wasn't a magic innovation that nobody else thought of, that reduced costs by orders of magnitude. When deciding between mostly like scenarios, it is more likely that the company lied than they found some industry changing magic innovation.
- pjfin123 2y agoIt's hard to tell if they're telling the truth about the number of GPUs they have. They open sourced the model and the inference is much more efficient than the best American models so it's not implausible that the training was also much more efficient.
- leetharris 2y agoIf we're going to play that card, couldn't we also use the "Chinese CEO has every reason to lie and say they did something 100x more efficient than the Americans" card? I'm not even saying they did it maliciously, but maybe just to avoid scrutiny on GPUs they aren't technically supposed to have? I'm thinking out loud, not accusing anyone of anything.
- mrbungie 2y agoThen the question becomes, who sold the GPUs to them? They are supposedly scarse and every player in the field is trying to get ahold as many as they can, before anyone else in fact. Something makes little sense in the accusations here.
- leetharris 2y agoI think there's likely lots of potential culprits. If the race is to make a machine god, states will pay countless billions for an advantage. Money won't mean anything once you enslave the machine god. https://wccftech.com/nvidia-asks-super-micro-computer-smci-to-investigate-how-its-chips-were-illicitly-procured-by-china-including-the-issue-of-forged-server-serial-numbers/ https://wccftech.com/nvidia-asks-super-micro-computer-smci-t...
- mrbungie 2y agoWe will have to wait to get some info on that probe. I know SMCI is not the nicest player and there is no doubt GPUs are being smuggled, but that quantity (50k GPUs) would be not that easy to smuggle and sell to a single actor without raising suspicion.
- rajhlinux 2y agoFacts, them Chinese VCs will throw money to win.
- rajhlinux 2y agoMan, they say China is the most populated country in the world, I’m sure they got loopholes to grab a few thousands H100s. They probably also trained the “copied” models by outsourcing it. But who cares, it’s free and it works great.
- rajhlinux 2y agoBro, did you use Deepseek? That shyt is better than ChatGPT. No cards being thrown here.
- latchkey 2y agoThanks to SMCI that let them out... https://wccftech.com/nvidia-asks-super-micro-computer-smci-to-investigate-how-its-chips-were-illicitly-procured-by-china-including-the-issue-of-forged-server-serial-numbers/ https://wccftech.com/nvidia-asks-super-micro-computer-smci-t... Chinese guy in a warehouse full of SMCI servers bragging about how he has them... https://www.youtube.com/watch?v=27zlUSqpVn8 https://www.youtube.com/watch?v=27zlUSqpVn8
- Leary 2y agoAlexandr Wang did not even say they lied in the paper. Here's the interview: https://www.youtube.com/watch?v=x9Ekl9Izd38 https://www.youtube.com/watch?v=x9Ekl9Izd38. "My understanding is that is that Deepseek has about 50000 a100s, which they can't talk about obviously, because it is against the export controls that the United States has put in place. And I think it is true that, you know, I think they have more chips than other people expect..." Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people.
- leetharris 2y ago> Plus, how exactly did Deepseek lie. The model size, data size are all known. Calculating the number of FLOPS is an exercise in arithmetics, which is perhaps the secret Deepseek has because it seemingly eludes people. Model parameter count and training set token count are fixed. But other things such as epochs are not. In the same amount of time, you could have 1 epoch or 100 epochs depending on how many GPUs you have. Also, what if their claim on GPU count is accurate, but they are using better GPUs they aren't supposed to have? For example, they claim 1,000 GPUs for 1 month total. They claim to have H800s, but what if they are using illegal H100s/H200s, B100s, etc? The GPU count could be correct, but their total compute is substantially higher. It's clearly an incredible model, they absolutely cooked, and I love it. No complaints here. But the likelihood that there are some fudged numbers is not 0%. And I don't even blame them, they are likely forced into this by US exports laws and such.
- kd913 2y agoIt should be trivially easy to reproduce the results no? Just need to wait for one of the giant companies with many times the GPUs to reproduce the results. I don't expect a #180 AUM hedgefund to have as many GPUs than meta, msft or Google.
- sudosysgen 2y agoAUM isn't a good proxy for quantitative hedge fund performance, many strategies are quite profitable and don't scale with AUM. For what it's worth, they seemed to have some excellent returns for many years for any market, let alone the difficult Chinese markets.
- matthest 2y agoI've also read that Deepseek has released the research paper and that anyone can replicate what they did. I feel like if that were true, it would mean they're not lying.
- aprilthird2021 2y agoYou can't replicate it exactly because you don't know their dataset or what exactly several of their proprietary optimizations were
- woadwarrior01 2y agoCEO of a human based data labelling services company feels threatened by a rival company that claims to have trained a frontier class model with an almost entirely RL based approach, with a small cold start dataset (a few thousand samples). It's in the paper. If their approach is replicated by other labs, Scale AI's business will drastically shrink or even disappear. Under such dire circumstances, lying isn't entirely out of character for a corporate CEO.
- deleted 2y ago[deleted]
- leetharris 2y agoCould be true. Deepseek obviously trained on OpenAI outputs, which were originally RLHF'd. It may seem that we've got all the human feedback necessary to move forward and now we can infinitely distil + generate new synthetic data from higher parameter models.
- blackeyeblitzar 2y ago> Deepseek obviously trained on OpenAI outputs I’ve seen this claim but I don’t know how it could work. Is it really possible to train a new foundational model using just the outputs (not even weights) of another model? Is there any research describing that process? Maybe that explains the low (claimed) costs.
- a1j9o94 2y agoProbably not the whole model, but the first step was "fine tuning" the base model on ~800 chain of thought examples. Those were probably from OpenAI models. Then they used reinforcement learning to expand the reasoning capabilities.
- mkl 2y ago800k. They say they came from earlier versions of their own models, with a lot of bad examples rejected. They don't seem to say which models they got the "thousands of cold-start" examples from earlier in the process though.
- echelon 2y agoI haven't had time to follow this thread, but it looks like some people are starting to experimentally replicate DeepSeek on extremely limited H100 training: > You can RL post-train your small LLM (on simple tasks) with only 10 hours of H100s. https://www.reddit.com/r/singularity/comments/1i99ebp/well_seems_like_the_cat_is_out_of_the_bag/ https://www.reddit.com/r/singularity/comments/1i99ebp/well_s... Forgive me if this is inaccurate. I'm rushing around too much this afternoon to dive in.
- weinzierl 2y agoJust to check my math: They claim something like 2.7 million H800 hours which would be less than 4000 GPU units for one month. In money something around 100 million USD give or take a few tens of millions.
- pama 2y agoIf you rented the hardware at $2/GPU/hour, you need $5.76M for 4k GPU for a month. Owning is typically cheaper than renting, assuming you use the hardware yearlong for other projects as well.
- buyucu 2y agoWhy would Deepseek lie? They are in China, American export laws can't touch them.
- echoangle 2y agoMaking it obvious that they managed to circumvent sanctions isn’t going to help them. It will turn public sentiment in the west even more against them and will motivate politicians to make the enforcement stricter and prevent GPU exports.
- cue3 2y agoI don't think sentiment in the west is turning against the Chinese, beyond well, lets say white nationalists and other ignorant folk. Americans and Chinese people are very much alike and both are very curious about each others way of life. I think we should work together with them. note: I'm not Chinese, but AGI should be and is a world wide space race.
- siltcakes 2y agoThe CEO of Scale is one of the very last people I would trust to provide this information.
- eunos 2y agoAlexandr only parroted what Dylan Patel said on Twitter. To this day, no one know how this number come up.
- deleted 2y ago[deleted]
- rajhlinux 2y agoDeepseek is indeed better than Mistral and ChatGPT. It has tad more common sense. There is no way they did this on the “cheap”. I’m sure they use loads of Nvidia GPUs, unless they are using custom made hardware acceleration (that would be cool and easy to do). As OP said, they are lying because of export laws, they aren’t allowed to play with Nvidia GPUs. However, I support DeepSeek projects, I’m here in the US able to benefit from it. So hopefully they should headquarter in the States if they want US chip sanctions lift off since the company is Chinese based. But as of now, deepseek takes the lead in LLMs, my goto LLM. Sam Altman should be worried, seriously, Deepseek is legit better than ChatGPT latest models.
- riceharvester 2y agoR1 is double the size of o1. By that logic, shouldn’t o1 have been even cheaper to train?
- wortley 2y agoOnly the DeepSeek V3 paper mentions compute infrastructure, the R1 paper omits this information, so no one actually knows. Have people not actually read the R1 paper?