5 ms·
Whistleblower: Huawei cloned Qwen and DeepSeek models, claimed as own
- tengbretson 1y agoIn the LLM intellectual property paradigm, I think this registers as a solid "Who cares?" level offence.
- mathverse 1y ago[flagged]
- oblio 1y agoDidn't know Sam Altman was Chinese :-)
- brookst 1y agoThe point isn’t some moral outrage over IP, the point is a company may be falsely claiming to have expertise it does not have, which is meaningful to people who care about the market in general.
- tonyedgecombe 1y agoNobody who pays attention to Huawei will be surprised. They have a track record of this sort of behaviour going right back to their early days.
- npteljes 1y agoWhile true, these sorts of reports are the track records which we can base our assessments on.
- deleted 1y ago[deleted]
- didibus 1y agoYa, the models have stolen everyone's copyrighted intellectual property already. Not sure I have a lot of sympathy, in fact, the more the merrier, if we're going to brush off that they're all trained on copyrighted material, might as well make sure they end up a really cheap, competitive, low margin, accessible commodity.
- lambdasquirrel 1y agoEh... you should read the article. It sounds like a pretty big deal.
- didibus 1y agoI did read the article, appart for that it sounds like a terrible place to work, I'm not sure I see what's the big deal? No one knows how any of the models got made, their training data is kept secret, we don't know what it contains, and so on. I'm also pretty sure a few of the main models poached each others employees which just reimplemented the same training models with some twists. Most LLMs are also based on initial research papers where most of the discovery and innovation took place. And in the very end, it's all trained on data that very few people agreed or intended would be used for this purpose, and for which they all won't see a dime. So why not wrap and rewrap models and resell them, and let it all compete for who offers the cheapest plan or per-token cost?
- esskay 1y agoIt is very hard to have any sympathy, they stole stolen material from people known to not care they are stealing.
- some_random 1y agoClaiming to care deeply about IP theft in the more nebulous case of model training datasets then dismissing the extremely concrete case of outright theft seems pretty indefensible to me.
- perching_aix 1y agoPar for the course for emotional thinking, I'm not even surprised anymore.
- Arainach 1y agoEveryone has a finite amount of empathy, and I'm not going to waste any of mine on IP thieves complaining that someone stole their stolen IP from them.
- mensetmanusman 1y agoIt’s theft in the way taking a picture of nature that you had nothing to do with is theft.
- Arainach 1y agoThis line of argument was worn out and tired when 14 year olds on Napster were parroting it in 1999.
- some_random 1y agoI'm not asking you to cry or even stifle a laugh, the only think I'm criticizing is an uneven application of claimed values. Edit: Or an argument as to why this IP theft is fine while that used in training isn't. I'm sure some of that training data was CC-SA licensed for instance ;)
- pton_xd 1y ago> dismissing the extremely concrete case of outright theft seems pretty indefensible to me. Outright theft is a meaningless term here. The new rules are different. The AI space is built on "traditionally" bad faith actions. Misappropriation of IP by using pirated content and ignoring source code licenses. Borderline malicious website scraping. Recitation of data without attribution. Copying model code / artifacts / weights is just the next most convenient course of action. And really, who cares? The ethical operating standards of the industry have been established.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- gausswho 1y ago"Saturday was a working day by default, though occasionally we had afternoon tea or even crayfish." Unexpected poetry. Is there a reason why crayfish would be served in this context?
- tecleandor 1y agoI understood like "even as they made us work on Saturday, we sometimes had the luck of having some afternoon snack", and I guess crayfish might be popular there. Or maybe it's a mistranslation.
- sui762o 1y ago[dead]
- alwa 1y agoImmensely popular, delicious, and very beautiful on a plate or in a bowl, both whole/boiled/stir-fried and as snack packs of pre-peeled tails! See, e.g., https://mychinesehomekitchen.com/2022/06/24/chinese-style-spicy-crayfish/ https://mychinesehomekitchen.com/2022/06/24/chinese-style-sp... So yes, I read it the same way you do: “They made us work weekends, but at least they’d order us in some pizzas.” (…and if you’re in the US, you can have them air-freighted live to you, and a crawfish boil is an easy and darn festive thing to do in the summer. If you’re put off by the crustacean staring back at you, and you have access to a kitchen that operates in a Louisianan style, you might be able to find a “Cajun Popcorn” of the tails seasoned, battered, and fried. Or maybe one of the enormous number of “seafood boil” restaurants that have opened in the US in recent years.) (I feel like those establishments came on quickly, that I notice them mainly in spaces formerly occupied by American-Chinese restaurants, and that it’s felt like a nationwide phenomenon… I suspect there’s a story there for an enterprising young investigative nonfiction writer sort.)
- tecleandor 1y agoOh! That sounds tasty. I'm in EU, but I'm gonna take note of both. Thanks.
- bigmattystyles 1y agoOld maps (and perhaps new ones) used to add fake little alleys so a publisher could quickly spot publishers infringing on their IP rather than going out and actually mapping. I wonder if something similar is possible with LLMs.
- Tokumei-no-hito 1y agoi have come across this one for example https://github.com/sentient-agi/OML-1.0-Fingerprinting https://github.com/sentient-agi/OML-1.0-Fingerprinting > Welcome to OML 1.0: Fingerprinting. This repository houses the tooling for generating and embedding secret fingerprints into LLMs through fine-tuning to enable identification of LLM ownership and protection against unauthorized use.
- NitpickLawyer 1y agoWould be interesting to see if this kind of watermarking survives the frankenstein types of editing they are presumably doing. Per the linked account, they took a model, changed tokenizers, and added layers on top. They then presumably did some form of continued pre-training, and then post-training. It would have to be some very resistant watermarking to survive that. It's not as simple as making the model reply with "my tokens are my passport, verify me" when you ask them the weather in NonExistingCity... Interesting nonetheless.
- Tokumei-no-hito 1y agoi have never used it and have limited understand of fine tune models. i only remember see this a few weeks ago and your comment reminds me. i am curious too.
- varispeed 1y agoI often say an odd thing on public forum or make up a story and then see if LLM can bring it up. I started doing that once LLM provided me with a solution to a problem that was quite elegant, but was not implemented in the particular project. Turns out it learned it from GitHub issues post that described how particular problem could be tackled, but PR never actually got in.
- matt3210 1y agoThe question is who really made the original models?
- hereme888 1y ago[flagged]
- jambutters 1y agoI don't think anyone cares about that. OpenAI ripped off of the internet and books. Deepseek distilled some of openAI and pushed the field forward
- kkzz99 1y agoRemember that there was a Huawei Lab member that got fired for literally sabotaging training runs. Would not be surprised if that was him.
- yorwba 1y agoI think the case you're talking about is this one: https://arstechnica.com/tech-policy/2024/10/bytedance-intern-fired-for-planting-malicious-code-in-ai-models/ https://arstechnica.com/tech-policy/2024/10/bytedance-intern... where it was a ByteDance intern.
- typon 1y agoLLMs are all built on stolen data. There is no such thing as intellectual property in LLMs.
- mattnewton 1y agoThat’s not the point IMO; the point was this was being used to display capabilities to train models with Huawei software and hardware.
- mensetmanusman 1y ago/robots that read books in the library are stealing/
- JPLeRouzic 1y agoThat's a very human and very honest report. It presents the confusion there is in some big companies and how the pressure by the management favors dishonest teams. The writer left the company. I hope he is well; he is a fine person.
- dworks 1y agoYes. In fact, this report should be written in the context of other farewell letters to employers that have been published recently in China. There has recently been one, by a 15-year Alibaba veteran, who decried the decline of the company culture as a cause of its now lacking competitiveness and inability to launch new products. The issues in this report are really about: 1. Lies about Huawei's capabilities to the country (important national issue) 2. Lies to customers who paid to use Huawei models 3. A rigid, KPI-focused and unthinking organization where dishonest gaming of the performance review system not only works but seems to be the whole point and is tacitly approved (this and the reporters idealism and loss of faith is the main point of the report as I see it)
- yorwba 1y agoI think the reporter's motivations would've come across more clearly if you had posted a paragraph-by-paragraph translation instead of the current abridged version. (I assume Dilemma Works is your Substack.) Lots of details that add color to the story got lost.
- mystraline 1y ago[flagged]
- knowitnone 1y ago"Google cloned Linux kernel, claimed as own." Link?
- owebmaster 1y agoAndroid
- egypturnash 1y agoLLMs are apparently completely incompatible with copyright anyway, so if you can train them without paying a single dime to anyone whose work you ingest, then you should be able to clone them for free. What goes around comes around.
- mensetmanusman 1y agoThey are naïvely incompatible, but lawyers will find a way to make it not so.
- throwaway48476 1y agoChinese efficiency. The west is held back by archaic IP laws.
- option 1y agoDoesn't feel like a healthy culture, IF true. Also, apparently current DeepSeek lab members aren't allowed to travel to conferences. This is all maybe good for execution but absolutely not for innovation
- option 1y ago"Organization: We belong to the “Fourth Field Army” initiative. Under its structure, core language large models fall under the 4th brigade; Wang Yunhe’s small-model group is the 16th brigade." - Lol, what? So is this literally a part of CCP military?
- tedivm 1y agoI don't think so. The Fourth Field Army doesn't exist anymore (and hasn't since 1955). My guess is the company named their LLM initiative after this for historic reasons, and that these are more like internal project code names than anything else.
- neurostimulant 1y agoHuawei has a militarily-style culture. Even their new employee orientation is run like an army boot camp. https://archive.is/wvbca https://archive.is/wvbca
- jauntywundrkind 1y agoMeanwhile Apple legitimately built on Qwen2.5-Coder-7B, adding some of their own novel ideas. It mostly seems like custom training for their own code examples, but notably if you turn the temperature up, it can write multiple blocks of code out of order. https://9to5mac.com/2025/07/04/apple-just-released-a-weirdly-interesting-coding-language-model/ https://9to5mac.com/2025/07/04/apple-just-released-a-weirdly... https://news.ycombinator.com/item?id=44472062 https://news.ycombinator.com/item?id=44472062
- maxglute 1y agoWriter somewhat naive. His Ascend team couldn't get comparable performance (gen1 910A NPUs) initially vs (I assume) Nvidia because obviously. Management supported teams that pivot to cloned alternatives that used GPUs that can be immediately commercialized. Internal office politics make this happen. Ascend team works out kinks (this is huge confirmation), but feel (are) mistreated, i.e. biased bureaucracy, lack of recognition. Many burnout / leave to other Chinese AI companies. HW strategy/culture has been burning tier1 talent since forever. I remember in the 90s When HW and other domestic PRC telco started poaching from Nortel, Siemens, Lucent etc... the talent (most Chinese diaspora used to comfy western office culture) did not have a good time fitting into an actual Chinese company with Chinese culture (but got paid lots). Many burned out too... yet HW, a particularly extreme outlier of militant work culture, has become dominant.. LBH, both HW post sanctions, is a strategic company, overlapping with semi fabrication, domestic chips, and AI is cubing their strategic value. They can get away with doing anything under the current geopolitical environment to stay dominant. The worthwhile take away from this farewell letter is HW threw enough talent at Ascend that it kind of works now, and potentially can throw enough talent at it to be competitive with Nvidia. AKA how it has always operated, like massive wankers. The intuition from the author and most of us is... you need to reward employees right, cultivate proper workplace environment blah blah blah... but look at HW for the past 30 years. They pay a lot of smart people (including patriotic suckers) A LOT of money, throw them at problems until they break. And win.
- rjzzleep 1y agoThis doesn't seem right at all, given that DeepSeek reported massive performance increase because the Huawei team helped them port their LLM to HW infra. I was willing to put it in the "I don't know, maybe, let's see" category, but that comment specifically makes it read like a propaganda piece.
- maxglute 1y agoAscend 910A was Q1 2022. Qwen 2.5 / Deepseek v3 was Q1 2025. Implied timeline seems to be it took ~2.5 years developement to figure out how to use Ascend somewhat competitively and HW may rival Nvidia with proper support. Where proper support in authors opinion is building a team and treating team well. My guess is Huawei can keep burning talent at problem. The Huawei+Deepseek > Nvidia for inference claims is based on HW CloudMatrix 384 supernode using all of the optical interconnects and power to allegedly out perform Nvidia cluster. Bottleneck of Nvida cluster is switching. Bottleneck of PRC cluster is chips on old node size (slower or more power hungry). CM384 workaround is 5x more chips 910c chips and 4x more power consumption and connecting cluster with full optical to compete vs GB200 on pure cluster performance. Hard to say what actual economics of that solution is.