6 ms·
As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, the
by rao-v 24d ago
As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.
I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.
They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.
- alchemist1e9 24d agoquant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.
- TacticalCoder 24d ago> quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take. It's quite crazy that it's Deepseek's background/original purpose. We already had very advanced stuff from the world of HFT, but now a frontier family of models from a private company that used to be (still is?) in HFT is plain bonkers. Is more known about them and the HFT background?
- natrys 24d agoAccording to an old interview, apparently they were always interested in AI. But finance is just where they had their first success. > Many of High-Flyer's original team members worked on AI. Back then, we tried a lot of fields before getting our big break in finance, which is complex enough. AGI is probably one of the hardest things we can do next, so for us it was a question of how, not why. It's a very good interview: https://www.lesswrong.com/posts/kANyEjDDFWkhSKbcK/two-interviews-with-the-founder-of-deepseek https://www.lesswrong.com/posts/kANyEjDDFWkhSKbcK/two-interv... Incidentally, Wenfeng is kind of reverse Hassabis. There were some rumours that: > Hassabis quietly assembled a team of around 20 researchers to develop high-frequency trading algorithms, without Google's approval. When the parent company found out, the project was disbanded. https://timesofindia.indiatimes.com/technology/tech-news/when-google-killed-deepmind-ceo-demis-hassabis-secret-team-building-algorithms-on-/articleshow/129976958.cms https://timesofindia.indiatimes.com/technology/tech-news/whe...
- throwaway85825 24d agoIn this case the finance model was used to parse lengthy, verbose and inscrutable yet very impactful chinese government pr statements and do sentiment analysis.
- gpt5 24d ago[flagged]
- rao-v 24d agoumm what are you talking about? Basically this crowd (esp. folks like me who run medium models locally) like open stuff and can be a tiny bit unenthused about opaque mysteries handed down from on high. You'll see people delighted with Gemma releases and heck even IBM's Granite models (boring architecturally though they may be) every time they come out. Heck I was chuffed about gpt-oss-120b for weeks. @sama give us another already!
- deleted 24d ago[deleted]
- kcocoa 24d agoNot Chinese/American models. We are talking about open-weight (and their detailed tech report) and close-weight (with non-sense restrictions)
- markasoftware 24d agoOr maybe, the "hacker" philosophy that this site is named after, is strongly opposed to the philosophies that the American labs seem to be operating on? anyways, remember HN rules: "Please don't post insinuations about astroturfing, shilling, brigading, foreign agents, and the like. It degrades discussion and is usually mistaken. If you're worried about abuse, email hn@ycombinator.com and we'll look at the data."
- gpt5 24d agoIt has nothing to do with open vs closed or "hacker" philosphy. See this the announcement of the closed Seedance 2.5 - https://news.ycombinator.com/item?id=49138302 https://news.ycombinator.com/item?id=49138302 Direct quote from the second top comment: > Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them. Compare that with the launch of ChatGPT Image of yesterday.
- porridgeraisin 24d agoThis is adapted from Microsoft research's YOCO. It was known for a while(2024!). Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM. Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".
- NooneAtAll3 24d agowhy didn't Microsoft scale its own invention?
- CharlieDigital 24d agoPolitics and profits. Deepseek delivers 1 product; Microsoft delivers dozens (or hundreds depending on how you want to count it) across various domains.
- porridgeraisin 24d agoThere are thousands of such techniques across different parts of the system. In ML, there are way too many ideas, and lots of people knowingly and unknowingly restate the same ideas. It's a new field, so even common language is not there. For an extreme example, so many improvements are restatements of 1960 signal processing techniques - obviously very few ML people have done DSP beyond the undergrad course. It is also highly empirical and many parts of deep learning (heck, even non-NN ML) are not understood yet. Thus, the reality is that most of these ideas become polished only when its actually deployed and it has to work outside of a PoC. Since LLMs are a high capex product, only very few people actually make non-PoCs. Deepseek is in the business of low cost, fast inference. So they are the ones actually polishing these efficiency-ish ideas and combining many of them (this one, then engram which is based on multiple previous ideas including google brain's ngrammer) to make a coherent system. Openai and anthropic's systems will also involve a polished combination of multiple ideas for each of their systems - Luna is likely a combination of a few efficiency-ish ideas. Shame they won't publish though. As for microsoft, they don't really sell models, they sell azure. So there is no reason for them to do the high capex scale out of these types of bags of techniques. In a sense, it did benefit them, others developed the model and now many US customers can serve DS4.1 Flash on Azure datacenters. If it is not clear, I am not understating anything. Combining these rough ideas and making them work actually involves real novel ideas on top and is what is much more difficult than the academic results that were built upon. This also does not mean that the academic results are useless, they are what give us useful priors at all in what is a highly empirical field.
- ainch 24d agoIt was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).
- deleted 24d ago[deleted]
- ungovernableCat 24d agoIts CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it. Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is
- cbg0 24d agoWhile typical investors in their last round are subject to a five-year lock-up and will not have voting rights, China's National Artificial Intelligence Industry Investment Fund also put money into it, retaining both voting rights and freedom from the lock-up. Nothing really new if you're aware of how involved the CCP is with companies of strategic importance in China. https://www.reuters.com/world/asia-pacific/chinas-deepseek-closes-over-7-billion-funding-with-unusual-deal-structure-2026-06-16/ https://www.reuters.com/world/asia-pacific/chinas-deepseek-c...
- ungovernableCat 24d agoOh of course, you’re not getting into positions of power by not playing by the party’s rules. And if you get notions that you can tell THEM what to do you’ll be swiftly dealt with. The company is doing well and providing great PR so the party is content to not meddle too much I imagine. My comparison with American labs is more that I think they have to deal with bean counters, creditors, investors etc which can shuffle incentives and aims (and is a big reason why they dont do open weights anymore)
- nozzlegear 24d ago> My comparison with American labs is more that I think they have to deal with bean counters, creditors, investors etc which can shuffle incentives and aims (and is a big reason why they dont do open weights anymore) And don't forget kowtowing to Trump.
- swed420 24d ago> Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is To expand, he also has knowledge and hands-on experience in this and related fields. Just like the founder of Xerox PARC.
- deleted 24d ago[deleted]
- nater5000 24d ago>it really amazes me how fearless Deepseek are Reel it in a bit, man.
- jrflo 24d agoThe circle jerking of Chinese models on this site never ceases to amuse me.
- camel_Snake 24d agoI assume it has something to do with being on a website called "Hacker News" and said models are, for the most part, the only ones being published as open weights.
- gf000 24d agoAs opposed to the western ones'?
- Freedom2 24d agoIf it were a YC company I'd understand, but for anything foreign, I'm always suspicious.
- nozzlegear 24d agoOpen is better than closed, simple as. Nothing to do with China versus America, I'll always stan the open models and cast my aspersions on the closed ones.
- orangeboats 24d ago>The circle jerking of Chinese models on this site never ceases to amuse me. As opposed to "ask the model about Tiananmen" which seems to be the site's favorite pastime about Chinese models. ;) -- Sarcasm aside, I don't think people are happy about _Chinese_ models making advances. They are happy about _open_ models making advances. It's just coincidental that China is the one making them. If some American lab were to develop a SOTA open model most people here will be equally excited. Although besides GPT-OSS-120B the American labs have been disappointing in this regard.
- 24d ago
- asdfman123 24d agoIt's because you have to be fearless as the challenger. As the established player you have more to lose.