8 ms·
The fact that AI models can be so easily distilled and replicated is such a stroke of luck. 10 or 15 years ago if one had asked me to envision a future where a
by eigenspace 1mo ago
The fact that AI models can be so easily distilled and replicated is such a stroke of luck.
10 or 15 years ago if one had asked me to envision a future where a private company invents artificial intelligence, I'd have thought for sure they'd have a massive moat, be very difficult to catch, and it would create an almost instant monopoly.
Rather, it seems that selling intelligence might end up as a race to the bottom.
Who woulda thought that just having access to enough textual inputs and outputs and a vaugely similar transformer architecture would be enough to copy-cat rather useful intelligence.
- td-andrew 1mo agoIt reminds me of the seo antics out there. The search results page is the engine, much like how distilling is the "intelligence" for your chinese room machine
- visarga 1mo agoFunny you mention Chinese Room and LLMs in the same response, I would say LLMs proved Searle wrong, agents now make cutting edge discoveries and meaningful problem solving. They not lookup tables though and you need to pay for inference, so the intuition of syntax doing the work of semantics without understanding was wrong.
- td-andrew 1mo agoWhether it's useful or not isn't the point. Chinese Room was a consideration as to what counts as knowledge. Chinese Room / LLM is not knowledge, it's pattern matching. Having a dude doing translations has always been useful even if they didn't understand the subject matter. Happened all the time pre LLMs
- deleted 1mo ago[deleted]
- make3 1mo agowell, a stroke of luck until the whole US stock market crashes & everyone's retirement funds get cut 40% I guess when people internalize this. it will have to happen sooner or later though I suppose
- eigenspace 1mo agoI'd take a market crash over a monopoly in the hands of a ghoul like Altman. The economy he and his ilk want to build is infinitely worse.
- fidotron 1mo agoIn truth it crashes either way.
- make3 1mo agointerestingly also, open weight models are also more effectively run in the cloud, so it creates a weird scenario where the frontier labs crash but the compute providers, not as much
- eigenspace 1mo agoI wouldn't be so sure about that. The popping of a bubble is usually just as irrational as its rise. If investors start fleeing from senseless businesses in the AI sector, that does not mean that sensible businesses will be spared. These things follow herd mentality, and the primary drivers of the herd are greed and fear, not fundamentals or business logic.
- cyanydeez 1mo agoAmerica is pretty close to rhyming with nazi germany circa 1929.
- aliasxneo 1mo agoOk, I'll bite. What's your rationale?
- Rover222 1mo agoI think the only moat in the future will be the scale of hardware deployment. If one company is able to deploy an order of magnitude more silicon, they'll have a firm grip on a SOTA model and massive inference usage. China or SpaceX seem like the 2 likely candidates in 5 years, but who knows.
- GMoromisato 1mo ago"Who knows" is the right answer, I think. If (a) demand for AI continues to increase, and (b) SpaceX can get to ~$100/kg to orbit, then they will have a ridiculously deep moat. Probably more like 10 years, though. But as you said, who knows.
- Rover222 1mo agoYeah, very hard to predict the future at this point. But the Starship + Terrafab combo will be this type of order-of-magnitude-moat IF it works out. Big if. If it doesn't work out, I think China's exponential terrestrial energy deployment will eventually give them the lead, IF they can get enough chips. Another big if.
- t0mas88 1mo agoThey will have moat in the satellite launching business, which is not useful in the AI datacenter market. You can put AI chips in datacenters in the desert for far less than $100/kg. With lots of solar power available, the option to easily access your hardware and far less radiation issues. The datacenter in space story really only exists to make it possible for Musk to sell X to SpaceX and make more money from the IPO. That's all. There is no engineering reason.
- noja 1mo agoAny cooling issues to be resolved?
- GMoromisato 1mo agoNo. Cooling is probably the easiest problem to solve, easier than power. And in both cases, the problem is solved by mass to orbit. All you need for cooling is a big f-ing radiator. Solar panels are chips, and not trivial to manufacture. But a radiator is just a hunk of metal with some pipes. That's why the cost of mass to orbit is the most important thing. You can solve almost any space problem by just throwing more mass at it.
- ElijahLynn 1mo agoAltman specifically has said in an interview that I listened to once that he envisions AI being as cheap as electricity.
- pseudony 1mo agoHe also wanted to do a non-profit. He even raised money on that premise. He is a pathological liar, so is Dario. Don’t rely on the benevolence or truthfulness of these people. They will say whatever is beneficial to say in the moment.
- pseudony 1mo agoTelling. Downvotes but actually no arguments. Fitting, because there are none. Referring to a baseless prediction by Sam Altman that AI will become like electricity without any push-back? Who really thinks Sam is working toward that future? He already worked to undo every early promise made (non-profit, open source models, strong governing board, strong ethics/alignment/security focus). He's flip-flopped on other things like first characterising Trump "an unprecedented threat to America", then contributing 1M USD to Trump's inaugural fund far exceeding his earlier political contributions. Lately OpenAI, under his supervision, has also been working with Anthropic to lobby regulators in Washington for restrictions on open weights models - why so if not to undermine a free market in favour of an oligopoly? Beyond that, you have the simple fact that most of his personal wealth and very probably the fate of OpenAI hinges on AI inference NOT becoming an interchangeable commodity. I mean.. Honestly. The naivete is downright astounding.
- deleted 1mo ago[deleted]
- ElijahLynn 1mo agoI believe Sam is working towards the future. I do believe in his good parts. I have listened to a fair number of hours of interviews with Altman and do believe he is sincere.
- 1mo ago
- colingauvin 1mo agoWhere is the actual evidence of distillation? I keep seeing this repeated ad nauseam but I must have somehow missed the evidence.
- voxic11 1mo agoDistillation a pretty well documented technique that actually pre-dates LLMs https://arxiv.org/pdf/1503.02531 https://arxiv.org/pdf/1503.02531 Here is a project that guides you through it if you want to prove to yourself that it works https://github.com/arcee-ai/DistillKit https://github.com/arcee-ai/DistillKit
- maleldil 1mo agoI took GP as asking for evidence that the reduced-cost Sol is actually a distillation of the previous-cost Sol. AFAIK, providers distilling or quantising models and offering them as the same model have not been proven.
- sebzim4500 1mo agoI doubt he was claiming that. He's probably saying that the ability of Chinese companies to be able to distill frontier US models has put downwards pressure on the price of all models.
- colingauvin 1mo agoI'm saying that it is unclear that without distillation this wouldn't still be happening. There is a massive narrative that no one but OpenAI, Anthropic, and Google can make a model without distilling. But there's basically no evidence of that.
- byzantinegene 1mo agoI think the (non-definitive) evidence is that the chinese models are always just slightly behind the publically released US models, and never ahead.
- state_less 1mo agoEven before LLMs, ML folks were already aware that you can use a model to teach another model. I doubt this is something AI companies put at the top of their investor materials, but it's been nice to see it play out. That said, there are other moat factors like, a US company needing to use a US AI provider, sticky customers due to corporate onboarding friction, and others. Not nothing, but not as large a moat as some imagined.
- eigenspace 1mo agoYes, but 10 or 15 years ago, I would have thought that there'd be more to it than just a slight modification on the ideas behind a CNN to get this level of AI. There were somewhat good reasons to think it needed more than just this data-driven ML approach.
- state_less 1mo agoThere's something startling about how (relatively) simple these networks are and yet how powerful they are. The main ingredient the AI darlings are using is vast amounts of compute and data. I don't want to take away anything from what the researchers came up with, but I suspect even they are surprised at how capable some of these models have become.
- chasd00 1mo agoearly on there was a lot of talk about "emergent behaviors" in the models where they were good at things that were unexpected or did not align to the training data. IIRC doing arithmetic is one example from early on. I think this is where the AGI craze took off, the labs were throwing more and more data in the training to see what other behaviors would emerge. The thought was with enough data and enough parameters AGI would surface on its own. Then i think tool use became a priority or at lest a sibling priority to more data/more params. Along with multiple specialized models communicating with each other which is sort of a special case of tool use. That pretty much brings us to today.
- 1mo ago
- lerchmo 1mo agoThe internet created lots of monopolies with network effects and economies of scale.a low margin commoditized business that still attracted a trillion dollars of investment to get off the ground was not how I envisioned it happening either.
- mullen 1mo ago> Rather, it seems that selling intelligence might end up as a race to the bottom. Personally, I came to this conclusion early this year. To acquire the data that AI Companies are using to train their models is low cost and once they have it, they can refine and store it. Creating the LLM takes a bit of money but it is not a serious blocker. Clearly, the Chinese companies can make AI so they will drive down costs. There is a need for good AI (Not just Great AI) and it is not cost prohibitive to make good AI (The same with specialized AI). My prediction is that AI will spilt into two categories, Great AI (High Cost) and Good Enough AI (Low Cost). Which for the long run of AI and companies that use AI, this is good.
- dlandis 1mo agoHmm, don't people think that if the frontier labs really put enough engineering effort into preventing distillation that they would be able to do that, or at least diminish it significantly? I'm sure there are variety of additional techniques they could use on top of what they already do, but I suspect it just hasn't been at the top of their priorities yet. Maybe that will change soon. Worst case they could add additional hurdles to account creation ("know your customer" type of thing).
- taf2 1mo agoi trained another AI on all my codex logs... it's pretty good actually
- chrsw 1mo agoOnly the Chinese authorities can stop Chinese labs from distilling from western labs. And they won’t do that, for obvious reasons.
- btown 1mo agoAt the end of the day, while you can do your best to obfuscate your reasoning tokens, it's a losing battle to hide actual user-visible output tokens. The very nature of API offerings is that you can't do KYC on where that API's output is going - there's a rich secondary market that's not going away. And with the sheer volume of data created from that, coupled with benign-seeming prompts like "plan out your reasoning in a document before implementing" that could never be patched without breaking existing customer workflows... there's more than enough for someone to distill on. Even if that only gets them to not-quite-frontier, if you're pushing the frontier every few months, they're only ever a few months behind you.
- zarzavat 1mo agoEven if it were possible it wouldn't change the outcome. China is capable of training frontier models even without distillation. Distillation is only an accelerant. The primary resource you need to train LLMs is money and China has plenty of that.
- miki123211 1mo ago
- maxgiraldo 1mo agoOpenAI could still have a significant moat. ChatGPT occupies most consumers’ minds when they think about AI and has become a household name. Google won because search became a habit-forming product people grew accustomed to using. Bing was once effectively indistinguishable from Google Search, yet still failed to achieve mass adoption because users had already become accustomed to “Googling” things. The same could be said for people "ChatGPT-ing" things. If OpenAI and Anthropic are smart, they will maintain similar pricing rather than aggressively undercutting each other, allowing the market to resemble Home Depot and Lowe’s, or cloud computing, where AWS, Google Cloud, and Azure coexist as highly profitable competitors. Unfortunately, I doubt OpenAI or Anthropic will pursue this strategy, as both companies appear to be acting as though the race to AGI is winner-take-all even if the market may ultimately support several highly profitable competitors.
- nozzlegear 1mo ago> If OpenAI and Anthropic are smart, they will maintain similar pricing rather than aggressively undercutting each other, allowing the market to resemble Home Depot and Lowe’s, or cloud computing, where AWS, Google Cloud, and Azure coexist as highly profitable competitors. Wouldn't that just be price fixing? If they arrive at their prices independently and they all happen to be similar, fine. But if they're all "smart" and coordinate so none of them undercuts the other, that's probably illegal.
- tccole 1mo agoIllegal for sure but rarely enforced.
- maxgiraldo 1mo agoMy understanding is that collusion among competitors is illegal (although I’m not a lawyer). I was referring instead to the prisoner’s dilemma that Bruce Greenwald discusses in Competition Demystified. In theory, competitors, like prisoners who are pitted against one another to rat on one another, are usually better off cooperating rather than turning against each other.
- petercooper 1mo agoIt reminds me conceptually of the idea of using a ST:TNG replicator to just give you another replicator of your own, or asking a stereotypical genie for "infinite wishes". The genie is indeed out of the bottle in many ways.
- demibabs 1mo agoI guess it’s more like asking the paid genie to give you a new cheaper genie.
- jaggederest 1mo agoAnd for a lot of non-frontier purposes these days, you can bootstrap via LLM-as-judge so your hyperspecific wakeword model or whatever can be trained with little to no human input, that aspect of it is fully terrific. The frontier models are a replicator that can give you another replicator which specifically produces tea, earl grey, hot, when you push the single button, and does nothing else.
- gfody 1mo agothe moat is real. the big expensive base models are like the data collected from huge particle accelerators - there's enough unknown structure to be mining for years. you can extract features with more and more generation loss but access to the raw weights is a real advantage, and literally a moat if the interesting behaviors are fenced off
- nonethewiser 1mo agoI wouldn’t quite call it a “race to the bottom” because the costs to produce the models aren’t actually decreasing.
- deleted 1mo ago[deleted]
- chrismsimpson 1mo agoIntelligence ended up being an equalising force. Kurzweil kind of predicted this, but SV was too obsessed with total world domination.
- Razengan 1mo agoWas it not obvious that the value and advantage was going to be in AI-adjacent services? The quality of the harness UX, and random fun crap like Sora, it's a shame that OpenAI killed that so soon, and also Group Chats in ChatGPT.. they risk running a Googlelike reputation at this rate Maybe ultimately whomever can be the "Apple of AI" will win
- tshaddox 1mo agoFor what it’s worth, “race to the bottom” typically refers to a scenario that we absolutely do not want as a consumer. We do want a highly competitive market that drives prices down, but “race to the bottom” specifically refers to a scenario where firms compete by minimizing quality, regulatory oversight, consumer/labor/environmental protection, etc.
- eigenspace 1mo agoIm sure that's well on its way.
- weird-eye-issue 1mo agoWell aside from quality it kind of seems like those other things are getting skirted by. We are very much in a phase of let's see if this is possible and exploit it rather than should we actually be doing this And I'm saying this as somebody that's made millions selling AI software in the last few years...
- redox99 1mo agoIt's a mistake to think only OpenAI and Anthropic are actually spending the big bucks on pretrain, and the others just distill that. The Chinese models are pretrained on large clusters just like OpenAI ones are. Yes, they use outputs of the frontier models to further improve the final model, but even without those outputs they'd still have very strong models. It's not like in a world without distillation things would be much different as you claim.
- giwook 1mo agoThey'd still have strong models without distillation, but strong enough to challenge frontier models and to claim the meaningful market share that they have? Probably not.
- truncate 1mo agoOr, they can figure out something else out? I recall couple years ago when China didn't have enough GPUs (still don't?), DeepSeek team figured out how to train with less computing. IIRC they made Mixture of Experts mainstream and made really optimized kernels and clever use of PTX instruction set.
- ACCount37 1mo ago"China does it in a cave with a box of scraps" is a myth. Chinese labs play the shell game to get their hands on a lot of compute outside China. Tricks like distillation save compute in the RL leg of the process - where a lot of the frontier labs puts their own training run compute.
- truncate 1mo agoI'm sure they do, but its not 0 or 1 thing.
- eigenspace 1mo agoI should have been more clear. While distillation is part of how we got lucky here, what I really think is that it's just surprising and lucky that such a heavily data-driven approach ended up being so powerful here. Transformers are like just a step or two removed from being fancy convolutional neural networks. I guess I'm just surprised that it didn't turn out to require more 'special sauce' with extremely elaborate internal architectures, and less of a big-data approach. Because the data is so central in building these LLMs, rather than some special insights or ideas in the model architecture, or very special hardware requirements, the field is much more open than I would have guessed some years ago. And it's the fact that the data is so central that makes distillation possible in the first place.
- MetaWhirledPeas 1mo ago> The fact that AI models can be so easily distilled and replicated is such a stroke of luck. Sort of. It means the country on the verge of monopolizing all aspects of hardware production (China) doesn't need to rely on outsiders for the software. So while that weakens one monopoly it strengthens another.
- byzantinegene 1mo agoChina is still quite far from monopolizing hardware production
- NegativeLatency 1mo agoCan you buy a laptop or phone not made in china, not made from parts from china? At a regular store not some weird nerd laptop for normies.
- Gud 1mo agoChina doesn’t produce the chips that AI runs on.
- jml78 1mo agoThat was true a year ago but is no longer true. Deepseek is training and running inference on Huawei chips. The times of China needing Nvidia chips is quickly coming to an end. I guarantee that China will scale faster, build faster, and ultimately produce far more chips than the rest of the world combined in 5 years. Our export controls sank the west. It would have been better to allow them to use Nvidia chips. Now they will have chip fabs that aren’t quiet as good but way more of them. Their investment into sustainable will make the power so cheap that the less efficient chips will not be a relevant issue
- brookst 1mo agoCan you name one with parts ONLY made in China? Not Taiwan, not Vietnam, but China proper?
- navaed01 1mo agoHow much of this ability to replicate is down to the openness of the science community and the paper on transformers being accessible by anyone?
- eigenspace 1mo agoEven if the transformer paper wasnt published, the info would have diffused out eventually. Its not that wild of an idea. It's not like e.g. chipmaking where even just knowing how things are done doesnt mean you can copy it.
- TZubiri 1mo agoOpenAI implemented measures to reduce reverse engineering, following the Anthropic lead. They disabled the temperature and seed parameters. There's still logprobs, so they aren't as closed up as Anthropic yet. I might write about an article of the history of LLM APIs, I used to think the ChatGPT was going to be a de facto standard like intel's 80866 mutated into x86, but it seems to be a bit more nuanced and diverse than that, vibecoding introduced so much complexity because the vibecoding product itself became vibecoded so the enshittification was accelerated, many such cases.
- novok 1mo agoMeh distillation doesn't mean you can create an existing model from scratch of similar quality. It's the AI equivalent of making a VHS copy of a video, it doesn't enable you to make your own movies very well and post training is the equivalent of video editing, which again, doesn't let you make your own movies very well. Your seeing the AI labs respond by never publishing chain of thought now and in the future, I see them not even publishing their top models as a general purpose API and instead using it to drive their own AI apps, which will obscure even more model output. Anthropic Mythos was internal only for many months for example.
- raincole 1mo ago> The fact that AI models can be so easily distilled They're not, right? If it's really easy why there are no counterparts of DeepSeek from the Europe or Japan?
- davrosthedalek 1mo agoBecause it breaks the TOS.
- matchagaucho 1mo agoWe're mostly paying for AI delivery, performance and SLAs at this point. GPT 5.6 Luna at $0.20 per 1M is pretty good for 80% of enteprise applications.
- fooker 1mo agoI think this indicates we are eventually going to stumble upon much more efficient ‘intelligence’. Probably not through what we call distillation now. There’s some magical technique hiding there, go find it!
- scotty79 1mo ago> The fact that AI models can be so easily distilled and replicated is such a stroke of luck. Why shouldn't it work though? It's just models teaching other models same way humans are.