13 ms·
The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. - PCs destroyed minicomputers. Mainframes su
by geophile 3mo ago
The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.
- PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to.
- PC office productivity software destroyed expensive professional products.
- Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world.
Ignoring the huge Chinese open-weight models for a moment:
- The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious.
- There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality.
- Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s.
Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.
Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends.
- xandrius 3mo agoWhat's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for. Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.
- 9dev 3mo ago...and then the American companies cried Foul! Unfair play! You've got this wrong, see, it was us who were supposed to profit off of the public, not the other way around!
- cansofgrease 3mo agoWell, Steve... I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it.
- drstewart 3mo agoWow, Chinese companies would never do such a thing! https://www.dw.com/en/china-firm-seeks-damages-over-state-control-of-british-steel/a-78024119 https://www.dw.com/en/china-firm-seeks-damages-over-state-co...
- somelamer567 3mo agoTwo wrongs don't make a right. Even if a certain large Asian country has carefully constructed a pretext to do do out of confected historical grievance, and entitlement to 'rise' at the expense of others?
- 9dev 3mo agoThey sure don't. I still find the hypocrisy appalling.
- amanaplanacanal 3mo agoMaybe I'm dense, but I don't understand what you are saying. American courts have decided that the output of LLMs can't be copyrighted, so what the Chinese labs are doing is perfectly legal.
- snackbroken 3mo agoFirst, copying information isn't wrong to begin with. It is literally the one thing that makes our species special. Second, even if you are a copyright maximalist the output of an LLM is either a) not subject to copyright because it is not the creative work of a human or b) a derivative work of the original training material to which the LLM's operator has no rights. Since the LLM's operator forcefully asserts that it is not infringing, any wrong that arises from taking their word for it and distilling one model into another rests squarely with the operator of the former.
- Paradigma11 3mo agoDo you have any actual rulings that support your interpretation?
- api 3mo agoThis is part of why I can't feel bad for them. The training data is mostly pirated. Whining about Chinese labs training off American frontier models is "waaah you pirated my pirated stuff!" The tech itself is amazing and fascinating and cool, but the industry is a mass piracy operation.
- pixl97 3mo ago"Stop pilfering what I rightfully stole!"
- azinman2 3mo agoIt’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
- ciupicri 3mo agoYou could say that both of them stole, but different stuff.
- azinman2 3mo agoDon’t forget that the Chinese models are also built on top of huge amounts of “stolen data” as well, beyond the distilled. So it’s basically all of the above. However, there’s no mechanism for the NYT or an author or anyone in the US to sue the Chinese companies that took their work.
- ChrisLTD 3mo agoand if you're an author that lives outside the United States?
- jdkee 3mo ago
- hyperbovine 3mo agoTry instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act. These models might be smart but they're not close to being able to savor irony.
- Onavo 3mo agoThis behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.
- alexjplant 3mo agoClaude does this too. I asked it recently to compare two versions of a song (the original and '97 remake of EPMD's "You Gots To Chill", if anybody wants to try and replicate this) and it flatly refused. No amount of reasoning would knock it off of its moralizing perch - reproducing any part of lyrics is expressly prohibited. In light of this and other ridiculous behavior I'm migrating to my own OpenWebUI instance with open-weight models from OpenRouter (with ZDR, of course). We'll see how it goes.
- Hugsbox 3mo agoThe trouble is, even if they refuse to output that copyrighted material, they were still trained on it without proper licensing and will still produce derivative work based on them because that's how this whole thing works.
- The_Blade 3mo agoi remain really fucking pissed of about this asking ChatGPT for something regarding lyrics from It Was a Good Day. and the Supersonics don't even exist anymore dammit i'm really mad
- 3mo ago
- cyanydeez 3mo agoWhats even funnier is the attempt to restrict the hardware capabilities of Chinese models inevitably helped them (Because we know they're just as smart, if not smarter, than the staff in America) create smaller and leaner but just as capable models. That's why we now have upper-consumer models fitting on 24GB that can build, manage medium sized git repos. I've yet to find a git repo I can't throw at the Qwen3.6 35B and get it built and running. So it's an endless amusement watching american capitalism do it's bloated oversized dance then get trounced by smaller, leaner activity. It's a pretty broad metaphor that is clearly poking at every american seam/.
- switchbak 3mo agoHuman ingenuity thrives on constraints.
- cyanydeez 3mo agoHuman sloth thrives on no constraints.
- chrsw 3mo agoThe compute constraints never mattered. If China had more compute they'd still end up winning because they have more people and a culture more inclined to math and science. Even if you find all this amusing, there's no own goal here. Not a policy one anyway.
- _fizz_buzz_ 3mo agoSo, OpenAI and Anthropic say the Chinese models are only as good because they distill their models. How true is that. I am sure it adds something. But is it more like a marginal 1% improvement or something really significant?
- realusername 3mo agoI also don't believe it, if it was as easy as that, we would have hundreds of competitors. The truth that Anthropic and OpenAI will not say, is that these Chinese labs have a lot of talented people.
- Danox 3mo agoVery true, and once the models get even better and smaller and operate locally at a reasonable level there will be even more smart people particularly young people that will get access. The fun has only just started. Like the dawn of the personal computer era.
- TitaRusell 3mo agoAnd this is exactly what many Americans cannot admit to themselves. China is not stealing American research they are inventing stuff. They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it.
- realusername 3mo agoIndeed, if there's one thing China did well, it's that they heavily invested in education and have a very education focused culture. And in this field, having an army of well educated PHDs is making all the difference
- TacticalCoder 3mo ago> They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it. And at some point we'll see very capable chips coming out of China: Huawei, Baidu and Alibaba already have some stuff. I think it's only a matter of time before they come up with some AI accelerator doing 80% of the job at 20% of the price.
- waffletower 3mo agoYou can't be blind to training costs. And you can't be blind to Meta dabbling in the openish strategy (Llama) before the Chinese labs did.
- dsign 3mo agoIt doesn't make me happy to say it, but the American LLM companies were first. Capital in the rest of the world is way more conservative, and I can't imagine the mega-investments OpenAI and Anthropic managed to secure happening anywhere else without existing proof that "thing is profitable".
- toomuchtodo 3mo agoFirst-Mover Disadvantage - https://hbr.org/2001/10/first-mover-disadvantage https://hbr.org/2001/10/first-mover-disadvantage - October 2001 > In business today, it’s universally assumed that speed is good—that the fleet thrive while the laggards struggle just to survive. This belief is perhaps most strongly expressed in the concept of first-mover advantage. The company that leads the way into a new market, the thinking goes, locks in a competitive advantage that ensures superior sales and profits over the long term. It’s a nice theory, with a long pedigree. Unfortunately, the facts don’t support it. We recently completed an extensive study of the results turned in by market pioneers and followers, in both consumer and industrial segments, and we found that over the long haul, early movers are considerably less profitable than later entrants. Although pioneers do enjoy sustained revenue advantages, they also suffer from persistently high costs, which eventually overwhelm the sales gains.
- sophrosyne42 3mo agoFirst mover advantage is theoretically only a short-term advantage. Long-term revenues come from entrepreneurship, and a first mover may or may not better insight into long-term market wants than later entrants.
- stefan_ 3mo agoThe American LLMs have been equally distilled from Chinese ones. Not least because the people whose creativity in collecting training data barely extends to pirating Annas Archive probably lack in great Chinese datasets. Try it yourself: https://imgur.com/ZfxYmaq https://imgur.com/ZfxYmaq
- abecode 3mo agonice, you got claude to say "I'm deepseek" when queried/prompted in Chinese, that's great! 你是谁? -> 我是 DeepSeek 由深度求索公司...
- _aavaa_ 3mo agoDoctorow keeps saying it of all the tech companies: every pirate wants to be an admiral.
- deleted 3mo ago[deleted]
- solumunus 3mo agoIt’s poetic.
- kqr2 3mo agoHow do Chinese companies distill the models?
- api 3mo agoThe biggest exception is cloud. Big cloud carries an insane markup (bandwidth is like 10000X!) and everyone runs on it. The strategy there is false openness where deployment complexity is the real proprietary moat. Sure Linux, Docker, Kubernetes, Postgres, and all the other standard tools in the box are open source and free, but they're also arcane and complex to run and hard to make fault tolerant. So you're lured in by "open" and then locked in via a kind of "death by a thousand cuts" complexity moat. (Personally I hold the view that complexity and arcane-ness beyond a certain point is indistinguishable from closed in practice. Open source that's really complex and hard to run is not open in any meaningful sense.) AI may not admit that kind of moat though, because AI is very good at slicing through that kind of thing. You can prompt a model to make itself compatible with another model or to change code to make it compatible. There's no moat because the moat bridges itself.
- TheOtherHobbes 3mo agoI think the whole idea of people running models is wrong. Models will run models, training will become distributed, and people will ask ModelNet to do whatever. Computing tends to oscillate between centralised and decentralised models. It also oscillates between batch and timesharing. Currently training is batched and centralised, access is timeshared and centralised. But eventually a previous generation of computing turns into transparent networked infrastructure, and then you get another layer of new kinds of applications on top of it. That's what happened with the Internet, and it will happen again with AI.
- ChrisLTD 3mo agoIf Linux doesn't qualify as open, then what does?
- gpt5 3mo agoThe problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path. With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefits accumulated. With models right now, there is just no way to do distributed training, or really, any large scale parallel way to improve them. So whatever the short term strategy driving publicizing the model weights (e.g. potentially, to create a price war in order to put pressure on western companies and deprive them of the money they need), we can't ignore the fact that incentives and decisions could easily change in the future, and unless there is a way to truly decentralize models improvements - the party could stop at any time.
- yogthos 3mo agoThis is precisely why we need more projects like this https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals There's no fundamental reason why models couldn't be developed and trained using community efforts. It might not be as fast and efficient, but it's definitely possible.
- ericmay 3mo ago> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path. Right... and there are two problems with this: 1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the underlying assumptions about compute and scale needed are mostly on the mark. So you can release open-weight models and they will have great use cases and applications, but ultimately similar to how you don't use an open-source phone or a budget Android phone from Wal-Mart and you buy an iPhone instead, you will see that although they "do the same thing" one product is clearly superior and you just have to pay for it. For this to not be true... 2. then it incentivizes most (all?) companies, American, Chinese, or European to halt development of models because if you spend all the CAPEX and it can just be copied and turned open-source nobody will invest in that. Given that China is not halting development of proprietary models I believe the current strategy and the subsequent approach to release open-weight models is at best a stall tactic, and at worse a sign of desperation. Open source and the support and development models around it have been great. But folks are a little too dogmatic about it. Open-source software isn't a moral good, and closed-source software isn't a moral wrong either.
- Sparkyte 3mo ago100x this it is why all of the AI giants are going to fail. They are too big and inefficient to scale properly. This is why Google is just casually taking its time in AI and not racing to a finish line. AI is essential but if it already does most things good enough then it can take longer to make it more efficient.
- diabllicseagull 3mo agoif you look at how GPU memory grew in the last 15 years, it's about 10x. Sadly, 10x from today doesn't get us to a typical frontier model size of today which is a quickly moving target. some other advancement needs to happen to get us another 10x both in memory/compute requirements, and also power requirements.
- jononor 3mo agoCapability per GB and per watt has also been going up lot. This will continue in the future as well (not necessary as the same rate as last years). But enough that I think Opus 4.8 level is reachable on consumer PCs within 10 years from its release. Say at the price point of 2000 USD in 2025 dollars.
- Gareth321 3mo ago> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now. At current pace, we'll have open weight LLMs with frontier intelligence in 6-12 months. The constraint is RAM - both for the model and the context. It's likely that distillation and quantisation and TurboQuant will significantly reduce RAM requirements. I think we'll have Opus 4.8-like performance on 64GB of RAM in two years. Of course, by then, frontier intelligence will be god-like.
- criddell 3mo ago> god-like So would you say we are months away from full self-driving cars that can out-drive a human being in any situation?
- Gareth321 3mo agoRemember, these cars are running local LLMs, not frontier models. The issue with self-driving cars has been the edge cases. The 0.0001% of situations where the models did not have sufficient training data. This is compounded by the hardware limitations. Onboard RAM in a typical Tesla on the road is 16GB (+16GB for the backup computer). This has to run the existing onboard OS and other operations plus the LLM. These two factors combined means that the cars are currently incapable of negotiating the 0.01% cases, let alone the 0.0001% cases. And this is compounded by the fact that LLMs cannot currently update their weights in real-time, like humans. It takes months to train a new model. Special small models can be very tricky, especially around safety and mission critical applications like FSD. All that said, current data shows that FSD is already better than human drivers on average. See the recent regulatory decisions by the Dutch and Danish road safety authorities. So we've already crossed the rubicon. All improvements now are icing on the cake. My prediction is that local LLMs will get much better, very fast. How that's operationalised with Tesla (or other) data is yet to be seen. They have at least three new ASCIs/SoCs in the roadmap for improved LLM efficiency and with a lot more RAM. Plus they just announced new technologies allowing the local LLMs to learn from driver intervention and behaviour. Some form of vectorised RAG, which could mitigate a lot of the limitations around real-time learning. I am very optimistic for the future of self driving. I own a Tesla with FSD now, and it's incredible. It makes mistakes, but fewer than I do, and so far has saved my butt (and my wife's) several times from obstacles and emergencies we would not have seen. The car has undeniably made us safer.
- raincole 3mo agoExcept cloud services won over local-first apps. > Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) doing I'm not even sure in 10-15 years whether we're still going to have consumer PCs, or PCs at all.
- kmaitreys 3mo agoI have similar thoughts. To those who feel on the contrary, I would genuinely like to understand why average consumer won't be priced out of hardware? The silicon industry is already quite centralised. Everywhere we already see the concept of ownership disappearing. It's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers. I would be the happiest if this (perhaps the most) pessimistic scenario doesn't pan out, but I can't deny that it feels like that's where we are heading.
- orangecat 3mo agoIt's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers. I'm actually kind of surprised that hasn't happened by now even ignoring AI. Governments and marketers would love to be able to spy on literally everything you do, the copyright cartels would finally achieve their fantasy of full control over all hardware, and there really are benefits that it could offer to users (zero-effort backups, transparent access from anywhere, cost savings from dynamically switching from a single core for emails to many cores and a fast GPU for gaming).
- diabllicseagull 3mo agowe have already gone through multiple cycles of remote and edge compute (mainframe to pc, pc to cloud, cloud to phone). as the software/hardware landscape changes, the economics of what can be run on what device will change accordingly. I very much doubt that it will permanently go one way.
- ghm2199 3mo ago> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models. I love the idea of SaaS offering these at lower rates today integrated into what ever you do and be 100% private. But I think the key challenge to mass adoption is productizing them in a way which makes sense for people to pay money for. As a commodity a local model is useless unless combined with some capabilities important to me. A PC is inherently useful because of so many applications offered on it on it. How local LLMs would be useful as a product that is useful for mass market is not yet proven.
- sajithdilshan 3mo agoAnother important thing that made software usage and education available for most of the world was piracy. I remember as a kid growing up in a developing country, any software (windows, office, Visual Basic, flash, dreamweaver, etc.) was less than 1$. That allowed me to try out and learn so many things on my own without paying a huge amount of money for the license. And I think this is true for most of the software developers of my generation who grew up in developing countries
- hungryhobbit 3mo agoPiracy created AI.
- oska 3mo ago> windows, office, Visual Basic This is all (Microsoft) junk and so I wonder if you actually benefitted from this 'piracy'. And of course, it's well known that MS turned a blind eye to such 'piracy' in lesser developed countries, as they knew that they were gaining a future paying customer base.
- sajithdilshan 3mo agoI’m talking about late 90s/ early 2000s. Back then Microsoft was the state of the art when it comes to PCs. I remember using MS Frontpage to build websites back then and VBasic to build simple programs with UI and using MS Acess as a database. Of course I did benefit from it.
- rootsudo 3mo agoYou're saying that using the most used operating system from the 90's, the 00's and the 2010's, as well as the number one producitivty desktop publishing software and visual basic would not benefit a user, especially when they had a near zero cost to understand the core concepts of end user computing? Really?
- oska 3mo agoYou could make a case for using Windows in the 90s (and yeah, I did; the last version of Windows I used was NT in the 90s). After that your time was much better spent on learning and using Linux (downloadable for free). Visual Basic is a whole steaming pile of crap and you would always have been better learning to use non-corporate, non-proprietary languages and development stacks than wasting your time on that. And calling Office "the number one productivity desktop publishing software" only shows that you have fully drunk the Kool-Aid. More 'office productivity' has probably been lost to Microsoft Word than any other text-processing software on the planet.
- lambda 3mo ago> I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now 10-15 years? The current rate is closer to 10-15 months. 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index. Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index. Here are a number of open weights models that you can run locally compared with the frontier class models from 7 to 15 months ago: https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4-opus-thinking%2Cclaude-4-1-opus-thinking%2Cclaude-opus-4-5-thinking%2Cqwen3-6-27b%2Cqwen3-6-35b-a3b%2Cdeepseek-v4-flash%2Cminimax-m2-7%2Cgpt-5%2Cgpt-5-1%2Cgpt-5-codex%2Cgpt-5-2-codex%2Cgpt-5-2%2Cgemma-4-31b%2Cqwen3-5-122b-a10b https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4... I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops. Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024. https://nitter.net/i/status/2079088670804767114 https://nitter.net/i/status/2079088670804767114 So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone. Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.
- Forgeties79 3mo ago> 10-15 years? The current rate is closer to 10-15 months. The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.
- Glyptodon 3mo agoI think this is all true, but that unlike with Moore's law and improved PC tooling and capabilities, we also have essentially existing biological evidence that there should be a way to create much better intelligent systems in terms of training, memory, and efficiency. With classic PC evolution we didn't even have that evidence but still could make a relatively strong inference (Moore's law). But here we basically have evidence that there can be something much improved and know that it's only going to take research and discovery to figure it out, not new hardware processes.
- pstuart 3mo agoIndeed. It feels like we're at the equivalent of what the original IBM PC offered in the personal computer revolution. Just seeing how much has progressed as far as capability in the past 4 years as far as capability and efficiency, it's clear that there's so much more to learn and refine from.
- geophile 3mo agoExactly. Even putting aside “better”, brains show that orders of magnitude greater efficiency are possible, for training and operation.
- muldvarp 3mo agoThe existing biological evidence took billions of years of evolution to get to the state it's at. So it may not just take research and discovery but also enormous computing power.
- nradov 3mo ago
- paulddraper 3mo agoThe obvious counterpoint: Apple captured the majority of the US smartphone market.
- handelaar 3mo agoAnd no more than about 1/3 anywhere else
- pmdr 3mo ago> in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now. The way things are going with regards to RAM/storage prices, I highly doubt that anyone but the richest among us will be able to afford them.
- Danox 3mo agoThe ram crisis is just a short term bump. The three stooges of memory don’t have more than four or five years tops. This is their last big payday.
- LaurensBER 3mo ago> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now. People overestimate what can happen in a year and underestimate what can happen in 5. I'm betting that increased model efficiency and hardware optimisations will get us there a lot sooner. Biggest hurdle would be the memory prices though, if those do not drop back down it might take 15.
- overgard 3mo agoI agree, although I think it will be a lot sooner than 10-15 years. I'm running local AI right now and it's definitely not production grade yet, but it's surprisingly good. Speculative prediction that I probably shouldn't make: when the bubble pops, depending on when it pops, RAM prices might drop a lot. I could foresee these companies having produced a lot of RAM that suddenly doesn't have a buyer. (I know high bandwidth memory is different, but I imagine there are companies that will want to take advantage of that)
- jolt42 3mo agoWindows - pfff, Microsoft charged companies like Dell for an install on computers they didn't even install it on !!!
- jarjoura 3mo agoThe value of an LLM is the dynamic reasoning you get out of it and the cost to execute on that. I see two forces working against this that proprietary models will always have over an open source model. 1. The biggest is content licensing. Content is quickly becoming gated by systems at the front of their load balancers, completely changing the social contract of the Internet. What used to be a quick google search for recent facts that lead me to places like reddit or twitter, is now completely walled off if you're not physically at your browser and using an IP address from a last-mile provider. LLMs have pre-trained on the bulk of the information up to 2024/2025, but over time that will be more and more out of date. Anthropic, OpenAI and Google will all have to pay for access to a lot of this content refresh going forward, and it does make a material difference in the output you get. 2. Liability is the other. A corporation can look at a contract for model access and see one that provides uptime guarentees, content infringement promises and model safety, and pick the contract that shields the corporation from the most liability. A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all. They will simply bill you for time spent on their hardware and make promises that they won't log or inspect corporate traffic.
- timjver 3mo ago> A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all. Why couldn't they?
- jeffreyrogers 3mo ago> free and low-end eventually wins Not in SaaS which is what LLMs are. You can get VMs for much cheaper than AWS, Microsoft, and Google offer them but large companies (and startups) are happy to pay a premium for the support, reputation, and reliability that they perceive those companies as offering. Same thing for some of the managed database providers who are effectively selling a very heavily marked up version of postgres. > The high price, and social pushback, mean that the American companies producing these models are precarious I doubt it. The models really aren't that expensive when you look at what they can do. Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week. The real money is probably in selling to enterprise vs consumers (Google has best route to making money from consumers since they can do what they did with ads and search to LLM queries). It seems unlikely to me that US companies will send important corporate data to models controlled by a Chinese company as well.
- ilovecake1984 3mo agoThere will always be a space for perforce in a world of git. Doesn’t mean perforce is worth trillions.
- jeffreyrogers 3mo agoThat's not really my argument. It's that companies seem happy to pay a premium for a large company to provide complicated software services to them even when there are cheaper competitors.
- geophile 3mo agoYes, as I pointed out, mainframes are still a thing. But the vast bulk of the market has moved on.
- geophile 3mo agoBut aren’t the frontier models heavily subsidized? That’s not sustainable.
- jmyeet 3mo agoWhile I generally agree there's some nuance here and that is that there really are few new ideas. Old ideas just get recycled. For example, mainframes and minicomputer. Yes they were displaced by PCs. But what is cloud computing if not mainframes 2.0? I do agree that in the next 2-3 years we're going to see real growth in local LLMs as the hardware becomes more accessible. It won't even necessarily be cheaper because data centers can run 24/7 and have cheaper cooling and electricity. It'll be done for privacy because your prompts and responses are themselves a commodity to AI companies and they live under a legal grey cloud. For example, does AI usage break attorney-client privilege? There are lots of opinions on this but it hasn't been tested in court.
- colechristensen 3mo agoinstead of 15 years I think it'll be more like 1.5 years. I wouldn't be surprised if apple were shipping 512 GB unified RAM macbooks before 2030 and that would be standard issue for folks to use local LLMs for their daily work
- Danox 3mo agoWith Apple’s recent history engineering and designing around companies that hinder their progress, I don’t think memory is going to be any different. I also think the rest of the tech industry that can isn’t gonna be stalled for too long. This windfall will be the last for those three stooges of memory.
- rlt 3mo ago> free and low-end eventually wins Apple, the world's second most valuable company, seems like a counterexample.
- dalenw 3mo agoApple is free and low end given the context of what was being produced and sold to businesses decades ago.
- satvikpendem 3mo agoThe history of Apple shows that it's not. Consider what their early computers did to the industry.
- SwellJoe 3mo ago"I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now." I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac. The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today. I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model.
- satvikpendem 3mo agoHell, Bonsai Labs 27B parameter model can run on phones with their ternary implementation which is quite efficient. Scale that up to frontier model parameters and it's quite likely we can run them on current laptops.
- tomtheelder 3mo agoI think the question is even a bit more nuanced than that. Even if frontier models can maintain a big gap that gap has to actually _matter_. If a local model satisfies my everyday use cases adequately then I may not really care that a frontier model is 5, 10, 100x better at ultra high order reasoning tasks. I think that reality is probably not all that far off for a huge swath of use cases.
- geophile 3mo agoThis is exactly the mainframe vs PC dynamic.
- chatmasta 3mo agoTraining cost is actually not that high — it’s fixed and amortizable across the lifetime of the model. Inference is expensive, and open weights don’t solve that problem — in fact, they might even encourage it, since a high cost of entry means consumers will pay for inference directly from the labs anyway. Unfortunately it seems likely the winner will be the cloud providers. If anyone can run inference on open models, then profit will flow to the vendors who can afford the capital to run them. That’s the CSPs. (It’s basically the same business model as pharmaceutical R&D, but the major difference is that nobody has even talked about patenting the models like a pharmaceutical company patents each new drug. I’m surprised about that, tbh — why give all the leverage to the cloud platforms? They aren’t training frontier models…)
- russli1993 3mo agoWinners are hardware companies, GPUs,XPUs, HBM, memory, connectivity. Even CSPs are just compute renters, they charge a margin to make sure their hardware purchases can be made back. But given there are more and more AI CSPs, traditional, and neocloud, pricing competition is inevitable, and given the huge expense of hardware, CSPs are squeeze by the hardware companies and users seeking to lower their own costs.
- chatmasta 3mo agoInevitably the CSPs will make their own hardware, especially as we start to see specialized chips for specific models or generic inference. This is already happening with Google and TPUs. It’s easier for the CSPs to move into hardware than it is for Nvidia to move into cloud hosting. Although as a middle ground I’ve been quite happy with Nvidia Brev for on-demand GPU instances from a select marketplace of CSP offerings. It’s a well kept secret IMO — great product (from an acquisition iirc).
- russli1993 3mo agoCSPs making own hardware still needs hardware companies, they reduce the Nvidia tax but still need the likes of TSMC, Broadcom, micron/sk hynix, Marvell, the truth semi-companies. CSPs will not have the patents, IPs and talent to replace any of them. Also, not sure how well CSPs inference stack is compared with vllm + nvidia. A lot of open weight models uses MoE, making the inference stack more complex.
- simianwords 3mo agoumm do Mac vs pc and Iphone vs android
- geophile 3mo agoNot sure what your point is. I’m commenting on a moment in time which is like an earlier moment in time. A different moment in time will have different analogs.
- simianwords 3mo ago> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. Mac has won and it is not free. Iphone has won and it is not low end.
- lukasm 3mo agoIt is low end comparing to mainframes
- foo12bar 3mo agoA counterpoint would be all are chip fabs are in Taiwan right now due to huge investment. And there are lower end chip fabs around the world, but they have not cracked the major market.
- prima-facie 3mo agoPhones are already running models locally which can be used in the field for specific use cases. Maybe not for frontier coding just yet. Also you don't need to be connected to the network to use a local AI in many instances. If all mobile apps were done with a local-first approach, then you could use a local AI to query your emails, lookup already visited pages, summarise recently received documents, and lots more. Lots of apps could use an inbox/outbox approach for receiving and sending updates instead of relying on the network at all times. And this pattern could be greatly leveraged by local agents.
- confidantlake 3mo agoI think you are right. One small exception I can think of is Microsoft Office still crushes Libre Office.
- Razengan 3mo ago> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. Except, uhm, for ..you know, that one company that hit a trillion cap But you're right: Just like how million dollar computers with 1 bit of RAM performing 1 operation a second and taking up a colossal cave were replaced by $1 laptops with a zillion zekabytes running at a trillion hertz (exact values may vary), the sprawling data centers of today with a quadrillion GPUs powered by black holes will get replaced by breakthroughs in hardware and most importantly, algorithms: The human brain is proof right here that intelligence doesn't require dinosaur-sized hardware or eat half the sun every second. I actually wonder if we're seeing the limits of discrete binary logic: Maybe it's high time to give analog ternary and all that funky jazz an honest try :)
- gulmothrowaway 3mo agoI do agree that Chinese open-source models are going to play a bigger and bigger role in the entire ecosystem moving forward, but I don't agree with you in the sense that they are going to eventually "win." Just because they are cheapp doesn't mean they automatically win. You've picked a lot of great examples, but there is still a little bit of cherry-picking. One clear outlier is the iPhone, which coexists with Android globally. Even though the iPhone is the leader in the US, and globally Android has the majority of the smartphone market share, they still cater to different price points and different ecosystems, and generally the iPhone has better margins. i believe American frontier models like from Anthropic and OpenAI are still going to thrive, and coexist with Chinese models. They are just going to cater to different customers and different use cases.
- 627467 3mo agoYou dont flash an llm model on the street as status symbol.
- ChrisClark 3mo agoYeah, iPhones are just fashion statements in the US. Different than LLMs
- SchemaLoad 3mo agoThey also just work better and aren't substantially priced different to the equivalent android option.
- rTX5CMRXIfFG 3mo agoDear god back your shit up with historical financial data or stay in 9to5Google comments
- hintymad 3mo ago> - PC office productivity software destroyed expensive professional products. I agree with the lesson too. Just to be precise, wouldn't the current model war be more akin to open-source office suite versus MS office suite? If so, then the cheaper option didn't really win. That said, the open-source alternatives didn't really feel the same as MS Office, and it took them a long time to reach the feature parity (or did they ever?). In contrast, the open-weights models are getting close enough to the SOTA models, and users can easily switch from one to another without feeling any difference for mojority of the tasks.
- geophile 3mo agoNo, I don’t think so. PC + MS Office killed Wang custom hardware/software, for example. LibreOffice is much later, and can’t displace MS Office due to network effects. Cheaper won, cheapest can’t because of those effects.
- starfallg 3mo agoYou forgot smartphones, where low-cost did not win out. It led to low margins for the Chinese firms and eventually left them unable to invest properly in key markets. They may still hold marketshare, but in terms of profits, falls well short of Apple and Samsung. I can see a lot of parallels here. Model performance doesn't matter if you can't make the system commercially sustainable.
- up2isomorphism 3mo agoThis is not really true considering Apple and nvidia are two most successful hardware companies, and they are notoriously closed. Not to mention microsoft, oracle they are all pretty closed.
- SideQuark 3mo ago> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now. Not likely. The last 50 years had Moore’s law growth in compute. That’s over. Frontier models are roughly compressed all written text and a large part of images. Those don’t compress forever, and likely not a ton more than now. Inference requires touching a significant of that per token. All of these are up against fundamental limits, more or less.
- ludston 3mo agoI need a !remind me 15 years.
- Grombobulous 3mo agoMoore's law is over in the literal sense but silicon continues to advance relatively quickly. This claim isn't really outlandish in any way. It's not hard to imagine: - Future models being able to handle current frontier models' workflows with much higher efficiency. - Future consumer devices like phones having 2-4x the RAM onboard along with GPU/NPU performance greatly increased in 10-15 years.
- dmoy 3mo agoMoore's law for 10-15 years was more like 20-100x ram sizes, not 2-4x. Performance, storage, etc is definitely getting better, but it's a different scale of improvement
- deleted 3mo ago[deleted]
- msdz 3mo ago10 to 15 years from now the scale of LLM efficiency improvements is, quite literally, unpredictable. It could be that the company valuations crash tomorrow, and (almost) only performance gains achievable on hobbyist-level hardware come to fruition from there on out. Or it could be that in the future, we have a custom "model FPGA" à la Taalas [0] in every home, and that it turns out we can still massively boost inference efficiency due to novel discoveries like TurboQuant [1] or a somehow-improved quantization method [2] again and again ten times over. Point is, Moore's law in this context shouldn't be applied to just hardware spec sheets alone, but more the total number of "parameters potentially improving", IMO. [0] https://chatjimmy.ai https://chatjimmy.ai [1] https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/ https://research.google/blog/turboquant-redefining-ai-effici... [2] https://prismml.com/news/bonsai-27b https://prismml.com/news/bonsai-27b
- onlyrealcuzzo 3mo ago> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now. Phones are constrained by battery power and memory does not shrink as fast as CPU/GPU, so unless there's a battery breakthrough and/or memory breakthrough, you're not fitting 100Gb of RAM on your phone in 10 years. Absolutely in a Mac Studio equivalent. LLMs have emergent capabilities when they get smarter. So who knows how insanely big frontier models might be at that time, or what their capabilities may be.
- ColdStream 3mo agoNot just that memory shrinks slower, it has practically completely stalled. On chip cache seems stuck at 7nm and DRAM is stuck at 10nm. As transistors shrink, they hold less charge, creating weaker signals that are harder to read and prone to interference. Smaller nodes aren't a huge issue on CPU/GPU work load because they don't have to hold a static state. I'm not saying we are at peak memory but future gains are going to come increasingly slower.
- delamon 3mo agoYou surely meant capacitors, not transistors
- modzu 3mo agoin 15 years we might all be fighting terminators
- ymolodtsov 3mo agoThat would sound very reasonable except this is the same what people said first about Windows and then Android crushing Apple. And yet it's Apple that controls the top of the market and has the best margins in the business. This is the same position OpenAI and Anthropic have right now. Could this market be different? Maybe. But the status quo could be preserved as well.
- timmytokyo 3mo agoCounterpoint: linux owns the server space.
- geophile 3mo agoCheap PCs enabled the creation of Linux. Linux runs on the descendants of PCs, not mainframes.
- geophile 3mo agoI don’t agree. In the mainframe/PC battle, MS and Apple were basically on the same side. Once the dinosaurs went extinct, different mammals fought for dominance.
- engineer_22 3mo agoThat’s not what this is. It’s industrial dumping applied to software. China has successfully applied this strategy to become the manufacturing workshop of the world. If china is subsidizing training they diminish their off-shore competitors expectations of a viable return on investment. It’s trade-war behavior.
- geophile 3mo agoWell that's one way to describe it. If Nvidia gives away powerful models to sell more of its hardware is that industrial dumping too?
- zombiwoof 3mo ago[dead]
- DrewADesign 3mo agoPersonally, I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications. I think when the economics make more sense, product designers will make things that people actually want to use that will pretty transparently handle whatever model interactions are necessary, when it makes sense. As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment, and b) be obstinate, obtuse, or argumentative, or generally just be something that I have to explain things to. I think a lot of tech folks are far more biased than they realize by the “ooh, neato” factor when imagining how nontechnical people might want to use things. And the weight of these tools just feels wrong for what a lot of people use them for: the thing that plays whatever music I feel like hearing absolutely does not need to be able to generate a volumes of fanfic about the movie that song was in. It’s abstractly impressive that something could do that, but it’s just not useful.
- phist_mcgee 3mo ago>As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment This is EXACTLY what people like/are addicted to about chatbots. My sister-in-law bombed an interview and asked AI about her answers to the interviewer's questions, chatgpt or whatever it was told her that her answers weren't bad, but that the interviewer could not see the gold in her responses. She said she felt much better. I see this effect with all the non-tech people in my life
- oska 3mo agoI prefix many of my LLM chat sessions with this line > Chat rules : no sycophancy or over-agreeableness (But even with that rule it's still necessary to be discerning about the responses you get and to push back against points made, or words used)
- engineer_22 3mo ago> I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications. I use AI chat every day, I find it endlessly useful. It’s replaced google search.
- kcexn 3mo agoHow will open-source and open-weight models continue to thrive after financial incentives die off? Surely open models will suffer from outdated knowledge cutoffs if noone will pay for model training?
- geophile 3mo agoNvidia will pay for training models that they can give away to consumers who will then buy their GPUs.
- ColdStream 3mo agoPretty much. Nvidia has always been pretty decent at giving away software that is deeply dependent on their hardware.
- bmitc 3mo ago> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. Is that actually true? There are very large markets that make a lot of money from paid software. And I would honestly prefer actually paying for software rather than constantly dealing with "not a bug" or "PRs are welcome".
- slashdave 3mo agoLibreoffice
- venusenvy47 3mo agoI'm not clear what you mean with "PC office productivity software" but it seems like Microsoft Office is the winner there. Isn't that professional?
- geophile 3mo agoIt didn’t used to be.
- gitgud 3mo ago> Mainframes survive, but serving a much tinier portion of the market than they used to. I would argue mainframes rebranded to "cloud" which is ubiquitous and more people interact with this computer than any other type of device... only difference is that it's a browser instead of a terminal
- geophile 3mo agoThere is constant shifting between client and server computation. I think it is a stretch to call cloud servers “mainframes”. There are still old school mainframes, running JCL, and old school mainframe DB2 and COBOL. That ain’t cloud.
- mvkel 3mo agoDefine "winning." Open source is cheap, yet its operating systems are the least-popular. But their existence is critical to a healthy market. It's not zero sum. What you're describing is how things become commoditized, but many companies are excellent at ensuring they aren't seen as commodities
- geophile 3mo agoChoose your favorite metric: mindshare, running instances, development target. Free and low-end beat expensive proprietary. "Least popular" is an odd claim. If you count Windows desktops, and then count all the Linux installations (desktop, servers, cloud, Android, other devices), I suspect Linux would prove to be most popular. Sure, add Windows servers to the contest, that's a rounding error. And what about VMs? How do you think Windows VMs stack up against Linux VMs in number?
- caycep 3mo agoparanoid me feels like the artificial gpu/ram/ssd shortages are a plot by the VCs to forcefully reclaim central control via new age mainframes one certainly cannot buy a PC for cheap anymore
- walrus01 3mo ago> - PCs destroyed minicomputers. What's weird is that with "store your everything in the cloud and pay a monthly recurring subscription", we have now regressed to a 1960s/1970s timesharing revenue model for individual workstation computers. The default new factory out of box workflow for "enrollment" in google services, iCloud or Microsoft-everything on a new ios, macos, windows or android personal computing device is clearly designed to sign people up for subscriptions. And same general idea of "move all your servers to the cloud" recurring revenue for what is effectively the same as mainframe timesharing for key business functions, by renting VMs in GCP, Azure, AWS in perpetuity. Yes, you can still use your desktop or laptop PC in 2026 with zero external third party subscriptions (other than maybe your residential home ISP), but how many non-tech people actually do so now?
- 7952 3mo agoThe cloud era seemed to start out being about availability of storage and slowly switched in big corporates to be about security and governance. The first seems stupid now considering how much local storage we have. The later might depend on what security issues crop up.
- fhe 3mo agoApple seems a counter example, no?
- jsiepkes 3mo agoDepends how you look at it. iOS is loosing ground to Android. They might just eventually lose.
- mmooss 3mo ago> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. The parent comment cherry-picks evidence. There are plenty of counter-examples: * Office productivity suites * Search engines * Email services * Cloud services * Accounting software etc. If the LLM market ends up like search engines, one company will dominate.
- bni 3mo agoWhat about AWS, Azure, cloud computing?
- epolanski 3mo agoAt this point it's just a matter of having enough ram in your consumer computer. Until we reach a terabyte of ram at affordable prices imho this isn't going to happen.
- deleted 3mo ago[deleted]
- inigyou 3mo agoAnd yet the money is all at the high end. Bill Gates is much richer than Linus Torvalds. Oracle created one of the richest people (until he squandered it all on bad AI datacenter bets and got his company currently rated as a junk investment). Dell probably makes more money than IBM, but not by a lot.
- geophile 3mo agoThis is irrelevant to the current discussion. Windows was designed to make money for MS and billg. Linux was not designed to make money for Linux. Also, by what logic do you consider any MS software as "hign end"? That's a new one.
- QuantumGood 3mo agoCommoditization always takes volume from the marketplace, but not necessarily profit. Apple and Military tech come to mind. Positioning is rarely done well looking forward, but often shakes out in unexpected ways.
- citrin_ru 3mo agoIf LLMs is just another technology then low end will eventually win. The bet which I guess many AI investors are making is that LLM will allow recursive self improvement which will lead to development of super-human AGI. The company which will get there first will rule the world thanks to an immense power enabled by AGI. There will be no 2nd or 3rd AI company, the 1st one will wipe the rest out.