8 ms·
This is what worries me. People become dependent on these GenAI products that are proprietary, not transparant, and need a subscription. People build on it like
by janwillemb 5mo ago
This is what worries me. People become dependent on these GenAI products that are proprietary, not transparant, and need a subscription. People build on it like it is a solid foundation. But all of a sudden the owner just pulls the foundation from under your building.
- GaryBluto 5mo agoLuckily local AI is becoming more feasible every day.
- ModernMech 5mo agoI love how it's just a tacit understanding that these companies' entire MO is to carve out a territory, get everyone hooked on the good stuff and then jack up the price when they're addicted and captured -- literally the business plan of crack dealers, and it's just business as usual in the tech industry.
- aleqs 5mo agoIndeed, I feel like we are in the early computer equivalent phase of AI, where giant expensive hardware is still required for frontier models. In 5 years I bet there will be fully open models we'll be able to run on a few $1000 of consumer hardware with equivalent performance to opus 4.7/4.6.
- whattheheckheck 5mo agoYou'll never have the power of what they have though. Cloud capital is insane. So you can run 1 agent locally on $1k to $3k hardware They can run a fleet of thousands
- aleqs 5mo agoI think intelligence per compute will go up significantly in the coming years, while the cost per compute will drop significantly. No way to know for sure, so I guess we'll see
- nozzlegear 5mo agoBut does one individual need a fleet of thousands of agents?
- Someone1234 5mo agoIt feels more and more like OpenAI/Anthoropic aren't the future but Qwen, Kimi, or Deepseek are. You can run them locally, but that isn't really the point, it is about democratization of service providers. You can run any of them on a dozen providers with different trade-offs/offerings OR locally. They won't ever be SOTA due to money, but "last year's SOTA" when it costs 1/4 or less, may be good enough. More quantity, more flexibility, at lower edge quality. It can make sense. A 7% dumber agent TEAM Vs. a single objectively superior super-agent. That's the most exciting thing going on in that space. New workflows opening up not due to intelligence improvements but cost improvements for "good enough" intelligence.
- echelon 5mo agoOpen Source isn't even within 50% of what the SOTA models are. Benchmarks are toys, real world use is vastly different, and that's where they seriously lag. Why should anyone waste time on poorer results? I'd rather pay my $200/mo because my time matters. I'm not a poor college student anymore, and I need more return on my time. I'm not shitting on open weights here - I want open source to win. I just don't see how that's possible. It's like Photoshop vs. Gimp. Not only is the Gimp UX awful, but it didn't even offer (maybe still doesn't?) full bit depth support. For a hacker with free time, that's fine. But if my primary job function is to transform graphics in exchange for money, I'm paying for the better tool. Gimp is entirely a no-go in a professional setting. Or it's like Google Docs / Microsoft Office vs. LibreOffice. LibreOffice is still pretty trash compared to the big tools. It's not just that Google and Microsoft have more money, but their products are involved in larger scale feedback loops that refine the product much more quickly. But with weights it's even worse than bad UX. These open weights models just aren't as smart. They're not getting RLHF'd on real world data. The developers of these open weights models can game benchmarks, but the actual intelligence for real world problems is lacking. And that's unfortunately the part that actually matters. Again, to be clear: I hate this. I want open. I just don't see how it will ever be able to catch up to full-featured products.
- MostlyStable 5mo agoI think that there will come a point when open source models are "good enough" for many tasks (they probably already are for some tasks; or at least, some small number of people seem happy with them), but, as you suggest, it will likely always (for the forseeable future at least) be the case that closed SOTA models are significantly ahead of open models, and any task which can still benefit from a smarter model (which will probably always remain some large subset of tasks) will be better done on a closed model. The trick is going to be recognizing tasks which have some ceiling on what they need and which will therefore eventually be doable by open models, and those which can always be done better if you add a bit more intelligence.
- fourside 5mo agoMaybe for folks who are deep into this, but it’s not exactly accessible. I tried reading up on it a couple of months ago, but parsing through what hardware I needed, the model and how to configure it (model size vs quantization), how I’d get access to the hardware (which for decent results in coding, new hardware runs $4k-$10k last I checked)—it had a non trivial barrier of entry. I was trying to do this over a long weekend and ran out of time. I’ll have to look into it again because having the local option would be great. Edit: the replies to my comment are great examples of what I’m talking about when I say it’s hard to determine what hardware I’d need :).
- root_axis 5mo ago> new hardware runs $4k-$10k last I checked Starting closer to 40k if you want something that's practical. 10k can't run anything worthwhile for SDLC at useful speeds.
- zozbot234 5mo ago$10K should be enough to pay for a 512GB RAM machine which in combination with partial SSD offload for the remaining memory requirements should be able to run SOTA models like DS4-Pro or Kimi 2.6 at workable speed. It depends whether MoE weights have enough locality over time that the SSD offload part is ultimately a minor factor. (If you are willing to let the machine work mostly overnight/unattended, with only incidental and sporadic human intervention, you could even decrease that memory requirement a bit.)
- SwellJoe 5mo agoYou can't put "SSD offload" and "workable speed" in the same sentence.
- zozbot234 5mo agoAs a typical example DeepSeek v4-pro has 59B active params at mostly FP4 size, so it needs to "find" around 30GB worth of params in RAM per inferred token. On a 512GB total RAM machine, most of those params will actually be cached in RAM (model size on disk is around 862GB), so assuming for the sake of argument that MoE expert selection is completely random and unpredictable, around 15GB in total have to be fetched from storage per token. If MoE selection is not completely random and there's enough locality, that figure actually improves quite a bit and inference becomes quite workable.
- politelemon 5mo agoFeasibility on commodity hardware would be the true watermark. Running high end computers is the only way to get decent results at the moment, but if we can run inference on CPUs, NPUs, and GPUs on everyday hardware, the moat should disappear.
- zozbot234 5mo agoYou can already run inference on ordinary hardware but if you want workable throughput you're limited to small models, and these have very poor world-knowledge.
- andyfilms1 5mo agoSure, but local AI is still a black box. They can be influenced by training data selection, poisoning, hidden system prompts, etc. That recent Wordpress supply chain hack goes to show that the rug can still be pulled even if the software is FOSS.
- root_axis 5mo agoNot really. The hardware requirements remain indefinitely out of reach. Yes, it's possible to run tiny quantized models, but you're working with extremely small context windows and tons of hallucinations. It's fun to play with them, but they're not at all practical.
- ac29 5mo agoThe memory requirements aren't that intense. You can run useful (not frontier) models on a $2-5K machine at reasonable speeds. The capabilities of Qwen3.6 27B or 35B-A3B are dramatically better than what was available even a few months ago. Practical? Maybe not (unless you highly value privacy) because you can get better models and better performance with cheap API access or even cheaper subscriptions. As you said, this may indefinitely be the case.
- root_axis 5mo ago> The capabilities of Qwen3.6 27B or 35B-A3B are dramatically better than what was available even a few months ago. Yes, a lot better, but still terribly unreliable and far less capable than the big unquantized models.
- nozzlegear 5mo agoI've been using local AI via LM Studio ever since I canceled my Claude subscription. It's obviously slower than Claude on my M1 Studio[†], but like someone else said, I use AI more like a copilot than an autopilot. I'm pretty enthused that I can give it a small task and let it churn through it for a few minutes, while I work on something alongside – all for free with no goddamned arbitrary limits. [†] The latest Qwen 3.6 whatever has been a noticeable improvement, and I'm not even at the point where I tweak settings like sampling, temperature, etc. No idea what that stuff does, I just use the staff picks in LM Studio and customize the system prompts.
- gip 5mo agoTrue. That is why it is key important to have open source and sovereign models that will be accessible to all and always on / local. Competition (OpenAI vs Anthropic is fun to watch) and open source will get us there soon I think.
- sdevonoes 5mo agoThe sooner you cancel the sooner you become independent of them
- derektank 5mo agoYou could say the same thing about your mobile phone bill. Most people still consider the benefits of roaming access to the internet greater than the downsides of being dependent on it.
- zdragnar 5mo agoThere's very few, if any, alternatives to roaming internet access. AI tools... do what you already do, sometimes faster, sometimes worse, usually both depending on the task. There's a massive gap of necessity between them.
- wongarsu 5mo ago[dead]
- tetha 5mo agoThe owner rug-pulls, or Broadcom buys the owner and starts squeezing.
- jjfoooo4 5mo agoBut these products are all drop in replacements for each other. I've recently favored Codex more than CC, just because rate limits got mildly annoying. I really didn't have to change anything about my workflow in doing that.
- Capricorn2481 5mo ago> But these products are all drop in replacements for each other For now. That doesn't really change the risk, that just means they are all hyper competitive right this moment, and so they are comparable. If one of them becomes king of the hill, nothing stops them from silently degrading or jacking prices. The only shield is to not be dependent in the first place. That means keeping your skills sharp and being willing to pass on your knowledge to juniors, so they aren't dependent on these things. Of course, many people are building their business on huge AI scaffolding. There's nothing they can do.
- conrs 5mo agoI'm curious - why for now? This stuff is practically commoditized. Trying to think of anything that ever successfully got back into proprietary land from there.
- zozbot234 5mo agoThe thing is that AI is still more akin to a glorified autocomplete than something that can really supersede your skills. Proprietary model suppliers are constantly trying to obscure this basic underlying fact, without much success (much of the unpredictable shifts you see in proprietary AI behavior ultimately boils down to this); so it becomes far more crystal-clear when using open models that really are a pure commodity.
- conrs 5mo agoyeah, I think there's the marketing and then there's the actual true utility. AI isn't a better computer program. It's not going to be able to do everything you want autonomously. But, it's pretty good at some stuff!
- SwellJoe 5mo agoAt least some of the investors in this tech are hoping for a monopoly position. They'd like to outspend the competition to get an insurmountable lead, at which point they can set their price. But, so far, competition remains fierce. Anthropic still has the best tools for writing code. That lead is smaller than it's ever been, though. But, honestly, Opus 4.5 is when it got Good Enough. If Anthropic suddenly increased prices beyond what I'm willing to pay, any model that gives me Opus 4.5 or better performance is good enough for the vast majority of the work I do with agents. And, there are a bunch of models at that level, now maybe including some discount Chinese models. Certainly Gemini Pro 3.1 is on par with Opus 4.5. Current Codex is better than Opus 4.5 and close to Opus 4.7 (though I won't use OpenAI because I don't trust them to be the dominant player in AI). I often switch agents/models on the same project because I like tinkering with self-hosted and I like to keep an eye on the most efficient way to work...which models wastes less of my time on silly stuff. Switching is literally nothing; I run `gemini` or `copilot` or `hermes` instead of `claude`. There's simply no deep dependency on a specific model or agent. They're all trying to find ways to make unique features for people to build a dependence on, of course, but the top models are all so fucking smart you can just tell them to do whatever thing it is that you need done. That feature could probably be a skill, whatever it is, and the model can probably write the skill. Or, even better, it could be actual software, also written by the model, rather than a set of instructions for the model to interpret based on the current random seed. Currently, the only consistent moat is making the best model. Anthropic makes the best model and tools for coding, but that's a pretty shallow moat...I could live with several other models for coding. I'll gladly pay a premium for the best model and tools for coding, but I also won't be devastated if I suddenly don't have Claude Code tomorrow. Even open models I can host myself are getting very close to Good Enough.
- agumonkey 5mo agoSome people are so dependent on it they can't even say it without twisting words to hide the fact that they're now stuck at zero
- blueone 5mo agoAnthropic sells due to unrelenting pressure and unachievable demand > new owner cuts costs > models become worse > new owner sells > the capitalistic cycle wins > we, the people, suffer
- fortyseven 5mo agoThis is why, despite enjoying all of this, I really want to focus on locally hosted models. If we don't host the technology ourselves, we're setting ourselves up for a hard fall down the line. Until very recently, local models been little more than brittle toys in my experience, if you're trying to use them for coding. But lately I've been running Pi (minimal coding agent harness) with Gemma4 and Qwen3.6 and I've been blown away by how capable and fast they are compared to other models of their size. (I'm using the biggest that can fit into 24gb, not the smaller ones.) In fact, I don't really need to reach for Claude and friends much of the time (for my use cases at least).
- 2ndorderthought 5mo agoImagine if anthropic and openai went bankrupt in the next 2 years. If you look at their financials its a real possibility.
- _the_inflator 5mo ago“In the future there might be the possibility that catastrophic event A could happen.” Not the best argument. Also there is nothing without dependencies. Loose coupling means coupling.
- pmarsh 5mo agoFor the sake of argument if you build on AWS is that any more of a solid foundation? You're beholden to Amazon, unless you have the bandwidth to be able to DR immediately to another provider.
- janwillemb 5mo agoThat is indeed a similar problem. Europe is now aware of this problem and starting to mitigate it, to reduce dependencies on big tech - how hard that may be.
- notjes 5mo agoSoon, a dented toaster will be enough to run decent models.