6 ms·
If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclama
by abixb 2mo ago
If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international market if a smaller model can match SOTA of just a few months ago.
There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.
- ComplexSystems 2mo agoIt is true. I don't care about having infinite frontier-level intelligence, and I don't care if Fable can one-shot frobnicate a klaxelzorp with a benchmark performance of 97%. I doubt most people do, in fact. I just want something that meets the baseline level of intelligence needed to be a really, really good pair programming agent. It shouldn't have any silly dealbreaker issues involving laziness or hallucinations, it should be smart enough to bounce ideas off of, and it should automate doing tedious boilerplate. And - most of all - I want to be able to afford using it as much as I want. That's what has happened here.
- abixb 2mo agoI wonder when we crossed the "99 percentile of intelligence for 99% of the usecases" threshold. At this point, the gains seem to be right at the very edge of bleeding edge for narrow and specialized use cases, and wonder if it'll be a sort of diminishing return from here on.
- horsh1 2mo agoIn April
- dukeofdoom 2mo agoProbably the best counterexample is the games they are able to design. It's still mostly AI slop, few would want to play.
- jqbd 2mo agoWith those cheap models the idea is you're still in the loop anyway so the more expensive model is a waste of time and money. In this case, that means you're steering the game to look like you want not how AI wants.
- bronson 2mo agoSame with the expensive models. Fable can't one-shot a good new game. Game dev is still human-in-the-loop no matter what model you're using.
- mdp2021 2mo ago> narrow and specialized use cases Such as Decision Making. /s You just can't set a high enough threshold of intellectual effort for critical decisions.
- inciampati 2mo agoIt is amazing how fast it happened. Right now one of my main projects is fully running on DeepSeek flash. My reason was that I was blocked by both of the main US AI labs from working on it because it involves viruses. DeepSeek flash has been killing it since I switched it on, completing the first phase of the project and setting up an iteration in another application space. It isn't the most brilliant model, but it is reliable and I don't have to manage my weekly token allowance. I just spend freely and end up spending only a few dollars a day. Intelligence is going to become a basic commodity. Only special stuff is going to drive us to use special models. And maybe not even that.
- monster_truck 2mo agoSeconding this, flash and especially pro are damn near batting 1000 for me in my experiments. In one instance it figured out that the poc was running on a much slower machine and locked affinity to a single core. Many devs who have never tried from either side build it all up in their head but it's almost always been a matter of thorough tedium, which LLMs are excellent at churning through, especially when there's api docs/headers/code comments. If you have the space, try mirroring your port at the switch level and capturing every packet then making it go through them all to look for whatever. We have NSA at home lol
- FernandoTN 2mo agoI think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and you can already see it with GPT-5.6 and Fable-5, to delegate complex enough tasks that will take more than 24 hours to get done and they will not lose track. I'm not talking about a loop, but the actual intelligence to recover from these compounding errors that accumulate in dumber models. We're still a long way from the intelligence needed to let one of these agents go ahead and supervise multiple layers of sub-agents underneath to do complex orchestration. The future looks very promising and exciting. Imagine having the possibility of a Frontier model orchestrating as many sub-agents as needed that are running on cheaper models like DeepSeek.
- copperx 2mo ago> these compounding errors that accumulate in dumber models While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive. I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).
- fuck_google 2mo ago[dead]
- svachalek 2mo agoThese are not 24 hours of inference with floating point errors accumulating; largely the system guards against errors compounding. Tool failures, compile failures, test failures, etc, push back against the model taking a wrong turn and force it to correct. Yes it's much easier to have a smarter model that goes straight to the correct answer first, but it may not be necessary or economical. There's a minimum bar for the model where it understands problems and knows the right step to correct them, and above that newer models give diminishing returns.
- deleted 2mo ago[deleted]
- anal_reactor 2mo agoIt's been true for almost every business. "Cheap and good enough" usually trumps "excellent but expensive". Ikea, McDonald's, Ryanair, AliExpress, Aldi - these brands prove that catering to poor people is more profitable than catering to rich people simply because there are so many poor people that their collective spending power outweights the one of rich people.
- odo1242 2mo agoWell, not universally. It’s a tradeoff. If what you said was universally true Apple wouldn’t exist; Spirit Airlines wouldn’t be bankrupt, etc.
- tsunamifury 2mo agoApple sells to half the American population. And by definition many of them are poor. Spirit was broken by oil prices which everyone pays the same for. (There is no cheaper jet fuel alternative). Not a good comparison to the point of wrong conclusions.
- odo1242 2mo agoApple doesn’t target the low end (they are a luxury handset maker, their low end is the mid-end at best). The fact that people buy their products anyways shows that cost isn’t the only factor in business success. Spirit was broken by oil prices because they target the low end customer with their ticket prices. Oil prices went up and they had no pricing headroom to charge more on tickets so they simply went kaput. Ryanair also suffered a fair bit. Other airlines did (comparatively) fine because they had the ability to increase prices since their customers are less price sensitive. It’s a classical business lesson that being a “cost-sensitive” vs a “value-sensitive” business (what this tradeoff is called) is a tradeoff. It’s notably recommended that startups don’t target the lower end in prices since you can’t compete on cost with a business that has more economy of scale than you; you have to compete on features. And being a cost-sensitive business means that you are affected much more than other businesses by changes in material/component prices, because a 13-cent increase in the cost of a component matters more the more product you sell, and if you increase prices too much customers will start to wonder if the “budget” brand is really a good value proposition over the mid-end or high-end ones anymore.
- horsh1 2mo agoIf programming in the US to become unconditionally 10x more expensive, then the exodus from the US is about to begin.
- stockworks 2mo agoI couldn't agree more, and think of all the wasted inference across accounts overpaying for their subscriptions.. Need a secondary marketplace for this stuff.
- thenthenthen 2mo agoCould we look at manufacturing, for example the car industry, to predict what will happen?
- ngl999 2mo agoOnly to to find out they end up in a much worse place
- swiftcoder 2mo ago> Only to to find out they end up in a much worse place Like Europe?
- horsh1 2mo agoLike Shanghai
- gigatexal 2mo agoYeah this is what I’m curious about. How good are they after the benchmarks. I’ve been told yeah they’re good but they’re just building to show off for benchmarks. The ByteDance folks are apparently training a mythos level model 10T params apparently. If they do would it still be subsidized at these cheap rates?
- rcpt 2mo agoThe bet isn't that people will be able to automatically reply on bugs and rack up API charges. The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model
- palata 2mo ago> You don't win the stock market or make the deadliest drone by switching to the cheap model Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"? I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?
- qup 2mo agoThere are more than two options. What about great engineers with great models?
- palata 2mo agoAgain, I was answering to: > You don't win the stock market or make the deadliest drone by switching to the cheap model The question is not "can you win with the best model?", it is "can you not win without the best model?".
- fuck_google 2mo ago[dead]
- btbuildem 2mo ago> only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models In what kind of sad and failed dystopia is this a "saving grace"? For whom?
- swiftcoder 2mo ago> "saving grace"? For whom? For Anthropic and OpenAI, presumably. And the rather large economic distortion field around them, that may or may not go very badly for all our retirement funds if those firms become insolvent...
- ekidd 2mo ago> If what you're saying is true and accurate, then US-based AI labs are in big trouble. I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook. It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers. It would absolutely have been a frontier model last December.
- chme 2mo ago> If what you're saying is true and accurate, then US-based AI labs are in big trouble. Well... I would think that the whole AI industry in the US are working towards public bailouts... Which I guess they'll get under the current administration... So they'll be fine... Nobody there really seems interested in actually creating a profitable business anyway...
- chii 2mo ago> There's no way large companies outside the US will pay the "US AI lab" premium this has already happened with manufacturing, so it isn't surprising that other industries follow. The US premium in engineering and scientific endeavors have been lacking for the past 30-40 years, and if it werent for tech and silicon valley, the US would have nothing state of the art. Even on that front, the US is falling behind given how much effort in tech has been diverted into privacy invading, and advertising. The US has been riding momentum, but eventually that momentum will stop. It will take half a century to get back up to speed, and by then, the US will have fallen behind so far that catching back up will seem impossible. The telling evidence would be if china has the first moonbase before the US does. I think this is highly likely looking at today's US administration.
- alfiedotwtf 2mo agoLaguna S 2.1 is American and seems to be on par (or better) than DeepSeek v4, is only 110B, and it's faster! Weird there's not much talk about it vs other models.
- apatheticonion 2mo agoI've been using DeepSeek exclusively since I've been playing with LLMs, largely due to the price. I joined a company that is an Anthropic shop and I am genuinely shocked. Sonnet 5 is a little better on long horizon tasks and headless unsupervised agent workflows - but for in-IDE workflows, it's virtually unusable. I am so used to flipping around my codebase at warp speed with DeepSeek flash. It's so fast and accurate I don't have time for parallel agents. It's a really rewarding workflow. Moving to Sonnet, you ask is something simple like "split this into a seperate file" "implement this method" "this is my schema, implement a repository for it". It'll spend 30 minutes thinking and charge like $12. And no token caching, what are you even doing Anthropic? It's unusable. DeepSeek are in a league of their own
- conmod278 2mo ago> If what you're saying is true and accurate, then US-based AI labs are in big trouble. As it is now clear by behaviour of companies and US govt, all these investments will be backstopped by US govt. No US AI company will go hungry, they are national champions.
- audunw 2mo agoThese companies know their value was never in the models. There’s a scramble to acquire as much hardware as possible, and to build as many services that people get soft locked into, such that they have a moat when the open weight models fully catch up. Doesn’t matter that the models are open weight, if you want access to the hardware you will have to pay. I dont see how the outlook is any better for the open weight companies. They’re in the exact same situation as the closed weight companies except they have had much less revenue, and built up less of a brand, leading up the the point where they are equal in terms of model quality.
- reacharavindh 2mo agoSample of 1… I downgraded my Claude subscription and delegated my Claude Opus access to serve the role of an Architect to brainstorm and plan every step of development. I leave the development to Deepseek. Claude gets to review at many layers. It is often just as good as if I let Opus develop it by itself(the architect session will find similar number/level of gaps). Deepseek flash as an architect and problem solver is not as thorough as Opus5 + high. Codex sol+ high is even better than Opus 5 at this moment for my needs.
- Art9681 2mo agoIt's not necessarily true. OP has not provided actual objective figures. The nuance is in tokens spent per successful completion. So sure, you can blast DeepSeek for 1 hour solving a hard task, or Opus-5 for 5 minutes on the same task. Opus would ultimately be cheaper because time IS money and solving things faster is ultimately cheaper. Sure, benchmarks will show cost to be lower but people who test all of these models in real world use cases know the benchmarks are not reflective of real world use cases and Chinese models choke on problems Fable or gpt-sol will breeze through.
- budsniffer952 2mo agoOne random says they are using DeepSeek and you proclaim it's all over for the top AI labs. Brilliant. First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses? Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American labs are SOTA. That might change, but let's not act like the American labs don't know what they are doing. The idea that people are going to use cheaper models for cheaper work isn't some novel revelation, it's completely obvious. People are doing it already, eschewing Fable.
- edot 2mo agoSubsidy doesn't matter once the weights are released. You're conflating subsidized training and research with subsidized inference. DeepSeek's (allegedly government-) subsidized inference is coming to an end soon per their docs, but that's fine because the weights are open. You can run this on-prem with very modest hardware. And you can fine-tune it to your business with your data.
- budsniffer952 2mo agoThen why are the major American labs in "deep trouble"? They can also continue to be subsidized and release the weights if they want to, if that's what you think it takes for survival. Seems like a low bar. The point is it takes money to keep developing models. Everyone is playing by the same rules. At this point, the US labs are trying to build businesses. I'm not really sure what the Chinese labs goals are. But I do know they aren't doing charity work.
- irpap 2mo agoThey are in trouble because they offer something 5% better for 2000% the price.
- FpUser 2mo ago>"To what extent are the Chinese labs being subsidized? Are they real businesses?" And who gives a flying fuck. I am a "real" business and I count my money. It is not my life goal to prop some fat cats crying crocodile tears. Granted I do not use Chinese models. I use Junie straight from my JetBrain's IDEs that in turn uses Gemini Flash. Very cheap and more than enough for my use.
- ExxKA 2mo agoIt is true. I have been using it since it got released and its as good for scoped coding tasks, as the other big models I use, but just soo much cheaper.
- taf2 2mo agoAs a business though I can spend 100k and run this model and support about 10 fte- not bad…