6 ms·
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Ba
by LarsDu88 2mo ago
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.
Baking models onto silicon would've been the next logical move to get a moat.
Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
- LPisGood 2mo agoI’m surprised Nvidia hasn’t partnered to make a Claude chip yet. It’s a win/win you can license them out, sell them when they become obsolete, etc.
- moshun 2mo agoConsidering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
- amelius 2mo agoNot sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
- sroussey 2mo agoOr do a hybrid
- tsujamin 2mo agoSurely that added flexibility negatively impacts the density/parameter count of the model you could etch?
- tliltocatl 2mo agoI think they already do that, except it's not 1980 so you don't fix the upper mask, you fix the lowest metal layer (the upper layer is very coarse and is only useful for power). But even a single mask is still quite expensive.
- amelius 2mo agoBut I suppose the interconnect masks don't have the resolution requirements of the masks for transistors. Therefore it could be a lot cheaper. (Yes, you could fix a number of masks, e.g. entire logic gates, of course).
- ray_v 2mo agoI could see this making sense when model development start to settle down ... it's going to settle down, right? ...
- mdp2021 2mo agoCompute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically. *(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)
- flyinglizard 2mo agoLook at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.
- alightsoul 2mo agoWhich is exactly what companies and shareholders want to increase sales.
- topspin 2mo ago"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
- mdp2021 2mo agoWell, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
- thfuran 2mo agoPhones were getting too thin anyways.
- preg_match 2mo agoYes but have we considered employing, like, a really big block of ice? Like old-timey surgeries? What if we put a big block of ice on the 2.5 cubic meter CPU what happens then?
- topspin 2mo agoYes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now. One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.
- mdp2021 2mo ago> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance* That will yield new and renewed hardware technologies. (Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)
- breuleux 2mo agoIf you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.
- try-working 2mo agoobsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it. i have written about this: "For device makers Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications." https://try.works/role-model-the-case-for-a-model-routing-protocol https://try.works/role-model-the-case-for-a-model-routing-pr...
- nomel 2mo agoNo, the point is inference speed and power.
- try-working 2mo agoyou don't understand what I wrote.
- nomel 2mo agoI do. The point is inference speed and power, making previously impossible local inference possible. A side effect of that hardware optimization is fixed capabilities. You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?
- try-working 2mo agoWhat I'm saying is that Apple will use these type of models etched into chips, and they will do it because it drives obsolescence, so they can shorten the upgrade cycle. They will do it because they figure out it's good for them.
- zxspectrum1982 2mo agoI'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
- Gigachad 2mo agoIt costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
- zxspectrum1982 2mo agoI'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).
- mdp2021 2mo ago> It costs something like $300,000 for the hardware to run a model of that size You did not compute that as the cost for a speculative card from Taalas, right?
- Gigachad 2mo agoIt's the cost of the current nvidia hardware used to run these models. Of course all bets are off if you are accounting for some future chip that doesn't exist yet which could cost less.
- lsaferite 2mo agoGiven the fact that Taalas was claiming 1/10th the hardware cost and 1/10th the power consumption, yeah, companies would absolutely jump at that. Now, whether they can achieve that in practice is yet to be seen.
- subroutine 2mo ago
- nowittyusername 2mo agoDepends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.
- desmaraisp 2mo agoThat's definitely super-enthousiast territory. Paying 80 bucks a month for AI is more than 99.99% of people would be willing to do
- nowittyusername 2mo agoThis will be considered very cheap within the year IMO. The value you get from AI is exponentially increasing and like all tech just takes some time to ramp up. Cell phones, internet and many other amenities when they came out many people were not willing to pay for but that all changed and considering how important AI tech is this will also be the case especially considering if its 100% private such as for that cartridge.
- brailsafe 2mo ago> The value you get from AI is exponentially increasing. Perhaps in some cases, but the value I personally and professionally got out of LLMs reached a limit a while ago and has since kind of fluctuated between that limit and a bit less. If the best model was instant, like the demo here, it could certainly provide more value, I guess, but I think the limit I'd quickly hit is the same one as now, which is how much of it do I want to produce, for what reasons?
- kennywinker 2mo agoThat's because the super-enthusiast will upgrade in 4 months when a better model is released. The casual user would keep it for years. A year of claude at the lowest plan is almost $300
- Iolaum 2mo agoDo think about b2b. Companies are already paying much more for AI. a new K3 (or similar model) every 6 months for a monthly rate of ~100$ per month is something MANY businesses would pay for. Then they could even sell them at half the price to consumers.
- LarsDu88 2mo agoIt depends on how quickly you can bake new architectures. Text diffusion might be a disruptor here, but let me just say the most cutting edhe form of image diffusion (JiT and DiT) right now is just a big fat stack of alternating attention and MLP matmulls. Not theoretically hard to bake
- bamboozled 2mo agoIt googles models suck
- anthonypasq 2mo agoPersonally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
- bsaul 2mo agoThat's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one.
- sebular 2mo agoBut this is already happening with iPhones. Apple is touting on-device AI and only the latest phones offer the full capabilities. Newer phones will be able to run better models, so the incentive is there as soon as someone makes the killer app that only makes sense when the model is running locally on your phone.
- amelius 2mo ago> as soon as someone makes the killer app that only makes sense when the model is running locally on your phone. I expect this to be around the time when we're finally ready to travel to Mars.
- superb_dev 2mo agoFrom what I remember, these chips are not mobile size yet
- bradfa 2mo agoA small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
- wolttam 2mo agoIt's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.
- speed_spread 2mo agoIf a model is good enough today, it's still gonna be good enough in a year. Except you'll be able to serve it 1/100 of the price. Or 100x the speed.
- wolttam 2mo agoI think we will eventually reach a point where this is the case, but at the moment it seems like you can throw virtually any non-trivial use-case at a model today and end up being more satisfied with the results that a model tomorrow gives. I may just be closed minded as to what use-cases we have that current models are truly "good enough" (i.e. won't be dissatisfied when comparing results of today's model to tomorrow's model)
- nine_k 2mo agoNot so if it's embedded in something smart enough for its intended purpose. Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second.
- deleted 2mo ago[deleted]
- anigbrowl 2mo agoThis is only true for people who are solely focused on performance. There is absolutely a market for acceptable performance combined with predictability.
- teraflop 2mo agoTrue, but predictability cuts both ways. We're all used to having to constantly update our browsers and phones to keep up with the security arms race. If a frozen model can't be updated, it will predictably remain vulnerable to any "exploits" or idiosyncratic quirks that people discover over time. Let's say, as somebody suggested in another comment, that you buy 100,000 of these chips and deploy them to run fast-food drive-thrus. And then somebody discovers the model has a fondness for goblins[1], and if you role-play convincingly enough, you can get it to accept payment in shiny buttons and rodent skulls instead of cash. What do you do then? I guess your options are to try and fix the behavior with a better prompt, or put some kind of filter in front of the model to catch attempted exploits. If the filter is cheap and dumb it probably won't work well enough, and if you use another model as a filter, you've negated the cost and speed benefits of putting the first model in hardware. Of course the real answer is to just never expose the model to situations where an adversarial input could possibly lead to an undesired output. But that drastically limits what you can do with it. [1]: https://openai.com/index/where-the-goblins-came-from/ https://openai.com/index/where-the-goblins-came-from/
- karmasimida 2mo agoA model can't be updated, and a chip that is only relevant for 6 months at max?
- askvictor 2mo agoPeople already buy new phones every year, this just creates even more reason to do so
- throwaway240403 2mo agoYour location/income bias is showing. Most people do not buy new phones every year.
- winrid 2mo agoI live in the bay area and buy a phone maybe every 3 years? Why do people waste so much money :D
- askvictor 2mo agoI never said most people. But it's not uncommon. I personally find it ridiculous, and hold onto my phone until it's unusable, but plenty of people in middle class Australia seem convinced that they need the new one whenever it comes out.
- Gigachad 2mo agoOutside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people.
- boelboel 2mo agoCloser to every 5-6 years these days and with ram prices going up it will be even longer. Especially with the low/mid range phones, which are most phones outside some developed countries, people will keep their phones as long as they can.
- mrtksn 2mo agoIsn’t that kind of useless for the stock? It sounds complicated, unlike having number of CPUs go up. It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom.
- alightsoul 2mo agoBecause Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.
- wmf 2mo agoOpenAI and Anthropic are both designing ASICs.
- alightsoul 2mo agoSo they have decided that putting a small LLM on a phone would backfire because people would have a negative perception of their cloud models. Pretty sure AMD will use these taalas chips in data centers, not phones
- giancarlostoro 2mo agoASICs is what took over Bitcoin mining, cheaper in all ways, and lasts longer than Nvidia GPUs for inference.
- SR2Z 2mo ago> cheaper in all ways, Bitcoin mining doesn't have large memory requirements, but does have huge compute requirements. ASICs work great there because it's very straightforward to add some circuits for computing hashes. If you _also_ have to add many GB of memory, then suddenly ASICs will cost as much or more than comparable off-the-shelf hardware and they won't be faster unless you've also invested in huge memory bandwidth.
- giancarlostoro 2mo agoMy understanding is an ASIC can last 10+ years, where are Nvidia enterprise GPUs are rated for 5...
- SR2Z 2mo agoMost enterprise GPUs are scrap after 5 years because they're so inefficient compared to newer models. It's entirely possible to make them last longer by undervolting them, people just don't because it doesn't make sense. Bitcoin OTOH has used the same PoW algorithm for a decade. Barring some really exciting discoveries about the nature of computation, new ASICs are not that much more efficient than old ones. BTC mining is also not exactly competitive anymore; the nature of the PoW algorithm means that it's dominated by a few large players who've set up shop next to a dam and who pay very little for electricity. New entrants are highly discouraged because the mining rewards are constantly halving, it's hard to find cheap power, and the price of BTC is now so volatile that a yearslong investment is very likely to lose money.
- CircuitSeuss 2mo agoApparently Anthropic is moving that way: https://arstechnica.com/ai/2026/08/anthropic-confirms-plans-to-build-an-in-house-silicon-team/ https://arstechnica.com/ai/2026/08/anthropic-confirms-plans-...
- mdp2021 2mo agoNot necessarily: it is relevant to Taalas only if it is a compute-in-memory architecture. The Jalapeño mentioned («Anthropic is not alone in walking this path») in the article is still a classical Von Neumann architecture. And Taalas' idea makes sense in a perspective of scale - producing a large number of cards; "for internal use" (a lower order of items) means a high production cost.
- throwaway27448 2mo agoYou need to find customers for several-generations-ago models before this makes any sense. AMD is a lot more incentivized to look than mr vanilla llm is
- la6479 2mo agoJust to see how fast it is try chatjimmy.ai
- mr_mph 2mo agoPretty incredible to see. It reminds me of when I first used the Groq chatbot, except in this case it's a full response instantly.
- tasty_freeze 2mo agoIt is really fast and ... really hallucinates. I asked "Does the Wang corporation still exist? If not, what happened to it?" and it replied (in part): "Yes, the Wang Corporation, the company that originally developed and marketed the Wang 2200 computer, still exists as a rebranded company under the name PPL (Precision Pencil and Label), but it has undergone significant changes and challenges over the years. Here's a brief overview of what happened: Founding and Growth: The Wang Corporation was founded by An Wang in 1969." In fact, Wang labs was founded in 1951. PPL seems to be a made up entity. But it did generate those "facts" in 0.033 seconds. If people value speed over accuracy then I can write an LLM that is 100x faster than chatjimmy.ai and make big bucks by responding one of N canned responses to any question.
- mickaelkerjean 2mo agotheir tech is a mere demo to open up a new path, the day we can have some asics running a Qwen3.6 27b, this would open up new doors
- jjcm 2mo agoto be fair, it's running an 8b model from like 2 years ago. Taalas just does the chip design, not the model architecture.
- UncleOxidant 2mo agoI guess I'm not understanding why this makes sense for AMD to buy Taalas unless they plan to get into hosting. It doesn't seem like a great fit.
- stingraycharles 2mo agoDidn’t Anthropic acquire Cerebras? Seems like a move into the same direction. I also think that etching models into ASICs may be a bit too inflexible for what OpenAI and Anthropic want.
- wyrdcurt 2mo agoNo, that's backwards. OpenAI are the ones investing in Cerebras. Part of the deal is that they can't sell to Anthropic.
- wraptile 2mo agoThis seems like a very bad and dangerous direction for our society.
- unsigner 2mo agoTheir thing is improving the models; it would be extremely counter-company-culture to bet on models plateau-ing. Maybe wise in terms of hedging, but still difficult to pull of as a company decision.
- Haven880 2mo agoChinese already start making DUV which can do the lower end 7nm. They are winning. Once that 7nm and up market cornered by Chinese, AMD Intel and TSMC and Samsung will have to burn thru bleeding edge depreciation faster perhaps from 7yr down to just 18mths. The CPU they generated will be incredibly expensive. Meanwhile Chinese just keep minting the AI cheaply and more efficiently and inching upwards towards 1.4nm.
- petra 2mo agoThey can do 7nm. But they use multiple patterning(printing the same pattern multiple times to get to 7nm), which is expensive. So it's not comparable on cost to western single-patterning 7nm, and of course not to the leading edge on cost/size/power.
- planb 2mo agoThey are: https://openai.com/index/cerebras-partnership/ https://openai.com/index/cerebras-partnership/ My guess is they only consider Luna "good enough" to justify the immense up-front investment to put it onto silicon, but Luna at 10x the current speed would be killer. If they're really pursuing live voice conversations with a hardware assistant, latency is more important than accuracy (for complex questions the assistant could always say something like "wait a minute, I need to think about this" and hand over to another model).
- ldng 2mo agoHow ? Do LLMs actually "know' when they don't "know" ?
- Certhas 2mo agoHow do humans?
- dgellow 2mo agoAlways the same trick of not answering the question and deflecting to „what about humans“. Can you folks not evaluate LLMs as the system they are, without vague gestures at how a different system behaves?
- Certhas 2mo agoEvaluating LLMs is incredibly difficult. They are categorically different from any other system we have intuition about. That said, I read the question I am replying to as a rhetorical one. If it was meant as a genuine question, curious about the question of meta knowledge, then I misread. Certainly the question is extremely interesting, for both LLMs and humans! But it's also obviously a very difficult one, as we don't even have a clear theory on how "knowing" works in the base case.
- planb 2mo agoThey do this all the time, I'm using ChatGPT in Instant mode and it auto updates to thinking if my question is complex. Most of the time this works. To answer your question: A large language model itself does not know this (afaik). But chatbots are not "just LLMs" but a whole bunch of systems (and models) around them.
- larodi 2mo agoWe don’t really known (from the outside) how exactly do they move. Besides it may have not been truly viable 1-2 years ago…
- elAhmo 2mo agoThey were busy buying open source frameworks and teams behind those.
- bjackman 2mo agoDwarkesh recently pointed out [0] that these guys are almost forced to spend most of their compute on training instead of inference. This is because they need to maintain the appearance (which may also be the truth) that future models will make current models obsolete and be much more valuable. Completely fixed-function HW can't be used for training, it's inherently a statement that "this model is Good Enough and we are now gonna start just extracting its value instead of extending it". So yeah it's an inference moat but it's not a growth moat. Makes perfect sense for a company trying to get into the compute business, not companies who wanna be in the creating-ASI business. Still, I guess/hope they have teams doing it in-house anyway. Just not something they'd wanna make a huge amount of noise about, it doesn't look good for To The Moon valuations. [0] https://www.dwarkesh.com/p/why-compute-might-get-10x-more-expensive https://www.dwarkesh.com/p/why-compute-might-get-10x-more-ex...
- troyvit 2mo agoI wish we lived in a reality where Framework was anywhere near rich enough to acquire them. I'd love to have models on a chip that I could swap in at a whim. That would do the opposite by eliminating moats. I guess there's a tiny chance AMD makes something like that happen. It seems like a great way to get people and orgs to pay a few hundred bucks every 6 months or so.
- vonneumannstan 2mo ago>Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference. Google is no longer a serious player in frontier AI. I doubt they will ever hit a SOTA model again.
- adityazero 2mo agoThey are busy capturing market first, and I think that makes more sense. 'premature optimization etc.'