4 ms·
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
by moshun 2mo ago
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
- amelius 2mo agoNot sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
- sroussey 2mo agoOr do a hybrid
- tsujamin 2mo agoSurely that added flexibility negatively impacts the density/parameter count of the model you could etch?
- tliltocatl 2mo agoI think they already do that, except it's not 1980 so you don't fix the upper mask, you fix the lowest metal layer (the upper layer is very coarse and is only useful for power). But even a single mask is still quite expensive.
- amelius 2mo agoBut I suppose the interconnect masks don't have the resolution requirements of the masks for transistors. Therefore it could be a lot cheaper. (Yes, you could fix a number of masks, e.g. entire logic gates, of course).
- ray_v 2mo agoI could see this making sense when model development start to settle down ... it's going to settle down, right? ...
- mdp2021 2mo agoCompute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically. *(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)
- flyinglizard 2mo agoLook at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.
- alightsoul 2mo agoWhich is exactly what companies and shareholders want to increase sales.
- topspin 2mo ago"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
- mdp2021 2mo agoWell, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
- thfuran 2mo agoPhones were getting too thin anyways.
- preg_match 2mo agoYes but have we considered employing, like, a really big block of ice? Like old-timey surgeries? What if we put a big block of ice on the 2.5 cubic meter CPU what happens then?
- topspin 2mo agoYes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now. One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.
- mdp2021 2mo ago> decade after decade That can only work when there is physical capacity for improvement though. > underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So > * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance* That will yield new and renewed hardware technologies. (Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)
- breuleux 2mo agoIf you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.
- try-working 2mo agoobsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it. i have written about this: "For device makers Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications." https://try.works/role-model-the-case-for-a-model-routing-protocol https://try.works/role-model-the-case-for-a-model-routing-pr...
- nomel 2mo agoNo, the point is inference speed and power.
- try-working 2mo agoyou don't understand what I wrote.
- nomel 2mo agoI do. The point is inference speed and power, making previously impossible local inference possible. A side effect of that hardware optimization is fixed capabilities. You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?
- try-working 2mo agoWhat I'm saying is that Apple will use these type of models etched into chips, and they will do it because it drives obsolescence, so they can shorten the upgrade cycle. They will do it because they figure out it's good for them.
- zxspectrum1982 2mo agoI'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
- Gigachad 2mo agoIt costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
- zxspectrum1982 2mo agoI'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).
- mdp2021 2mo ago> It costs something like $300,000 for the hardware to run a model of that size You did not compute that as the cost for a speculative card from Taalas, right?
- Gigachad 2mo agoIt's the cost of the current nvidia hardware used to run these models. Of course all bets are off if you are accounting for some future chip that doesn't exist yet which could cost less.
- lsaferite 2mo agoGiven the fact that Taalas was claiming 1/10th the hardware cost and 1/10th the power consumption, yeah, companies would absolutely jump at that. Now, whether they can achieve that in practice is yet to be seen.
- subroutine 2mo ago
- nowittyusername 2mo agoDepends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.
- desmaraisp 2mo agoThat's definitely super-enthousiast territory. Paying 80 bucks a month for AI is more than 99.99% of people would be willing to do
- nowittyusername 2mo agoThis will be considered very cheap within the year IMO. The value you get from AI is exponentially increasing and like all tech just takes some time to ramp up. Cell phones, internet and many other amenities when they came out many people were not willing to pay for but that all changed and considering how important AI tech is this will also be the case especially considering if its 100% private such as for that cartridge.
- brailsafe 2mo ago> The value you get from AI is exponentially increasing. Perhaps in some cases, but the value I personally and professionally got out of LLMs reached a limit a while ago and has since kind of fluctuated between that limit and a bit less. If the best model was instant, like the demo here, it could certainly provide more value, I guess, but I think the limit I'd quickly hit is the same one as now, which is how much of it do I want to produce, for what reasons?
- kennywinker 2mo agoThat's because the super-enthusiast will upgrade in 4 months when a better model is released. The casual user would keep it for years. A year of claude at the lowest plan is almost $300
- Iolaum 2mo agoDo think about b2b. Companies are already paying much more for AI. a new K3 (or similar model) every 6 months for a monthly rate of ~100$ per month is something MANY businesses would pay for. Then they could even sell them at half the price to consumers.
- LarsDu88 2mo agoIt depends on how quickly you can bake new architectures. Text diffusion might be a disruptor here, but let me just say the most cutting edhe form of image diffusion (JiT and DiT) right now is just a big fat stack of alternating attention and MLP matmulls. Not theoretically hard to bake