4 ms·
I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips th
by mchusma 1mo ago
I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves.
For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.
I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
- sebzim4500 1mo agoMy guess is we only see this once they start saturating computer use benchmarks. That's a use case which would be extremely valuable at the right costs/speed, but the current models just aren't there yet.
- bmulholland 1mo agoProbably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a month, start to finish, for the physical processing. Maybe once LLM improvements asymptote further?
- kushie 1mo agotapeout could shrink but days per mask layer (DPML) does not have much margin..
- smallmancontrov 1mo agoI'm not in industry, is DPML (which I assume is the time required to make a mask?) set by electron beam scan time or something?
- vineyardmike 1mo agoHow much of that 16mo is design versus just production? If there was a “plug and play” chip where you just BYO weights, how long would it take? The bigger issue seems to be that these chips can’t hold that many weights at the moment. (I’m curious if chips with large weights in them would be more tolerant or less to yield issues. If you flip a few bits in the weights, does it really matter at scale?)
- RealityVoid 1mo agoTalaas, from what I understand is building stuff just like that. The infra is the same and the weights layer is all you need to change. I guess you could half etch the chips and then finish them with the weights only. I think their turnaround is 6-8 Weeks. The size of the models fitting on the chips at the moment is llama 3 I think?
- derefr 1mo ago> I guess you could half etch the chips and then finish them with the weights only. Basically a https://en.wikipedia.org/wiki/Gate_array https://en.wikipedia.org/wiki/Gate_array. (The non-field-programmable kind.)
- jeremyjh 1mo agoI think Sol is already good enough though.
- basilgohar 1mo ago"640k (token context) should be enough for anyone."
- jerf 1mo agoI know what you're saying, but modulo things like losing track of what year it is as time passes by, a current frontier model is going to continue to be useful for many tasks for many years, even moreso if it's 5-10x faster due to the chip architecture. It's not that it would be the best forever, it's that it would be useful for plenty long enough to be worthwhile, even if there was better stuff available. In exactly the same way that this computer I'm typing this message on is not the latest and hottest cutting edge stuff. A 7 year old CPU, 7 year old Intel integrated graphics, an older NVMe disk, a mere 32GB of RAM... ok, that's one spec that's still pretty modern although it is slower RAM... but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge.
- throwuxiytayq 1mo ago> but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge While it’s still too early to tell, I don’t think that’s how intelligence scales. Better models get you better solutions even to trivial problems. The ceiling for getting it done better is very high even if you’re not doing anything complicated. And difficulty isn’t uniformly distributed anyway - it seems to me that “mostly simple” tasks often have annoying 1% tails that low-intelligence models struggle with. I think we’ll see people chasing the top models for quite a while, or indefinitely - depending on the cost curve.
- dgently7 1mo agoexactly, but the "goes out of date" is bad when we talk about software.. but this isnt software, its hardware. the youd have to buy a new one to get a better model is a FEATURE not a bug. like if im apple... and i can put a sol level llm in an iphone, market it as privacy first you own your data personal assistant, integrate it all over the os... and then when there is a better model/siri make all the users buy a new phone... thats how they "win" ai. the old standbys of better screens thinner cameras and batteries arent enough anymore. its basically tapped out. all modern phones are as thin as they need as big as they need as fast as they need and last all day on a battery... apple needs a new number to up thing that people can actually feel/see. model generations could be it... every year faster, smarter, more capbilities and integrations.
- kurthr 1mo agoThe metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-extraordinary-token-speed/ https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-insane-ai-inference-performance/ https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even from 6nm to 3-4nm one could embed a 20-30B model as part of a MoE (or only a subset of activated layers) on a single reticle die (note B300s are already multi-reticle), with a separate predictive/dispatch model controlling them each on a separate chip. This is without even stacking CiM ROM die. Moving the layer activations (and KV cache etc) between die requires relatively high speeds (and low latency), but distributed with multiple die in parallel might well be doable even with standard multilane PCIe. Of course KV cache prefill could also be handled by external GPUs. I'm sure AMD will make some reasonable choices.
- MBCook 1mo agoBut that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.
- geysersam 1mo agoWhy would it be useless in 3 months?
- imtringued 1mo agoBecause on hacker news the only thing that matters is being in the current news cycle and not whether your business is profitable.
- Certhas 1mo ago
- thoughtbefore 1mo agoIt may not matter. Think about why SOTA model companies are exploring chips. What do chips offer? If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.
- calebkaiser 1mo agoGoogle rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020. If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
- edgyquant 1mo agoNeither of those companies core business model was serving llms
- calebkaiser 1mo agoWhat? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute https://www.anthropic.com/news/anthropic-amazon-compute Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.
- edgyquant 1mo agoNotice how I used the words “core business” but that they didn’t do business at all
- tintor 1mo agoThey could etch the model architecture, without the weights into the chip. This way newly post-trained model can be loaded and served the same day.
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- nerdsniper 1mo ago> Maybe once LLM improvements asymptote further? Maybe! But it also doesn't require the rate of improvement to slow down. As long as some current model is eventually "good enough" for general use, it could still be a market-killer at a very low marginal price thanks to ASIC. Even if slower, much more expensive models are 10x better, that doesn't actually diminish the utility of the ASIC model, as long as it's "good enough".
- lelanthran 1mo ago> Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). Yeah, with any luck it would put pressure on Nvidia to charge less, and not just to OpenAI. With a little more luck, we would see all the other players do the same thing, driving down the price of actual GPUs from GPU manufacturers.
- thesz 1mo agoI made some analysis half a year ago: https://news.ycombinator.com/item?id=47109252 https://news.ycombinator.com/item?id=47109252 It appears that to have working ASIC with the LLM baked into it we need to place and route macroblocks, and not a great variety of them. These macroblocks can be pre-placed-and-routed, available as masks already and shared between different LLMs. Thus it appears that the tapeout delay can be substantially lower than a year.
- htrp 1mo agoetched tried this.... it didn't go very well
- anukin 1mo agoI would assume asic based llm would work really well. Why did it not go well?
- striking 1mo agohttps://chatjimmy.ai/ https://chatjimmy.ai/ runs Llama 3.1-8B on an ASIC as a demo by https://taalas.com/ https://taalas.com/ I believe. That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
- mchusma 1mo agoYou are correct. I think this is the bull case. It seems like this would be useful right now for some things (eg moderation).
- fl0id 1mo agoisn't that what they are doing with cerebras?
- mkl 1mo agoNo, Cerebras holds the weights in SRAM - they are changeable, not baked in.
- andy_ppp 1mo agoYes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes.
- porphyra 1mo agoAlso right now Sol 5.6 Max is super slow but if it were way faster on a chip (like Taalas' Llama 8b demo) then it would be an extreme value multiplier. But the model is so large that "baking it onto a chip" doesn't seem straightforward.
- Caracas288 1mo agoMan wouldn’t it be cool to be able to slot a massive ROM AI chip into the external AI drive of the pc…
- xyzsparetimexyz 1mo agoIt'd just be pcie probably
- pantelisk 1mo agoIt should look like a NES cartridge! That you have to blow on its end to clear any dust and it should do a satisfying click when it slots in. Cooling might be an issue though...
- redox99 1mo agoThat'd be ungodly expensive.
- structural 1mo agoKeep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts. Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower. It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it.
- Aurornis 1mo ago> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon. Taalas needed a giant chip (6nm) for an 8B model. At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.
- greenknight 1mo agoNope. But we are hitting some pretty impressive levels with 128B models. The other thing is, a lot of the time, model performance is improved with more 'thinking' time. The thinking time is just more tokens... but instead of say 1000 tokens or 10,000 tokens worth of thinking its 1,000,000... how does that improve model performance? Could a 128B model hit levels of GPT Sol?
- cherioo 1mo agoThinking generates a ton of tokens. These baked in chips tend to not have a lot of memory for context. I am not sure taalas supports Thinking at all. The more problem like these they solve the more they will look like GPU.
- nextaccountic 1mo agocouldn't one just add some hundreds of GB of HBM?
- kimixa 1mo agoYeah, but then there's the size of KV cache needing to be read through that HBM interface for each token, putting a hard limit on the tok/s based on the memory bandwidth. On some models a large context can be a notable proportion of the size of the weights themselves. For example, qwen 3.8 27b uses ~64kb/token for the kv cache - so for a 256k token context that's ~16gb of the kv cache for a ~54gb model (assuming 2 bytes-per-param/f16 for both). So if the current non-baked-in chip is already memory bandwidth bound, as is often the case for current hardware and models, and the "only KV cache in HBM" chip has the same total memory bandwidth, it can only ever be (54/16)=~3.4x faster for the baked in-silicon model. EDIT: I guess actually (54+16)/16=~4.3x faster, as the current implementation would need to read that KV cache too :)
- mf_tomb 1mo ago"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal
- twobitshifter 1mo agoIt depends when the good enough level hits. Pretty sure we are almost there for most common applications of AI.
- dgacmu 1mo agoThat's only half the problem. OpenAI is contractually obligated, if you will, to believe that models will continue improving at an impressive rate for the foreseeable future (otherwise their valuation makes no sense). If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to end up serving that Sol-class model several times more expensively than it will take someone who comes along a few months later. (such as what recently happened with DeepSeek's update.) And under that assumption of continuing advancement, baking things in doesn't make sense in general - it's a play you'd make if you think things are slowing down a lot. Which may be right but it's not OpenAI or anthropic's play.
- adventured 1mo agoAssume a $800 billion valuation. $100 billion ad network. $30 billion op income. 26x price to op income ratio. It's right there for them to grab, or someone else to grab. Their valuation does make sense if you believe: 1) they can retain a massive user base and 2) a massive user base can be monetized. Future value is almost always pulled forward these days for high growth tech companies. An LLM the size of Google search in users is even more valuable than Google search. The ad market for LLMs will be even larger than search was (no matter what HN prefers). The monetization part is the easier part. Silicon Valley understands extraordinarily well how to build ad networks. If OpenAI maintain their gigantic user base, a $100 billion ad network is a given bolt-on. They'd have to screw that up in an epic way to not get there. Facebook - Insta - WhatsApp is an absolute dogshit tandem with a gigantic user base. $228 billion in ad sales and still expanding 10% per year. Google knows this is what's happening, that's why they don't care about chasing Anthropic very much. They're busy completely remaking how their core search business works.
- andsoitis 1mo ago> For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. but you trade updatability, which I don't think is worth it yet.
- mchusma 1mo agoMaybe! (1) Would SOL level intelligence be useful 3 years from now? 5 years? (2) would dedicated chips be the most affordable way to run this model in 3-5 years? I suspect the answer to both of these questions is yes right now, but I agree it’s borderline.
- andsoitis 1mo ago3 years is an eternity.
- imtringued 1mo agoLoRAs are a parameter efficient way to update models. The fine tuning stopgap only has to work long enough to extend the lifespan of a model until the next chip comes out.
- lqstuart 1mo agoEventually, someone is going to do this in Minecraft
- bigfishrunning 1mo agoLol i'm surprised this hasn't happened already. Minecraft is a great platform for scripting as long as you think every single other platform is too fast and convenient.
- raincole 1mo agoIt won't happen until IPO. If they do it now it'd be signaling that AI isn't improving fast.
- freakynit 1mo agoIn case anyone's interested in these niche startups like taalas, here are a few more: 1. https://matx.com/ https://matx.com/ 2. https://www.d-matrix.ai/ https://www.d-matrix.ai/ 3. https://www.etched.com/ https://www.etched.com/ 4. https://www.positron.ai/ https://www.positron.ai/ 5. https://hyperaccel.ai/ https://hyperaccel.ai/ 6. https://axelera.ai/ https://axelera.ai/ 7. https://www.enchargeai.com/ https://www.enchargeai.com/ 8. https://furiosa.ai/ https://furiosa.ai/
- deleted 1mo ago[deleted]
- grackasthebig 1mo agoYeah this feels like Altman not knowing engineering well enough to realize where the focus should be. Like he is optimizing to keep providing a vanilla token factory when weighted chips are coming and local models will supplement. My head canon is savvy chip execs will be etching architecture his OpenAI pioneered into their flagship products while trying to minimize how much foothold he can get in hardware. Murica done offshored it. Not ours to control.
- miki123211 1mo agoI really want this for Whisper, particularly in some kind of a power-efficient, portable form factor. I think people underestimate how much of a revolution having an always-on, privacy-preserving personal notetaker / secretary would be.
- mike_hearn 1mo agoAI is bigger than just LLMs. The real place model-hardwired chips will find value is in robotics. There you need very local, low latency inferencing with relatively stable models to handle motor control and navigation tasks. Higher level reasoning can be delegated to LLMs and run asynchronously.
- walthamstow 1mo agoI don't think any underestimates how powerful it would be. It's the social aspect that is difficult. I wouldn't want to be sat in a pub with you and your always-on notetaker.
- fooker 1mo agoComputer Science history has taught us that this decision was almost always wrong. Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.
- aleph_minus_one 1mo ago> Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant. This often happened, though not always. An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned into modern GPUs. Even today, GPUs live on as separate hardware in a computer instead of having become integrated into, say, the CPU. Intel's attempt to do something like this with the Larrabee project [1] was discontinued. --- [1] https://en.wikipedia.org/w/index.php?title=Larrabee_(microarchitecture)&oldid=1368694949 https://en.wikipedia.org/w/index.php?title=Larrabee_(microar...
- fooker 1mo agoIt didn't happen for the original fixed function graphics though, CPUs kept eating their lunch every year or so with fun rendering techniques. We still use some of these techniques for 2d graphics. Only when they became more general with shaders, and then added support for GPGPU, did it truly take off. I think this generality is the lesson here, not the fact that GPUs are not CPUs.
- meowers1 1mo ago[dead]
- nkotecha6 1mo ago[dead]