6 ms·
I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on
by TechTechTech 2mo ago
I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device.
"Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption.
Probably this will also happen for software engineering. Some usb-powered AI accelerator with Kimi K3 (and in future even better) performance running at 10K+ tokens/sec under 50W of power purchasable for almost no cost. Need a better model? Buy the new hardware. Old hardware is probably still fine for a lot of other use-cases. I expect China to be a big player here, it fits their open-model and hardware-manufacturing strategy.
- mixermachine 2mo agoScaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.
- formerly_proven 2mo agoIt's quite telling that the 8B Taalas chip was already reticle-sized on TSMC N6. I mean, we're talking about a process that does ~100 MTr/mm², ROM needs about one transistor per bit, but can probably be packed more densely than general logic. Something like, say, 150 megabit/mm² is not a lot. N6 has a 850 mm² reticle limit. This roughly tracks, the article says the chip has 8B parameters and apparently spends about half the area on ROM. There's a reason AI accelerators just use a ton of silicon area (each HBM3 die is >1000mm² of silicon). I imagine this is not terribly viable unless they make it a lot more space efficient e.g. using MLC ROM if they don't already, or use stacked dies with a ROM-optimized process. And then we're back to not cheap, though reticle chips were never in the cheap area to begin with.
- Tuna-Fish 2mo agoTaalas exploits the low cardinality to store one 4-bit weight with one transistor. (They are using metal layer traces for the ROM, and connecting an access transistor to light up one of 16 options.) Their system is honestly very efficient for the weights, the problem is the KV-cache. That's why HC1 only supports such short context, they use SRAM for that and spend most of what's left of the die for it. The recent advancements that made attention more efficient are probably going to be very useful for them.
- jrflo 2mo agoDo you really need a KV cache if inference is that fast though?
- Tuna-Fish 2mo ago... Yes. Quadratic is really bad for large enough n, and you need that big context for useful work.
- adityazero 2mo ago[flagged]
- phonon 2mo agoA full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion parameter model. So roughly DeepSeek-V4-Pro class. Imagine that running a factory, for example.
- briansm 2mo agoCerebras have special techniques to work around etching errors / bad cores on their wafers. This is possible since their wafers are effectively hundreds of identical copies of redundant cores. Can't do that for a globally unique model. Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.
- Closi 2mo agoScaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in half a second rather than a few minutes. Plus some sort of slow smart + fast dumb combo architectures might also work really well for different classes of problems.
- paulryanrogers 2mo ago> Where you can request it looks at hotel options in Montreal, and it starts answering in half a second Yet the answers will get outdated quickly whilst the silicon is fixed.
- fennecbutt 2mo ago>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.
- paulryanrogers 2mo agoI was told updates require replacing at least two layers of metal, though not whole thing. Was that not accurate? Can you say more?
- mdp2021 2mo agoYou are talking about two different things. Yes, to update the blueprint for new models two layers will be updated. That is the NN. To instead update the data on which to operate you could use a RAG to query. (As in "the Pathfinder 2.0 NN is on the chip; the geodata is in the OpenGeoMaps dump-DB-nightly" - not really overlapping with LLM+RAG but may give an idea in a different scenario.)
- Closi 2mo ago
- jermaustin1 2mo agoI have found that for some personal prose-related projects, QWEN 3.6 35B A3B is an amazing model even quantized down to 4 bits. I actually find it's "writing" style as a GM for an LLM-powered solo text adventure game, better than even some of the faster/dummer frontier models like GPT-5.6-Luna or Haiku 4.5, and it runs (slowly) on a 3090 with a 80k context. So I have faith in these embedded LLM chips when it comes to fun projects like that. I have not personally found my quantized QWEN good at agentic tasks, though, and it LOVES to make shit up when asking questions about documents in the prompt.
- dpedu 2mo ago350x is only about 10-20 years of improvement, using CPU FLOPS as the benchmark.
- 8note 2mo agowouldnt you want gpu, fpga, or dsp as the benchmark? its lots of parallel calculations, rather than one blazing fast one
- paulryanrogers 2mo agoHaven't CPUs largely plateaued? They're just getting bigger, more power hungry, and multiplying cores. Physics has hard limits and Moore's law is long dead.
- girishso 2mo agoAFAIK these don't have to be CPUs. ASICS are well suited for this purpose. This will be the most cost effective way to run LLMs.
- momojo 2mo agoThere's certainly incentive to do so. And its only an engineering problem haha.
- empath75 2mo agoYou could imagine different layers on different chips, though, i think...
- cmrdporcupine 2mo agoConsider a model like https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0 https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0 which just came out. 0.9B parameters and very accurate for doing a very specific task: document features classification. Now imagine you have a chip which is just that model, but can do it at absolutely insane speed. Like tens of thousands of documents a second. Same for things like text-to-speech or speech-to-text. Think of the accessibility wins if subtitling becomes insanely accurate and fast and omnipresent. There are all sorts of domains like that, and the trend has been such that smaller models are getting smarter and smarter. If you can stick them in parking meters, traffic lights / street crossings, mobility aids, etc etc I just see so much potential win.
- Frannky 2mo agoYes, it also opens up a faster recurring revenue model for hardware companies, faster model obsolescence than how often you change a computer, a server, or a GPU card. I hope they can figure out trillion-parameter models rapidly. Nvidia happened to be the best option for AI after building machines for graphics, so it makes sense they weren't the best idea from scratch for this specific use case. Especially given the scale of the demand and the possibility of recurring revenue, I hope a lot of smart people will try to solve it, compete with each other, and deliver us extremely fast and cheap intelligence. And one can say LLMs are not as smart as a human, but a lot of the reasoning humans do for product and service generation isn't smart at all—it's just a bit of fuzzy input/output plus some reasoning rules. And then if you hook a robot up to the LLM, you can get results in atoms instead of bits. I'm very excited about the future. I also hope it will stop money from flowing to bureaucrats who are incentived to keep the problems open to keep the money flowing, and instead facilitate sharing directly with the people (for example, no money to the state to solve homelessness—instead, spay instead with intelligence output to build a house and provide food as part of taxes.)
- joshspankit 2mo agoI think you’re making another good point as well: Specific traces for specific inferencing will mean that some generations get deprecated. Look at H265.
- jameshart 2mo agoCommoditize your complements - still a winning strategy. If you make chips, you want models to be free.
- orangeberrytea 2mo ago[dead]
- HPsquared 2mo agoIt can ponder the meaning of its existence.
- joseda-hg 2mo ago"My Job Is To Open and Close Doors" [1] [1] - https://www.youtube.com/watch?v=49t-WWTx0RQ https://www.youtube.com/watch?v=49t-WWTx0RQ
- steve1977 2mo agoIt can insult the plumbing with "you're a dumb pipe".
- RobertDeNiro 2mo agosame way that having wifi does. by providing no actionable value, but boosting marketing materials
- imhoguy 2mo agoWiFi is fine for notification although Bluetooth would be enough to not get forgotten cloths stuck there for days. But why it calls home and why we have to create accounts to just get a notification.
- OneDeuxTriSeiGo 2mo agoAssistive technology. Imagine an energy efficient IC for a small multimodal model that can do voice to text, text to speech, question/answer, tool calling, and structured output. Wire that up to a microcontroller that parses the structured output to constrain the model (rather than giving the model direct hardware access). Now you have an assistive tech mode for supporting vision impaired users without requiring them to configure an app on their phone, pair devices, etc. And so now the user can just speak to the washing machine to tell it what to do. And because models are getting better and better at multi-language support, you can rely on a single model to cover a wide range of spoken languages. And therefore you don't need a bunch of variants of this chip for a single product line. TLDR this gives a path to replace "always online" and "wifi enabled" devices with fully on-device capabilities without being forced to abandon assistive technology support.
- bheadmaster 2mo agoIt's possible in the future we will have Rick and Morty style AI in literally everything just because it's so easy to add it. Sentient Switchblade: "Hi Beth! You've gotten taller! Shall we resume stabbing?"
- rektomatic 2mo agoNightblood? is that you?
- therealpygon 2mo agoI think this combined with a bit of memory and something like the “high-bandwidth flash” they just announced (if it works out), could be an interesting thing for some resident (burned) experts + active moe streamed from HBF. I expect one day having small lower power demand drive sized devices with proprietary burned-in models that are quite fast running on-device in robotics and such. Commoditizing LLMs via burned and locked hardware seems likely when LLMs have stabilized (when we reach a year between releases again) and the hardware is capable and “disposable” enough. “Buy a robot and upgrade it forever* (5 years) with newer models (sold separately)”, at least until the planned hardware obsolescence that the “interface has changed to support newer hardware, so you’ll need to upgrade (again) to use the latest features”. The plans basically write themselves.
- lsaferite 2mo agoThe Taalas chip already had a memory region to keep fine-tune weights.
- joe_the_user 2mo agoI like "good enough" LLMs for search and quick trivia. But "cars, appliances, etc" is exactly where LLMs are between noxious and dangerous. My food processor could use a self-cleaning feature. It could only be made worse by some system that, IDK, changes the setting based on off-hand comments about "I don't know what his beef is.."