4 ms·
Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would
by mixermachine 2mo ago
Scaling a model on a chip is quite hard.
ChatJimmy is based on Llama 3.1 8 billion.
Kimi K3 has 2.8 trillion parameters.
That are 350x more parameters.
I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.
- formerly_proven 2mo agoIt's quite telling that the 8B Taalas chip was already reticle-sized on TSMC N6. I mean, we're talking about a process that does ~100 MTr/mm², ROM needs about one transistor per bit, but can probably be packed more densely than general logic. Something like, say, 150 megabit/mm² is not a lot. N6 has a 850 mm² reticle limit. This roughly tracks, the article says the chip has 8B parameters and apparently spends about half the area on ROM. There's a reason AI accelerators just use a ton of silicon area (each HBM3 die is >1000mm² of silicon). I imagine this is not terribly viable unless they make it a lot more space efficient e.g. using MLC ROM if they don't already, or use stacked dies with a ROM-optimized process. And then we're back to not cheap, though reticle chips were never in the cheap area to begin with.
- Tuna-Fish 2mo agoTaalas exploits the low cardinality to store one 4-bit weight with one transistor. (They are using metal layer traces for the ROM, and connecting an access transistor to light up one of 16 options.) Their system is honestly very efficient for the weights, the problem is the KV-cache. That's why HC1 only supports such short context, they use SRAM for that and spend most of what's left of the die for it. The recent advancements that made attention more efficient are probably going to be very useful for them.
- jrflo 2mo agoDo you really need a KV cache if inference is that fast though?
- Tuna-Fish 2mo ago... Yes. Quadratic is really bad for large enough n, and you need that big context for useful work.
- adityazero 2mo ago[flagged]
- phonon 2mo agoA full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion parameter model. So roughly DeepSeek-V4-Pro class. Imagine that running a factory, for example.
- briansm 2mo agoCerebras have special techniques to work around etching errors / bad cores on their wafers. This is possible since their wafers are effectively hundreds of identical copies of redundant cores. Can't do that for a globally unique model. Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.
- fulafel 2mo agoThere's several ways to engineer around that as the errors are detectable. There's a big literature on how to trade off speed or transistors for error correction. [1] (Is Cerebras doing something novel? CPUs and memory blocks have been doing those things for a long time too, since the error rate is otherwise too high for normal size chips as well) [1] see eg https://www.vlsimentor.com/dft/redundancy-bisr https://www.vlsimentor.com/dft/redundancy-bisr to get some basic concepts
- phonon 2mo agoA few hundred bad bits/transistors in a trillion+ parameter model would compromise its abilities not one iota...the models are inherently lossy and resistant to "brain damage"...
- wtallis 2mo ago> (each HBM3 die is >1000mm² of silicon) Did you mean that each HBM3 stack is that large? Because it only takes one glance to see that the memory chips are much smaller than reticle-sized GPUs they sit next to.
- Closi 2mo agoScaling is definetly hard - but there is no absolute requirement to put huge flagship models into this technology (although it might be possible over time). A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in half a second rather than a few minutes. Plus some sort of slow smart + fast dumb combo architectures might also work really well for different classes of problems.
- paulryanrogers 2mo ago> Where you can request it looks at hotel options in Montreal, and it starts answering in half a second Yet the answers will get outdated quickly whilst the silicon is fixed.
- fennecbutt 2mo ago>Yet the answers will get outdated quickly whilst the silicon is fixed. Bro is living in 2020 before rag was widely introduced.
- paulryanrogers 2mo agoI was told updates require replacing at least two layers of metal, though not whole thing. Was that not accurate? Can you say more?
- mdp2021 2mo agoYou are talking about two different things. Yes, to update the blueprint for new models two layers will be updated. That is the NN. To instead update the data on which to operate you could use a RAG to query. (As in "the Pathfinder 2.0 NN is on the chip; the geodata is in the OpenGeoMaps dump-DB-nightly" - not really overlapping with LLM+RAG but may give an idea in a different scenario.)
- Closi 2mo ago
- jermaustin1 2mo agoI have found that for some personal prose-related projects, QWEN 3.6 35B A3B is an amazing model even quantized down to 4 bits. I actually find it's "writing" style as a GM for an LLM-powered solo text adventure game, better than even some of the faster/dummer frontier models like GPT-5.6-Luna or Haiku 4.5, and it runs (slowly) on a 3090 with a 80k context. So I have faith in these embedded LLM chips when it comes to fun projects like that. I have not personally found my quantized QWEN good at agentic tasks, though, and it LOVES to make shit up when asking questions about documents in the prompt.
- dpedu 2mo ago350x is only about 10-20 years of improvement, using CPU FLOPS as the benchmark.
- 8note 2mo agowouldnt you want gpu, fpga, or dsp as the benchmark? its lots of parallel calculations, rather than one blazing fast one
- paulryanrogers 2mo agoHaven't CPUs largely plateaued? They're just getting bigger, more power hungry, and multiplying cores. Physics has hard limits and Moore's law is long dead.
- girishso 2mo agoAFAIK these don't have to be CPUs. ASICS are well suited for this purpose. This will be the most cost effective way to run LLMs.
- momojo 2mo agoThere's certainly incentive to do so. And its only an engineering problem haha.
- empath75 2mo agoYou could imagine different layers on different chips, though, i think...
- cmrdporcupine 2mo agoConsider a model like https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0 https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0 which just came out. 0.9B parameters and very accurate for doing a very specific task: document features classification. Now imagine you have a chip which is just that model, but can do it at absolutely insane speed. Like tens of thousands of documents a second. Same for things like text-to-speech or speech-to-text. Think of the accessibility wins if subtitling becomes insanely accurate and fast and omnipresent. There are all sorts of domains like that, and the trend has been such that smaller models are getting smarter and smarter. If you can stick them in parking meters, traffic lights / street crossings, mobility aids, etc etc I just see so much potential win.