5 ms·
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery
by anthonypasq 2mo ago
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
- bsaul 2mo agoThat's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one.
- sebular 2mo agoBut this is already happening with iPhones. Apple is touting on-device AI and only the latest phones offer the full capabilities. Newer phones will be able to run better models, so the incentive is there as soon as someone makes the killer app that only makes sense when the model is running locally on your phone.
- amelius 2mo ago> as soon as someone makes the killer app that only makes sense when the model is running locally on your phone. I expect this to be around the time when we're finally ready to travel to Mars.
- superb_dev 2mo agoFrom what I remember, these chips are not mobile size yet
- bradfa 2mo agoA small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
- mdp2021 2mo ago> A small model would be [mobile size] A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)?
- teaearlgraycold 2mo agoIs that analogue or are they baking floating points into the silicon?
- AlotOfReading 2mo agoIt's entirely possible they're using something like block floating point, where most of the hardware is simply fixed point. AMD's NPU does this, for example.
- wmf 2mo agoNope, a small model would be larger than the whole iPhone SoC.
- adgjlsfhk1 2mo agoI don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster
- teaearlgraycold 2mo agoMy question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them.
- RussianCow 2mo agoThat likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question.
- spijdar 2mo agoI don't know. As others have said, the Taalas chip wasn't small, or particularly low power, so it's hard to "imagine" what that tech in an cell phone chip might look like. But if the basic premise of "good enough LLM at insane throughput" holds, I think it could qualitatively change local uses of LLMs. At a certain speed point, you're able to move from request -> response to a cascade of tool calling and "subagents", which could allow a small model to be much more useful, if provided with a lot of local data and tool calls. That said, this is assuming you could stuff a "good enough" model into a phone with Taalas-like technology. The Taalas tech demo was an 8B parameter model and required hundreds of watts (IIRC) to run. The efficiency was good given the speed (as I understand), but it's not clear at all that the approach scales small enough to be a sensible coprocessor on an iPhone or whatever.
- intrasight 2mo agoBox that plugs into my desktop would be fine. Or perhaps in SSF form factor.
- Melatonic 2mo agoThe Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )
- chorizo 2mo agoBaking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective. Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
- adrianN 2mo agoIt is my understanding that just baking the model itself into silicon only gives moderate gains because memory bandwidth remains a bottleneck.
- chorizo 2mo agoThe big benefit is ROM cells require fewer components than DRAM. So the chips would be tiny, dense, cheap and consume far less power.
- klodolph 2mo agoI thought DRAM was pretty dense already. Is mask ROM that much denser?
- chorizo 2mo agoYes, each rom bit can be a transistor or even a diode with a decoder circuit. Simplest Dram cell is capacitor+transistor - and you need a clock, refresh circuit etc. Someday, I imagine model weights could even be encoded as analog resistors (memristors or similar) for even greater density
- whatsThisBtn4 2mo ago[flagged]
- deleted 2mo ago[deleted]
- makeitdouble 2mo agoSlightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.
- krisoft 2mo agoIt is not a “should”. At least not in the “we wish it were so” sense. It is more that there are multiple reasons why this idea (burning an LLM into silicone and deploying it into a device in people’s pockets) requires huge piles of cash and the kind of engineering chops only a few company posesses. Of course i would like it if a small upstart would do this, but it doesn’t seem likely as a posibility. They won’t have the funds to fab the IC. They won’t have the funds to train and validate the model before burning it into silicone. They can’t absorb the risk of the first tape out going wrong. They can’t absorb the risk of the model being faulty in some subtle way. They don’t have a device to integrate the IC into. They won’t have the funds to develop one. If they somehow would make a device they don’t have the marketing and sales channels built out to get the device into people’s hands in sufficient numbers to justify the development cost. Basically this idea feels ruinously expensive. Apple has deep pockets, they already have working well-regarded phones, and an ethos of privacy preserving innovation. This is why this idea feels well suited for them and not many others. Do i want the winners to keep winning? No. But not many others can pay for a moonshot crossed with a manhattan project. They just can’t.
- ricksunny 2mo agoYes, it's an interesting register (sorry for the claudism; blame lesswrong-weighted training) for the use of the word 'should'. I agree with your assessment and it is rarely articulated. Sometimes I think that the HN set is abused by big tech both from above on the employer side and the consumer usage side (all the T&C's, VC incentives and M&A taking away once-good-things). So they adopt the only sliver of agency-salving language available, like 'big company that I have no scope over should X'.
- makeitdouble 2mo ago
- freekh 2mo agoIt would be cool if the future was a standard fairphone like module system where you could replace the model chip when you felt like it without having to shell out 1-2k $$$s for a new phone
- dzhiurgis 2mo agoIts wild but if chip is something like $30 and provides frontier intelligence then just throwing them away every 3 months isn't that big of a deal when a lot of us pay $50 to $150 to $1.5k per month on AI tools. I don't think it needs to be on phone per-se. It can keep chugging in cloud - plenty of people use cheaper older models. And I suspect the growth will slow eventually making taalas interations slower.
- koiueo 2mo ago> if you could burn a gemma4 class model into an iphone ... you would still have a mediocre phone with half-assed barely working features driven by locked down proprietary software