3 ms·
What hardware advances would we need to see for that to happen? It feels like everything in that arena has kind of plateaued.
by Mistletoe 2mo ago
What hardware advances would we need to see for that to happen? It feels like everything in that arena has kind of plateaued.
- sudo_cowsay 2mo agoIt could be on software side too. OpenAI has certainly not plateaued.
- bobbylarrybobby 2mo agoThe models themselves have far from plateaued. Maybe someone finds a way to get a really capable model down to, say, 12GB of ram. Then we'd be in business.
- swiftcoder 2mo agoAgreed. We've just seen DeepSeek post-train their ~300 billion parameter flash model to outperform their 1.6 trillion parameter pro model, in the space of a few months. There would seem to still be quite a few opportunities on the table to bring big model smarts down to the smaller models
- deleted 2mo ago[deleted]
- CircuitSeuss 2mo agoA lot of this will come from co-optimizing hardware and low level machine code for this specific use case… something apple is coincidently very good at. Apple has worked very hard to make unified memory a feasible approach, and the benefits of that are pretty clear in apple silicon- that efficiency not only results in power and therefore thermal gains, but also in a significantly faster full loop per process: or a faster time to token. This is why even their single core mobile chips in the budget line Neo out perform PC processors with several times more threads and RAM[1]. Turns out, unified memory lets you have a whole lot more control over things like RAM bussing and core use for specific workflows. Speculatively, a unified memory approach could also allow you to more easily integrate things like ReRAM to solve the current memory swapping bottleneck. Let’s say a friend of mine works hardware at apple and works on exactly this… on device processing is the future I’m betting on. [1] https://youtu.be/x26A28DoT-w?t=605 https://youtu.be/x26A28DoT-w?t=605
- jsjohnst 2mo ago> This is why even their single core mobile chips in the budget line Neo That’s a six core processor. It’s an A18 Pro in the Neo, same chip as on the Iphone 16 Pro
- ac29 2mo ago> Apple has worked very hard to make unified memory a feasible approach, and the benefits of that are pretty clear in apple silicon- that efficiency not only results in power and therefore thermal gains, but also in a significantly faster full loop per process: or a faster time to token. This is why even their single core mobile chips in the budget line Neo out perform PC processors with several times more threads and RAM[1] Unified memory has existed for decades in the PC space, Apple didnt invent it. And the test you linked to has nothing to do with unified memory, its a web browser benchmark (almost entirely constrained by single threaded CPU performance that Apple better than competitors at).
- ethersteeds 2mo agoI think a major factor is memory bandwidth. Apple has raised it steadily for each M series generation, and that hasn't plateaued. Nvidia leads in bandwidth and specialized architecture, but local inference takes off when it's usably fast at much lower cost and power consumption.
- harrouet 2mo agoI could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens. Who needs memory when your model is set in silicon ?
- KeplerBoy 2mo agoYou don't an FPGA if you're taping out your own chips. But that is just a MMA accelerator with decent memory bandwidth. No secret sauce here.
- harrouet 2mo agoI am talking about reconfigurable gates to implement an LLM in silicon, i.e. an FPGA...
- xprnio 2mo agoWhat might the “parameters/layers to gates” ratio look like? My naive and uninformed guess would be that 1B+ parameter would also need a 1B+ gate FPGA, but according to google they typically range from tens of thousands to several million (which would still be a fraction of a billion).
- dgently7 2mo agowhy would apple make a chip that could be updated to improve the model when they could just sell you a better chip in the next years device? on device llm gives apple the new "better camera" "better screen" race they need to keep people coming back for the latest. for average users everything else is tapped out... screens, cameras wifi... all the core stuff is good enough now its hard to feel/see the difference model year to model year. embedded llm would let them ship something new and the on device ecosystem advantage is huge. especially as the gpt and claudes get ads and enshittified... the apple on device even if its less "capable" would be so compelling.