4 ms·
The future of inference is likely in ASICs, so we'll get the inverse, a bit less capable than frontier but super fast models. Like this 14k tok/s beast https://
by batperson 23d ago
The future of inference is likely in ASICs, so we'll get the inverse, a bit less capable than frontier but super fast models. Like this 14k tok/s beast https://chatjimmy.ai/ https://chatjimmy.ai/ from Taalas (who got acquired by AMD recently).
GPT-6-astra runs at like ~40 tok/s, I have a hard time imagining what could be accomplished with that type of model at 10k+ tok/s when in the hands of the public. Will certainly make cybersecurity a challenge for older systems.
- domhudson 23d agoThis is incredible! Are there other big players in this space (freezing models to silicon)?
- HeWhoLurksLate 23d agotake a look at Cerebras, who are doing wafer-scale compute
- timcobb 23d agoI imagine Astra is/will soon will be on Cerebras?
- SPascareli13 23d agoLike how crypto used ASICS but then didn't because the scaling of consumer hardware made it obsolete?
- actionfromafar 23d agoAm I missing some joke here?
- wtallis 23d agoTo the extent that cryptocurrency moved off ASICs, it was because of interest shifting to different cryptocurrencies that were specifically designed to be harder to mine on an ASIC than Bitcoin's compute-heavy, memory-light hashing. I'm not sure there's any reason to expect a similar shift from LLMs. The hardware used for training doesn't dictate what hardware needs to be used for inference, and nobody's going to design an LLM architecture with an overt intention to make it better suited to GPUs and hard to target with ASICs.
- SPascareli13 23d agoYet it doesn't seem that ASICs will have any particular advantage over consumer hardware since AI is very memory heavy, which is (right now) expensive no matter how you package it. And the compute is just simple matrix multiplication, which is almost entirely what GPUs were meant to do anyway.
- infecto 23d agoGo back and correct your idea that consumer hardware made asics obsolete. Then we can figure out if asic or asic like devices for inference will have no advantage.
- andy_ppp 23d agoExcept Taalas is much faster than GPUs, orders of magnitude so. They aren’t going to get 100x faster at inference any time soon!
- deleted 23d ago[deleted]