4 ms·
Instant AI Response
- nacs 7mo agoWhat model and hardware powers this? Is this a Google T5 based model?
- pella 7mo ago3bit hard-wired Llama 3.1 8B ( https://taalas.com/the-path-to-ubiquitous-ai/ https://taalas.com/the-path-to-ubiquitous-ai/ )
- cyansmoker 7mo ago3bit is a bit ridiculous. From that page I am unclear if the current model is 3 or 4bit. If it’s 4bit… well, NVIDIA showed that a well organized model can perform almost as well as 8bit.
- personalcompute 7mo agoThis is a demo of Taalas inference ASIC hardware. Prior discussion @ https://news.ycombinator.com/item?id=47086181 https://news.ycombinator.com/item?id=47086181
- deleted 7mo ago[deleted]
- pella 7mo ago- https://news.ycombinator.com/item?id=47086181 https://news.ycombinator.com/item?id=47086181 - https://taalas.com/the-path-to-ubiquitous-ai/ https://taalas.com/the-path-to-ubiquitous-ai/ - https://www.nextplatform.com/2026/02/19/taalas-etches-ai-models-onto-transistors-to-rocket-boost-inference/ https://www.nextplatform.com/2026/02/19/taalas-etches-ai-mod...
- Kuyawa 7mo agoIf this is possible, why not all online AI engines work like this?
- yomismoaqui 7mo agoThis is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.
- sixtyj 7mo agoIt is really a crazy speed. 15k tokens/second.
- sixtyj 7mo agoI have tried it again. This is the future of chat UI, imho. Generated in 0,074s • 15 754 tok/s
- sbrother 7mo agoDo we understand how to scale up the hardware to the point it can run a frontier model? Because this is insane. It will be a game changer for agent systems making 10-100+ calls.
- alansaber 7mo agoI love seeing optimised SLM inference. Is there a current use-case for this? Edge CNNs make sense to me but not edge SLMs (yet).
- notronic 7mo agoimagine a model like opus 4.6 at that speed, that would be insane
- OutOfHere 7mo agoImpressive, but this particular underlying LLM is objectively weak. I'd like to see it done with a larger and newer better model.