3 ms·
It's crazy. Are they doing any precomputing as you type, I wonder if you paste a block of text is it the same speed.
by jcul 2mo ago
It's crazy.
Are they doing any precomputing as you type, I wonder if you paste a block of text is it the same speed.
- nerdsniper 2mo agoI pasted and instantly hit enter on this prompt: "I generated a filter set using REW v5.31.3 using real-world sweep tone measurements from the room I'm listening in . How can I use it as my MacOS output equalizer so that my spotify music is adjusted for this room and speakers" and it gave a very reasonable answer in non-perceptible time.
- jaggederest 2mo agoThat's a hell of a prompt, you should make a post about what you're doing maybe, because I want that for myself.
- weiliddat 2mo agoidk if you're still actually looking for a solution (I got nerdsniped heh), but I found a seamless way to have it locked into the speakers + room, instead of a specific source/software, is using an iLoud subwoofer that routes to any speaker setup and has room correction + EQ that lives on the sub.
- throwuxiytayq 2mo agoNo, it really does take ~0.03s to generate the answer. Try your browser's developer tools and watch the requests.
- HDBaseT 2mo agoIsn't it crazy that we can send a message, across the world near instantaneously and have a coherent reply, generated by a computer, sent back to your screen, in under 500ms in most circumstances. I find myself getting caught up in the sheer speed of modern computing and networking. The fact I can play an online game with 10 other people is just insane.
- jaggederest 2mo agoI got into Rust development via LLM last year, and being able to do things budgeted in nanoseconds is a heady feeling indeed. Real time video and audio analysis? Totally doable, plenty of time budget. 16.6ms is a long time, it turns out.
- Cort3z 2mo agoThe magic is in the fact that they essentially have an ASIC llm device. There is no other trickery. The problem they will face is that it is actually locked in silicon, so upgrading models will be difficult, and likely require new hardware each time.
- ch4s3 2mo agoAt some point I imagine you’d add a software layer on top that holds more current training and can be called as needed trading off for slower responses. There’s already work out there splitting models across networks. You could have the base on silicon, some stuff in memory on the machine, and another frontier tool in the cloud.
- jaggederest 2mo agoTaalas actually already support LoRA, basically doing exactly what you say. The other thing that I think is really interesting about all of this, is that LLMs are already perforce behind the times with their knowledge cutoff, so adding an additional ~3 months for bake into silicon isn't such a huge deal, I think, for the ~10x more efficient and faster you get.
- ch4s3 2mo agoYeah, really a fascinating time in computing. Once these chips start to become more common I'll be curious how people find ways to use them. One can imagine a world where really simple inference is available on dirt cheap chips found in toys and other low cost consumer electronics.
- theshrike79 2mo agoThere are already studies proving that a "stupid" model with a good harness + tool calling will outperform a "smart" model. Things like this give me hope for a system that can be fully local and private, but also with the ability to be almost infinitely extendable with tools.
- 2mo ago
- dyauspitr 2mo agoIt’s crazy that you asked the very question whose answer spawned the sub thread, within the sub thread.