5 ms·
200ms Voice LLM
- unraveller 2y agoI hope this catches on and can be transplanted into any LLM soon. Instant voice integration is poised to unlock many UX things like auto-pilot in cars making you a real powerful co-pilot where you can suggest fine-tune settings on the fly for the current situation or anything that changes lots of settings at once with visual feedback.
- ricardobeat 2y agoNio’s built-in assistant uses ChatGPT with Whisper (I think) and is almost real time already. It can change settings like that and rarely misses a beat. I expect once GPT 4o becomes available it will be awesome (even more if they “unlock” it to be asked generic questions and hold a conversation).
- mungoman2 2y agoYeah this is awesome. Keep reducing that latency, that's the path to the killer assistant.
- Gys 2y agoCannot wait for an instant translator. Something that can translate synchronously (!) one language it hears into a language that I understand. Getting closer!
- ofrzeta 2y agoI don't understand how that could possibly work. You will always need some time window, don't you? Also sometimes the translation of the first word can only be inferred when you process the last word in a sentence (simplified example).
- Gys 2y agoYou mean the job of a translator is by definition impossible? Edit: Yes, for sure there will be a delay of a few words or even one or two short sentences. Just like human translators. Not a problem I think. Edit 2: Very curious why my first comment was downvoted?
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- ofrzeta 2y agoObviously translation is possible but not "synchronously" as you wish. From what I understand 200ms is "time-to-first-token", so I still wonder how that works because, as I wrote in my comment above, typically there is no one-to-one correspondence of words/tokens from one language to another. (I didn't downvote your comment)
- ein0p 2y agoYou can’t translate word by word. You need entire phrases, sometimes pretty long ones, to understand the meaning.
- woleium 2y agosometimes, but in a constrained situation, e.g. at a hotel reception desk, maybe not.
- 2y ago
- akreal 2y agoCool! HF Transformers is great for prototyping and research, but should not an interactive tool like this be based on something more speed-focused, like llama.cpp? Any plans for languages beyond English?
- juberti 2y agoWe're running it on vLLM and are working with others in the community to bring it to other optimized inference frameworks.
- demarq 2y agoThat was amazing and so productive. Looking at the transcript I was able to cover so much more ground than if I was stuck typing all my questions with a keyboard. I can imagine a future in offices where people have Ai rooms where you go not for a meeting with other people but to have a convo with ai.
- morjom 2y agoIn the same vein is there any good, low latency, speech to text to speech (STTTS?) capable programs that are making use of LLMs or AI?