3 ms·
I wonder if we're at a point where you could build a voice assistant like that, except almost-realtime and streamed end to end: User speaks and speech to text
by Spiwux 3y ago
I wonder if we're at a point where you could build a voice assistant like that, except almost-realtime and streamed end to end:
User speaks and speech to text starts streaming text while the user is still speaking. That text stream is piped into a LLM, which also streams its output text. That output text is streamed to text-to-speech, which also generates audio in a streaming manner.
- WiSaGaN 3y agoAny existing stream api for llm input?
- everforward 3y agoYou can do the "almost-realtime" part, all locally. I tinkered with a Python script for a few hours that used Whisper to speech-to-text, fed that into a local Mistral model (don't recall which), and then piped the output into text-to-speech. It wasn't really streamed, though. Audio input was buffered, fully evaluated to a string, then fed into the LLM and the full text was converted back to audio. The Whisper speech-to-text was pretty real-time, the LLM was not. I was barely scraping by on hardware specs, though.
- canadiantim 3y agoyou try using ESP box?
- evilantnie 3y agoTTS and STT models have decent support for streaming in chunks, but the accuracy drops the smaller the chunk size. Current state of LLMs are pretty limited in their ability to handle streaming inputs due to attention window constraints. There is some emerging research into attention sinks and caching initial tokens that look promising. I don't think we're quite there yet though.
- modeless 3y agoI implemented this! All local models. And I packaged it up so people can install it with one click: https://apps.microsoft.com/detail/9NC624PBFGB7 https://apps.microsoft.com/detail/9NC624PBFGB7 The speech recognition part needs work for sure, but when it works you can see the potential. It's very different from the way it feels to talk to Siri or even ChatGPT's voice mode. It won't be long before we are having real conversations with our computers.
- bjelkeman-again 3y agoCould you record a demo of this?
- modeless 3y agoI really should! I'm not the type to publish videos of myself usually, but it really does need a video demo.
- 3abiton 3y agoBut how realtime is it?
- modeless 3y agoThe end-to-end response latency is around 1 second typically. It listens continuously, there are no buttons to press, and you can interrupt it while it's talking.
- adroitboss 3y agoThis has happened already. It was maybe about 7 months ago and I believe it was a twitter link posted here. They took it further and streamed it to twilio to create a live phone call.
- fudged71 3y agoThe one I tried was called Vocode
- zaptrem 3y agoAvailable as a phone line API (https://www.vocode.dev https://www.vocode.dev) and OS project (https://github.com/vocodedev/vocode-python https://github.com/vocodedev/vocode-python)