4 ms·
IME I can get about 0.2s to get the first chunk from Mistral (i.e. Mistral API, using Mixtral model (`mistral-small`), not Mixtral on Groq) (and note the Mistra
by tomp 3y ago
IME I can get about 0.2s to get the first chunk from Mistral (i.e. Mistral API, using Mixtral model (`mistral-small`), not Mixtral on Groq) (and note the Mistral sends larger chunks, unlike ChatGPT which sends individual tokens)
and another 0.6s or so to get first voice chunks from PlayHT
measuring STT latency is harder, I need to implement a local VAD model first to properly measure it, but I think it's on the order of 0.5s
So this has nothing to do with Groq, really. ChatGPT is just slow (too slow for realtime voice communication).
- hobofan 3y agoUnless the only thing you want to do with the robot is talk, you need to do a lot more reasoning and execution planning first (= multiple LLM round trips; tool calling) before you even know whether talking is the correct action to take. So the naive time-to-first-chunk estimate will be way off.
- andoando 3y agojust add a hmmm before every response
- fragmede 3y agoWhich we humans do all the time. Okay, like, so, hear me out, alright? See, what’s really going on, yeah, is…
- fennecbutt 3y agoWhich is so cool because it's an evolutionary/language thing. Why do we add junk words while we think? I think it's probably because we're social animals, we want to hold that person's attention as we think as periods of silence are likely to make them become disengaged. But who knows really.