4 ms·
In order to feel like a human, cues should not be a pre-programmed phrase, the system should continuously listen to the conversation, and evaluate constantly if
by Valgrim 3y ago
In order to feel like a human, cues should not be a pre-programmed phrase, the system should continuously listen to the conversation, and evaluate constantly if speaking is pertinent at that particular moment. Humans will cut a conversation if it's important, and such a system should be able to do the same.
- localhost 3y agoTotally agree with your take. But a pre-programmed phrase would work today and hopefully wouldn't be too difficult to implement. I would imagine that higher latency would be more tolerable as well. But in the fullness of time, your approach is better. When I'm listening to someone else talk, I'm already formulating responses or at least an outline of responses in my head. If the LLM could do a progressive summarization of the conversation in real-time as part of its context this would be super cool as well. It could also interrupt you if the LLM self-reflects on the summary and realizes that now would be a good time to interrupt.