4 ms·
One could bolt such a system on top of an LLM. An LLM is "just" a document completer. Given text, it predicts the following text. So there is nothing stopping
by gamegoblin 3y ago
One could bolt such a system on top of an LLM.
An LLM is "just" a document completer. Given text, it predicts the following text.
So there is nothing stopping you from bolting on a system which works like:
- Given the current streaming input, continually update a pool of N likely continuations (predict what the user will say)
- For each of those N likely continuations, pre-generate responses
- If the actual user continuation matches one of the continuations, use your pre-generated response
This is trading off compute for latency, as you won't use at least N-1 of those pre-generated responses.
- robbintt 3y agoThe loops is also quite large right now. Put a lot in, get a lot out, put a little in, get a lot out. I suppose there is some work from OpenAI to have a well formed response in a certain number of tokens or something, I am not sure how the "document size" part of their chat completion works.
- eru 3y agoThat's exactly what chess programs do when they think on your time.
- 8organicbits 3y agoWhat sort of latency are people seeing from LLMs? I type faster than an LLM outputs tokens, so predicting my continuation would be slower than waiting for me to type it. Most people speak faster. An LLM that knows when it should interrupt me, as the parent comment mentions, would be really cool but I don't think it has the ability to determine when an interruption would be helpful. I'd be annoyed if I was asking "what is one plus three plus five" but the LLM interrupts with "1+3=4", for example.
- Waterluvian 3y agoLike humans, context matters. “This reminds me of… that movie where uhhh… the big guy goes to death row and he like… he can help people with his special powers, like that mouse he names or.. uhhh…” You’d probably interrupt me and it would be appropriate and welcomed.
- 8organicbits 3y agoSure, but would it interrupt with auto-complete nonsense too soon or would it know when an interruption would help? > This reminds me of >> A cool spring day? > that movie where uhhh >> Star Wars is a movie. I'm human, so I can understand when I have enough info for a good guess (The Green Mile?) but that's a much different skill than what an LLM does, right?
- lucubratory 3y agoIn principle, understanding when it is appropriate to interrupt is the same sort of problem as understanding when it is appropriate to use a specific tool or API call. Both are deviations from natural language that the model has to decide to employ based on the context of the input, so yes it should be very possible for LLMs to do it. You're asking if they would do it well or if they would instead get it completely wrong and interrupt you all the time with irrelevant nonsense - I'm going to say that depends entirely on implementation. The bigger issue, in my opinion, is that current models running on current hardware at reasonable costs simply aren't performant enough to do the faster-than-speech prediction that is required to execute the concept well.
- villgax 3y agoI'm getting sub 2 second inference on V100 for Flan-UL2(20B)