3 ms·
Exactly. In the long run we want a cheap/fast LLM layer that we can call millions of times in a search/conversation-with-itself pattern. It may be that we'll
by waldrews 2y ago
Exactly. In the long run we want a cheap/fast LLM layer that we can call millions of times in a search/conversation-with-itself pattern. It may be that we'll interact with it the same way, and GPT-6 will be lots of GPT-5's talking to each other transparently to the user, but the path to next-level intelligence needs to get past generating the next token.
Since we don't know the real GPT-4 architecture, it may be doing some of this already.
- sebzim4500 2y agoI think we can be pretty confident GPT-4 isn't doing something like this based on its performance characteristics. i) The complexity of the prompt/answer does not effect the time per token. ii) We start getting a response pretty quickly (for short prompts) and then get new tokens at a roughly constant rate I don't think either of these properties would hold if there were a bunch of models that had to coordinate before they could start writing the output. Certainly techniques like ToT violate them.