3 ms·
Yeah that’s not how next token prediction works. To actually do multiple passes you’d need to do that yourself, making multiple calls and feeding the responses
by block_dagger 2y ago
Yeah that’s not how next token prediction works. To actually do multiple passes you’d need to do that yourself, making multiple calls and feeding the responses back to the model.
- IanCal 2y agoWhy? The very nature of next token prediction means it's entirely capable of having that. It's not multiple passes, it's just one pass. You making multiple calls is just inserting fixed tokens then asking it to carry on completing.
- bloomingkales 2y agomaking multiple calls and feeding the responses back to the model. By asking it to reconsider half its generated response, aren’t I essentially asking it to formulate the second half of its response from the first half internally? I’m bypassing the manual process of feeding in the extra prompt. We are constantly having to tell the LLM close, but no cigar, iterate again, more or less.