3 ms·
This is a great example of issues with LLM design, which is that they apply a constant amount of compute to both sentences (modulo a token or two).
by TaylorAlexander 2y ago
This is a great example of issues with LLM design, which is that they apply a constant amount of compute to both sentences (modulo a token or two).
- croemer 2y agoDon't LLMs do some degree of search and backtracking if they end up in a dead end (with limits, a couple of tokens at least). Sometimes you'd see ChatGPT hanger words it had already output.
- TaylorAlexander 2y agoActually I am not very familiar with the internals, I am mostly repeating what Yann Lecun said in an interview a few months ago about autoregressive models. https://www.youtube.com/watch?v=1lHFUR-yD6I https://www.youtube.com/watch?v=1lHFUR-yD6I
- croemer 2y agoSo beam search _does_ do some sort of search: https://towardsdatascience.com/temperature-scaling-and-beam-search-text-generation-in-llms-for-the-ml-adjacent-21212cc5dddb https://towardsdatascience.com/temperature-scaling-and-beam-...