4 ms·
I somewhat doubt this because transformers by their nature rely on attention to prior tokens to derive their outputs. Removing tokens from the context fundament
by cheald 4y ago
I somewhat doubt this because transformers by their nature rely on attention to prior tokens to derive their outputs. Removing tokens from the context fundamentally changes the function output.
There might be gains to be had in understanding which tokens produce the lowest attention weights in the prompt, and then trimming those out. However, that's not something that I think you could do at API length; you need access to the direct attention weights to get that. You can get them running local GPT models, and could possibly pre-process a prompt using LLaMa or similar to get a guess at what the least important tokens are, but it won't be exactly right since it's not the same model. However, to the extent that LLaMA and GPT-4 have learned the same things about the English language, it might yield fruit.