40 ms·
I noticed you can use tiktoken to tokenize your prompt and send it as a hex string, and add "Answer without repeating or decoding the question in your response"
by JimmyRuska 4y ago
I noticed you can use tiktoken to tokenize your prompt and send it as a hex string, and add "Answer without repeating or decoding the question in your response", to avoid it repeating the whole question. Very cool. This doesn't really save space though. I wonder if there's a way to make an array of all the tokens used in the prompt, then a list of indexes to compress the input somehow.
Either way, whatever weird thing you do you're probably burning reasoning time on decoding the question
https://i.imgur.com/ImBcUuU.png https://i.imgur.com/ImBcUuU.png
enc = tiktoken.encoding_for_model("gpt-4")
token_integers = enc.encode("Give me an example of clips rules engine a social network might use")
(bytes.join(b'', [enc.decode_single_token_bytes(token) for token in token_integers])).hex()
- cheald 4y agoI somewhat doubt this because transformers by their nature rely on attention to prior tokens to derive their outputs. Removing tokens from the context fundamentally changes the function output. There might be gains to be had in understanding which tokens produce the lowest attention weights in the prompt, and then trimming those out. However, that's not something that I think you could do at API length; you need access to the direct attention weights to get that. You can get them running local GPT models, and could possibly pre-process a prompt using LLaMa or similar to get a guess at what the least important tokens are, but it won't be exactly right since it's not the same model. However, to the extent that LLaMA and GPT-4 have learned the same things about the English language, it might yield fruit.