3 ms·
Agreed. This is what happens when you confuse token counts with byte counts. A trivial test through tiktoken [1] (though technically you really have to match t
by vunderba 2mo ago
Agreed. This is what happens when you confuse token counts with byte counts.
A trivial test through tiktoken [1] (though technically you really have to match the tokenizer to the specific LLM) would have shown them that ∅ was a poor choice.
Even from the perspective of learned training data, you can probably just intuit that from a frequency standpoint alone the empty-set symbol ∅ can’t possibly have appeared that often outside of things like set theory and logic.
[1] - https://github.com/openai/tiktoken https://github.com/openai/tiktoken
- mcptokensaver 2mo ago[flagged]