3 ms·
Perhaps we can use a generalized Unicode-like encoding space for LLM text tokens. A text tokenization scheme uses up a few hundred thousand entries, with say on
by OutOfHere 6d ago
Perhaps we can use a generalized Unicode-like encoding space for LLM text tokens. A text tokenization scheme uses up a few hundred thousand entries, with say one thousand new entries added annually. These can be called amojis, meaning AI mojis.