4 ms·
Why would that be annoying? It’s much easier to understand, predict and truncate appropriately than having to explain all of these different tokenization scheme
by ntonozzi 3y ago
Why would that be annoying? It’s much easier to understand, predict and truncate appropriately than having to explain all of these different tokenization schemes to devs.
- rcoveson 3y agoYeah, everybody agrees on what a character is, right? It's just {an ASCII byte|a UTF8 code unit|a UTF16 code unit|a Unicode code point|a Unicode grapheme}.
- ntonozzi 3y agoI’m not saying it’s easy but it’s much better than tokens IMO. I think bytes would be understandable too.
- geysersam 3y agoAt least there are standards for characters. Nothing like that for tokens.
- sheepscreek 3y agoAnd we think tokens solve that problem? Spoiler alert: they don’t https://www.reddit.com/r/OpenAI/comments/124v2oi/hindi_8_times_more_expensive_than_english_the/ https://www.reddit.com/r/OpenAI/comments/124v2oi/hindi_8_tim...
- est31 3y agoThey don't but Google could have been more precise with which of the definitions listed by GP they mean by "character".