4 ms·
Claude's tokenizers have actually been getting less efficient over the years (I think we're at the third iteration at the least since Sonnet 3.5). And if you pr
by curioussquirrel 6mo ago
Claude's tokenizers have actually been getting less efficient over the years (I think we're at the third iteration at the least since Sonnet 3.5). And if you prompt the LLM in a language other than English, or if your users prompt it or generate content in other languages, the costs go higher even more. And I mean hundreds of percent more for languages with complex scripts like Tamil or Japanese. If you're interested in the research we did comparing tokenizers of several SOTA models in multiple languages, just hit me up.
- arcanemachiner 6mo agoI would encourage you to post a link here, and also to submit to HN if you haven't already. :)
- curioussquirrel 6mo agoWill do! Thanks for the encouragement
- curioussquirrel 6mo agoHere you go! https://news.ycombinator.com/item?id=47847282 https://news.ycombinator.com/item?id=47847282