3 ms·
It completely butchers Greek. No wonder it charges some much for so little output. Every Greek character is a token. I wonder if there is space for innovation
by v4dok 4y ago
It completely butchers Greek. No wonder it charges some much for so little output. Every Greek character is a token.
I wonder if there is space for innovation there. I would imagine that it similarly difficult for other non-English languages as well-known. I fear for the effect this will have on them.
- 0xDEF 4y agoIt's crazy that OpenAI hasn't fixed their tokenizer yet. They are leaving the door wide open for some Chinese big tech company to capture the non-Latin script parts of the world. i18n (and accessibility) was something American tech companies were serious about in the 90s and early 2000s. That is how they captured most of the global market. US tech dropping the ball on this leaves the door wide open for Chinese competitors.
- hombre_fatal 4y agoDoes OpenAI’s tokenizer issues cash out into having worse results for Greek rather than just being more expensive for gpt-4? (gpt-3.5-turbo already costs peanuts) If not, then this response seems overblown. The competitive advantage in LLM at this point probably is not tokenizer optimizations and more about having results worth a damn.
- v4dok 4y agoThe usability is worse. Token limits are so much easier to reach. It's like using a model with dementia.
- hombre_fatal 4y agoI didn't think about that but now it's obvious. Good point.
- Fauntleroy 4y agoIt could be that they are actively working on this problem, but the product has not yet been released.
- goldfeld 4y agoThere is a market opportunity here for a GPTesque thinking machine who actually masters and knows their greek ancients well. I knew it it was a lack of refined Platonic understanding when ChatGPT said it could not comment further on the Russian war.
- resters 4y agoIt seems to be an accidental advantage of the messy hodgepodge that is English. There are semantic clues everywhere in word order patterns.