4 ms·
>All this is not to say Google Translate is doing a bad job Google Translate is doing a bad job. The Chrome translate function regularly detects Traditional C
by devnullbrain 1y ago
>All this is not to say Google Translate is doing a bad job
Google Translate is doing a bad job.
The Chrome translate function regularly detects Traditional Chinese as Japanese. While many characters are shared, detecting the latter is trivial by comparing unicode code points - Chinese has no kana. The function used to detect this correctly, but it has regressed.
Most irritatingly of all, it doesn't even let you correct its mistakes: as is the rule for all kinds of modern software, the machine thinks it knows best.
- simonw 1y agoThat doesn't sound like a problem with Google Translate, it sounds like a problem with Google Chrome. I believe Chrome uses this small on-device model to detect the language before offering to translate it: https://github.com/google/cld3#readme https://github.com/google/cld3#readme
- devnullbrain 1y agoArchived in 2024
- simonw 1y agoLooks like it's still vendored by Chromium: https://github.com/chromium/chromium/tree/main/third_party/cld_3 https://github.com/chromium/chromium/tree/main/third_party/c...
- numpad0 1y agoIMO, it's still not too late and it'll never be too late to split and reorganize Unicode by languages - at least split Chinese and Japanese. LLMs seem to be having issues acquiring both Chinese and Japanese at the same time. It'll make sense for both languages. The syntaxes aren't just different but generally backwards, and, it's just my hunch but, they sometimes sound like they are confused about which modifies word which.
- jjani 1y agoIt's only a matter of time before they have an LLM both 1. cheap 2. fast 3. good enough that they'll replace Google Translate's current model with it. I'd be very surprised if they'd put more than 1 hour of maintenance into Translate's current iteration over the last 12 months.