5 ms·
I've been working on a language-learning app, and gpt-4 has made things doable that didn't seem to be doable without it. For example, translating to lesser know
by keymasta 3y ago
I've been working on a language-learning app, and gpt-4 has made things doable that didn't seem to be doable without it. For example, translating to lesser known languages like Cree. Online you can find a dictionary of words, or phrases, but I wasn't able to find a translator for arbitrary statements. It's not listed on google translate, nor yandex, et al. But, it's possible with gpt.
My strategy in automated translation is to translate, and then translate back to English. That way we can be fairly sure the translation is accurate. Of course if this app "has wings" I would open source the translations to allow corrections, and/or hire native translators for languages I don't speak (which is the vast majority of languages).
Last tip is that for me, I was able to make a pretty good automated system for this. What I did was spell out a monstrosity of a system-prompt which ensures a few things. One, it will always give me the most descriptive romanization for non-ASCII languages. Two, it will give me output that is essentially .csv data. Three, if it encounters a made up language like Pirate, it will try (and do a pretty good job). The rest is just parsing my final translation file to find prompt/language pairs which aren't saved yet and piping that to the monstrous prompt which queries for translations.
Then you can run a separate program for validation. See if it passes my game of telephone test.
This way I and others could learn common phrases in languages which are hard to otherwise access. It even works for dead languages/dialects/character-voices, etc. For me, it's pretty amazing actually. Please note that I never said perfect. But it's pretty damn close actually.
All in all this workflow allows what would have been unimaginable even two years ago.
In doing tests today I found it interesting/useful to note how gpt-4 is "thinking" about translating the word "settings". Here is its thoughts. This is not the exact system prompt I'm using for my app.
system_prompt = "You are an expert in translating lesser-known languages. When translating you will include both the native writing system, and the romanization into the latin alphabet. When you romanize text you always include any accents or pronunciation marks."
user_prompt = "Translate the following into Cree. [Hello, goodbye, settings]"
# Output
In Cree language, your words will be translated into:
1. Hello - ᑌᔭᔭᑎ, romanized as "Tānsi"
2. Goodbye - ᐊᔭᙱᐂᒥᑎ, romanized as "Ayāwāw"
"Settings" is a bit more complex, because it implies technological context that doesn't necessarily have a direct equivalent in Cree. However, a possible option is:
3. Settings - ᓂᐹᕗᓂᑕᐚᓇᐠ, romanized as "Nipāvunitawānāk", which might refer to "adjustments".
- arthurcolle 3y agoThis is such a cool example. "Adjustments" makes so much sense, doesn't have the very modern "settings" implication of a device UI/machinery
- keymasta 3y agoYeah. Hopefully it actually makes sense!
- arthurcolle 3y agohave you had any luck with old Akkadian scripts?
- schoen 3y agoI think there's some hallucination creeping in here! (1) This HN discussion is the only Google search result for each of these syllabics strings. (2) I tried using https://syllabics.atlas-ling.ca/ https://syllabics.atlas-ling.ca/ to transliterate these to Roman letters, and none of these was transliterated in the same way as the GPT-4 output (although the third one is somewhat close). (3) I searched and found that "hello" in Cree is likely written ᑖᓂᓯ (not ᑌᔭᔭᑎ), while correctly romanized as "tān[i]si". Your approach is clever, but I think the language model is still ultimately overconfident (and confused) here somehow.
- keymasta 3y agoYeah, for Cree it is definitely more suspect than trustworthy. Another thing I noticed was that on another attempt I actually received different translations, so.. it's hard to say how this is going to be refined to be usable, or if it indeed is at all. And wow, yes we are all alone on google results for those strings. EDIT 1: Another thought occurs to me, if it's getting the transliteration right, and not the syllabics, maybe I seperate the tasks and go english -> transliteration -> syllabic. I will have to see if that approach works better. Another idea might be to use that syllabics site to bring it from transliteration -> syllabic. I noticed that they were correct if translated there. EDIT 2: By updating the system prompt I was able to get it to translate properly. I had to remind it to be correct! You are an expert in translating Cree. When translating you will include both the native writing system, and the romanization into the latin alphabet. When you romanize text you always include any accents or pronunciation marks. You use syllabics properly and in the modern usage Hello - ᑕᓂᓯ (Tânisi) Goodbye - ᐅᑲᕆ (Okaawii) Settings - ᐅᑌᕁ ᐟ (Otēw with Roman orthography)