4 ms·
“it seems translation is a simple string replacement. I don’t think that’s enough to cover all languages.” Honestly, simple string replacement isn’t good enoug
by hhas01 6y ago
“it seems translation is a simple string replacement. I don’t think that’s enough to cover all languages.”
Honestly, simple string replacement isn’t good enough to cover any languages. Homonyms and synonyms, anyone? I suspect machine translation will be harder with code simply because there’s a lot less contextual information around individual words compared to ordinary prose with which to make a best guess as to meaning.
“I think AppleScript tried to do better, but can’t find examples”
Dr Cook’s HOPL paper on AppleScript gives an example of French and Japanese dialects (p20; can’t it paste here as it’s an image):
www.cs.utexas.edu/~wcook/Drafts/2006/ashopl.pdf
Not really the same thing as Citrine as it relied on manual localization of an application’s resources (custom keywords defined in AETE terminology, vs labels and tooltips on GUI controls). And, as you say, it did not do anything to localize user-defined words.
In any case, AppleScript is a really good example of how NOT to design an accessible language syntax. All the rigidity and tolerance of a machine language, with all the complexity and ambiguity of a natural one faked on top. (Plus all the fun that comes with arbitrary keyword injection—argh.)
This is why I say artificiality is good. Humans aren’t after Shakespeare, they’re after high-level understanding of program code.
- Someone 6y ago“Homonyms and synonyms, anyone?” I think, for this case, that’s mostly solvable by being careful to not make assumptions. For example, you should create separate translation strings for the ‘for’ in “for…in” and in “for…each”, and for the ‘define’ in “define function” and “define procedure”. Also, when (not if; if your language is successful, it will happen) translators report problems, you should ‘simply’ not hesitate to introduce new ‘clones’ of to-be-translated words. Even ignoring languages with different word order, that wouldn’t make things perfect. For example, there will be languages where the correct way to say “define” depends on the (perceived) plurality or gender of the function name. I think the correct way to handle this is by letting the translator produce a grammar for your language that produces the same tokens as the ‘original’. That probably is beyond many would-be translators, though.