5 ms·
As far as I can tell, the best tool for localisation almost nobody is using is http://www.grammaticalframework.org/ http://www.grammaticalframework.org/. Licens
by dkbrk 12y ago
As far as I can tell, the best tool for localisation almost nobody is using is http://www.grammaticalframework.org/ http://www.grammaticalframework.org/. Licensing is a mix of GPL, BSD and MIT pieces.
It's a high-level functional programming language with a dependent type system specialised for operating on language ASTs. It's resource library, to quote "covers the morphology and basic syntax of currently 29 languages: Afrikaans, Bulgarian, Catalan, Chinese, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hindi, Japanese, Italian, Latvian, Maltese, Nepali, Norwegian bokmål, Persian, Polish, Punjabi, Romanian, Russian, Sindhi, Spanish, Swedish, Thai, Urdu."
In essence, once it has the language-independent AST, it can produce output in all its supported languages with the correct tenses, genders, inflections, etc.
It also seems to have tools for assisted parsing, so you could have an english document and interactively parse it into the correct AST. In addition, the text can be parameterised semantically, so if you changed the gender of a person, that could propagate to all the correct locations and update the translations as required.
While it seems the upfront cost may be quite high in having to learn such a complex system, I think the benefits of having reproducible, high-quality outputs into n languages for free could make this highly advantageous in many applications.
- canjobear 12y agoI'm very skeptical that this would work outside of toy examples, though it depends on what is meant by language-independent AST. For example, the best way to translate Spanish "X dió un golpe a Y" would be "X hit Y". But my naive idea of what the AST for the Spanish sentence would look like would be something like `(GIVE (X HIT Y)`, which when naively transduced to English would be the "X gave a hit to Y", which is either unidiomatic or means the wrong thing altogether. In order to avoid this problem, the AST would have to be a more abstract representation of the semantics. And coming up with a sufficiently expressive, tractable, and neutral representation of natural language semantics is an unsolved problem that people are still devoting their whole careers to. I was briefly involved in a very early stage startup that was considering using systems like this for better machine translation. We ran into problems like the above, and also: ambiguity, and the fact that the hand-written grammars and semantic representation systems were just very brittle and incomplete.
- nitrogen 12y ago"X dealt a blow to Y" seems like it could be a more grammatically similar translation.
- rmc 12y ago> which when naively transduced to English would be the "X gave a hit to Y", which is either unidiomatic or means the wrong thing altogether What's interesting, is that there are dialects of English (Hiberno-English spoken in Ireland) where "X gave Y a hit" would be a way to say "X hit Y". :)
- schoen 12y agoThat sounds like a great technical approach, although it can't necessarily remove the problem of idiomaticity that the article mentions (with the example of "I didn't search any directories"). Probably better examples are possible, in any case where the most idiomatic way to express something isn't a literal translation of that thing from other languages. Maybe like English "I don't care" Portuguese "tanto faz" (literally 'so much does') German "[das] ist mir egal" (literally '[it] is equal for me')
- taejo 12y agoBTW, egal doesn't mean equal in German (though it's borrowed from French égal, which does). It means "irrelevant", "all the same", "unimportant".