3 ms·
Google should create a translation module for dead languages. Why stop at scanning old books? Scan some ancient tablets Sergey!
by jamesk2 17y ago
Google should create a translation module for dead languages. Why stop at scanning old books? Scan some ancient tablets Sergey!
- jimmybot 17y agoModern statistical machine translation works by training on large amounts of already translated data, called parallel corpora. Monolingual data is helpful and can help you fix grammar, choose words that are more likely, but it won't work without that parallel corpora. Basically, you need a Rosetta stone, and if you had that, the humans could slowly deduce all sorts of things from only a small amount of data that would probably be useless for an MT system. Actually, scanning, or OCR, is a similar process--you train first on text that already has been properly transcribed, then you have a model that can be used to do OCR on new data.