7 ms·
Building a language translator from scratch with deep learning
- jeffreyrogers 8y agoThis is very cool. One thing I wonder about though is whether small companies will be able to compete with large ones like Google in ML in the future. One reason Google's translator is better is because they have way more data. In the past they digitized tons of books so they have an excellent dataset that has been translated by professional, human translators. This data collection is effectively cross-subsidized by Google's primary business: advertising. Since most competitors to Google offerings aren't going to have a hugely profitable core business with which to fund all the data collection and normalization that goes into building a high quality ML system, the future for poorly capitalized competitors to compete seems bleak to me. This seems to support some of the growing rumblings about enforcing antitrust laws against the large tech companies. Edit: better, not bigger.
- saip 8y agoAccess to parallel corpora is a limiting factor in general. A good way to train a language translator is to use an open source dataset (several here http://opus.nlpl.eu/ http://opus.nlpl.eu/) to train a base model, and then fine-tune it with a smaller dataset specific to your domain. In this case, the author claims pretty good accuracy, almost on par with Google Brain's! On my test set of 3,000 sentences, the translator obtained a BLEU score of 0.39. This score is the benchmark scoring system used in machine translation, and the current best I could find in English to French is around 0.42 (set by some smart folks as Google Brain). So, not bad.
- jeffreyrogers 8y agoWow, missed that part when I read it. Pretty incredible that using open source data you can outperform the state-of-the-art machine translators of a few years ago.
- zawerf 8y agoFor a historical perspective check out stanford's nlp course: https://youtu.be/IxQtK2SjWWM?t=1267 https://youtu.be/IxQtK2SjWWM?t=1267 Deep learning only started beating tradition methods in 2016!
- samlynnevans 8y agoYes, it shouldn't have been far off Google as the model in the article is Google's itself. In fact you see in some of the examples how Google Translate's output and this model are almost exactly the same. Pretraining is a very cool idea, I have also seen some good results with pretrained embeddings after doing language modelling. Fast ai discusses this and I think even has some pretrained embeddings available in their library!
- l9k 8y agoDeepL had a lot of good press when it came out last year. Some saying it was better than Google. https://www.deepl.com/en/translator https://www.deepl.com/en/translator
- thirdsun 8y agoWhen someone recommended DeepL to me I almost didn't take it seriously expecting mediocre translations that aren't bad, but hardly usable without heavy editing. However after trying it I'm very impressed with its results and in those cases where you have a better translation in mind the interface offers an easy way to suggest and replace expressions. It's impressive.
- akie 8y agoWow, thank you for mentioning that. I cannot believe how good the translations are! My native tongue is Dutch and I threw in some (long!) English, French and German texts and honestly, they read like they were written by a native speaker. Hugely impressive.
- batterseapower 8y agoThe value of large corpora for translation may be diminishing.. In particular, Facebook have achieved impressive results using unsupervised ML for translation: https://code.fb.com/ai-research/unsupervised-machine-translation-a-novel-approach-to-provide-fast-accurate-translations-for-more-languages/ https://code.fb.com/ai-research/unsupervised-machine-transla... The basic idea is to use word vector embeddings to build a source<->target dictionary, then combine this with a language recognition model to iteratively bootstrap a set of source<->target training examples for use with a conventional ML approach.
- sooheon 8y agoSo the value of a large corpus remains, it's just this one happens to be generated, as opposed to collected.
- pixelHD 8y agoAnother perspective in similar veins would be the rise of AutoML. Given its absurdly high computational cost, I'd think only enterprises with massive computational power at their disposal would be able to use it.
- virgilp 8y agoEU helps with this too, accidentally. All official documents are translated in all EU languages, with very high quality translators. And all these documents are public.
- kome 8y agoDeepl is already MUCH better than Google translate: https://www.deepl.com/translator https://www.deepl.com/translator
- samlynnevans 8y agoDespite 'The unreasonable effectiveness of RNNs', it's seeming CNNs and the solely attention-based models are managing to perform the same tasks, but with better results and faster speed!
- pixelHD 8y agoThe transformer paper was quite influential in machine translation space. This resource [0] posted here a while back is a good place to learn and get a better idea how it works. [0]: http://nlp.seas.harvard.edu/2018/04/03/attention.html http://nlp.seas.harvard.edu/2018/04/03/attention.html
- lucidrains 8y agoone of the best visual tutorials on the transformer I came across http://jalammar.github.io/illustrated-transformer/ http://jalammar.github.io/illustrated-transformer/
- pixelHD 8y agoWow, that does look really good! Thanks!
- lucidrains 8y agoyou're welcome! :)
- i_made_a_booboo 8y agoMachine translation has made some pretty impressive progress over the last decade. Unfortunately no methods will ever cover the very last mile as languages don't have perfect 1 to 1 mappings. Though it is amusing watching the machines try.
- jchw 8y agoIf we ever do solve the last mile, it would probably be one of the less interesting consequences, as it would probably imply we've built an algorithm capable of learning and thinking to a similar degree of a human. To that, though, I'm definitely not holding my breath :)
- felipemnoa 8y ago>>To that, though, I'm definitely not holding my breath :) Nature was able to do this. Sure, it took a couple billion years of evolution to get to this point, but it is doable. I'm betting that the chances of us inventing Strong AI within the next 100 years is almost a near certainty.
- i_made_a_booboo 8y agoThe last mile isn't solvable. Some languages contain concepts, set phrases, vocabulary and pop-culture references entirely unique to that language. There isn't a translation in every single case. Machines however will always try to come up with one and the results are amusing. Also people make the assumption that as soon as we make strong AI comparable to a human we will be to translate anything and everything (let's say we are excluding the last mile for arguments sake). That assumption ignores an important fact that sometimes translation is a team effort where certain words, phrases or concepts are debated among multiple translators to reach a consensus. It's not always done by a single intelligence. Some people might argue that's because people have far more limited capacity to consider all the examples in the corpus whereas a machine can consider all of lightning fast and thus can arrive at the right answer. A perfect edge case that illustrates why that doesn't matter and where multiple human intelligences will often grapple with how something should be translated would be what name to give to a movie you are translating to an international audience. The same movie often has quite different names depending on which language it gets translated into. There isn't actually a correct answer there is just answers that are deemed 'good enough'.
- psergeant 8y agoThe grammar correction in Google Translate is a little too good. I was trying to create some broken Russian phrases to send a Russian friend, but I’d put in weird or bad English as an input and get very good Russian as an output!
- simonveith123 8y agohttps://www.digistore24.com/redir/34473/SimonV/ https://www.digistore24.com/redir/34473/SimonV/
- skookumchuck 8y agoI find that google translator does very well when the text to be translated has no spelling errors and is grammatically correct. Add any errors, and it falls to pieces, even though a human reader doesn't have any issues with it.