3 ms·
So much of machine translation is the quality of your parallel corpus and your preprocessing. I've got a theory that Google actually has too much data, and mayb
by jen_h 7y ago
So much of machine translation is the quality of your parallel corpus and your preprocessing. I've got a theory that Google actually has too much data, and maybe too much imperfect data. Also, Deepl came out of Linguee (https://en.wikipedia.org/wiki/Linguee https://en.wikipedia.org/wiki/Linguee), who appear to have a massive number of parallel texts...I really believe whoever who has the (well-curated) texts has the (NMT) world.
As an aside, I don't know who needs to know this, but we are really lucky right now in that we now have tools that enable pretty much anyone to train a language model and translate with it, not just giant or small companies. It's not necessarily going to be any good, but the tools are all right there for us plebes and it's pretty fun.
I pulled together a tutorial just this weekend walking through doing this on Google Colab (https://sevenminuteserver.com/post/2019-10-17-machine-translation-on-a-budget-google-colab-and-aws-spot-instances-with-opennmt-py/#opennmt-py-with-google-colab https://sevenminuteserver.com/post/2019-10-17-machine-transl...) and EC2 Spot instances (https://sevenminuteserver.com/post/2019-10-17-machine-translation-on-a-budget-google-colab-and-aws-spot-instances-with-opennmt-py/#opennmt-py-on-gpu-enabled-ec2-spot-instances https://sevenminuteserver.com/post/2019-10-17-machine-transl...) if anyone wants to play around with it.