8 ms·
Hi, I'm Alexander Rush (@harvardnlp), one of the project leads on OpenNMT and an assistant prof at Harvard. Feel free to ask me anything about the project.
by srush 10y ago
Hi, I'm Alexander Rush (@harvardnlp), one of the project leads on OpenNMT and an assistant prof at Harvard. Feel free to ask me anything about the project.
- j0e1 10y agoHi, This is a great effort! Really excited about what you are doing. I'm curious about the minimum size of the dataset that would be required to get any reasonable output. I understand that this maybe dependent on the language pair, yet some concrete numbers (like for a few language-pairs shown in the demo) would help me get an idea. Also, what can I do to make it work for language pairs which have very small parallel/aligned data available. I would be grateful for any pointers regarding this.
- srush 10y agoIt's tricky to give you an exact answer. For translation, the minimal size we have been using is about 1 million aligned sentences, although people often report on smaller data. There are also lots of tricks to get around small dataset problem. (1) You can pre-initialize your model with monolingual word embeddings. We recommend using the Polyglot embeddings which exist for many different languages. See http://opennmt.net/Advanced http://opennmt.net/Advanced for details. (2) You can train your model with a nearby language. For instance we have a model that uses data from all the Romance languages simultaneously. (3) You can use monolingual data in other ways. For instance, if you are translating into English, you can combine with a standard language model or pretrain on the English data. There are a bunch of other approaches in the literature, but these are some of the more common tricks.
- dharma1 10y agogreat work! I've wondered if NMT models could be used with other types of data, like music notation (midi or ABC). The use case I had in mind was "translating" a monophonic melody input to a polyphonic ouput, ie. auto-arranging melodies. Of course assuming there is an available dataset of input-output pairs to train with.
- srush 10y agoNon-standard datasets like these are very fun to play with. If you can produce the training data (source => target aligned text files), it is relatively simple to try it out. Some mappings that people have recently published on: code => comments, ingredients => recipes, bad writing => good writing. I haven't seen the application that you are describing, but it would be pretty interesting. Note though that NMT is particularly helpful for variable-length output. If you know the target is the same length as the source, then there are likely easier ways to go.
- amelius 10y agoHow large should a training set be, typically?
- srush 10y agoYou can often get something started with ~10000 examples. It's very problem specific though.
- dharma1 10y agoThanks for the reply, that sounds like a lot of fun! Do you happen to have links to the projects you mentioned for non-standard mappings? Would love to see the results and insights from them, before embarking on assembling training sets for my use case
- bjmrey 10y agoHi, thanks for offering your help! A couple of questions: - Which resources (books, courses, tutorials...) would you recommend to learn how to use OpenNMT? I am a programmer, with vary basic knowledge of NLP concepts. - I see a dictionary integration in the Systran demo. I assume this is a Systran product, not something included in OpenNMT? Or am I wrong? Thanks!
- themedvedev 10y agoHi Alexander! I'm a trilingual mobile software engineer and I've always been fascinated with breaking down language barriers (particularly interested in Russian-English and English-Russian translation). In the context of these languages, what are the scenarios where OpenNMT performs poorly and how can I contribute to improving performance in these areas?
- srush 10y agoWhat's interesting about neural machine translation is that the core model is completely language pair independent. So we roughly use the same code for Russian-English, English-Russian, and Chinese-German. That being said the errors in Russian are quite different than those made in other languages due to case endings. For instance for a similar size data set there are often 5x more unique Russian words than in English. But if you want to get involved more generally our gitter is http://gitter.im/OpenNMT http://gitter.im/OpenNMT and our forum is at http://forum.opennmt.net http://forum.opennmt.net.
- kmicklas 10y agoThis seems like an argument for a character based rather than word based network. It just so happens that English and Chinese, the two languages which have the most machine learning research, are relatively analytic, with a low morpheme per word ratio. But many world languages have a much higher ratio (Russian wouldn't even rank that high!) and acquiring training data covering all unique "words" is essentially impossible.
- srush 10y agoI'm glad you mentioned this. There is a lot of interest these days in character-based machine translation, including several papers in review at ICLR. The current practical consensus (at least in OpenNMT) is that character-only models are not really worth the efficiency loss. A simple compromise is to use Byte-Pair Encoding as a preprocessing step in morphologically rich languages and allow the model to produce sub-word chunks. This is implemented in OpenNMT as a preprocessing option (see http://opennmt.net/Advanced http://opennmt.net/Advanced).
- arthur2e5 10y agoAs a terrible human being, I tested Chinese-to-English translations on f-word-based profanity on the demo page (https://demo-pnmt.systran.net/production https://demo-pnmt.systran.net/production). The results were arguably not as good as Google Translate, which can recognize the colloquial form 操[^1] but not the more orthodox 肏. PNMT's results were great in another way as they form more fluent English. [^1]: Everyone hijacked the character. On the other hand, OpenNMT appears able to translate the "f* you!" sentence to "你妈!" in Chinese, a form of "操你妈" with the F word not spoken but implied. The single-f-word sentence generates amusing variations when used with different punctuations, although many ("...", "?", "!", "") fail by reducing the sentence to a normal call for one's own mother. These experiments make me curious about what kind of corpora the system was trained on. * * * The test sentences used were: 操你十八辈子祖宗![^2] 操你十八辈祖宗! 操你妈! [^2]: This one is technically wrong. For the orthodox version, replace all instances of 操 with 肏. As an additional experiment, I tested some variations of "mother" profanities without "操": 1. 你妈! 2. 他妈的! 3. 他妈的。 Between #2 and #3, PNMT appears more sensitive to punctuations than human do. #1 is only included as a round-trip validation. Edit: (#1) ... and is not supposed to be translated back into profanity due to high false positive rate.
- liushuyu 10y agoI think that's kind of uncertainty of the nature of neutral network? Or the training data is not very trustworthy. IMO, the latter is more probable.
- srush 10y agoWhile this type of translation is heinously understudied, the opposite problem of controlling the politeness forms of translation is actually an important area of research. For instance Controlling Politeness in Neural Machine Translation via Side Constraints (http://homepages.inf.ed.ac.uk/abmayne/publications/sennrich2016NAACL.pdf http://homepages.inf.ed.ac.uk/abmayne/publications/sennrich2...) is something that OpenNMT can support.
- m_ke 10y agoHey Sasha, I took the NLP course with you at Columbia when you covered for Michael Collins 2-3 years ago. I vaguely remember either you or Michael mentioning that it wouldn't really be feasible to do multi lingual translation by mapping the input text to some universal intermediate language (embedding space) and then decoding into the target language (this was in the context of pharse based systems). It's great to see that the MT field has made such great progress in the last few years and that the latest NMT models are doing just that.
- srush 10y agoHey! Yeah I taught that class three years ago... Several papers that really demonstrated this was possible at scale came out around over the next year, the most well known is "Sequence to Sequence Learning with Neural Networks" (https://papers.nips.cc/paper/5346-sequence-to-sequence-learning-with-neural-networks.pdf https://papers.nips.cc/paper/5346-sequence-to-sequence-learn...). It's been quite fun watching something I assumed was too hard at the time, become essential to the field.