4 ms·
Non-standard datasets like these are very fun to play with. If you can produce the training data (source => target aligned text files), it is relatively simple
by srush 10y ago
Non-standard datasets like these are very fun to play with. If you can produce the training data (source => target aligned text files), it is relatively simple to try it out. Some mappings that people have recently published on: code => comments, ingredients => recipes, bad writing => good writing. I haven't seen the application that you are describing, but it would be pretty interesting.
Note though that NMT is particularly helpful for variable-length output. If you know the target is the same length as the source, then there are likely easier ways to go.
- amelius 10y agoHow large should a training set be, typically?
- srush 10y agoYou can often get something started with ~10000 examples. It's very problem specific though.
- dharma1 10y agoThanks for the reply, that sounds like a lot of fun! Do you happen to have links to the projects you mentioned for non-standard mappings? Would love to see the results and insights from them, before embarking on assembling training sets for my use case