Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
srush
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
srush
10y ago
It's not there yet, but the improvement has been quite significant in aggregate. Key table from the GoogleNMT paper, empirically showing a 60% relative improvement on this task: PBMT GNMT Human Relative Improvemen
92.
▲
by
srush
10y ago
It's tricky to give you an exact answer. For translation, the minimal size we have been using is about 1 million aligned sentences, although people often report on smaller data. There are also lots of tricks to get around small dataset
93.
▲
by
srush
10y ago
I'm glad you mentioned this. There is a lot of interest these days in character-based machine translation, including several papers in review at ICLR. The current practical consensus (at least in OpenNMT) is that character-only models
94.
▲
by
srush
10y ago
Unfortunately we do run into these types of "robustness issues" in our experiments. They're rare enough that they don't really change many of the performance metrics of the system, but they are quite embarrassing. It is
95.
▲
by
srush
10y ago
Hey! Yeah I taught that class three years ago... Several papers that really demonstrated this was possible at scale came out around over the next year, the most well known is "Sequence to Sequence Learning with Neural Networks" (
96.
▲
by
srush
10y ago
The code implements a pretty close model to what Google Translate has published ( https://arxiv.org/abs/1609.08144 ). However the two systems are likely trained on very different datasets.
97.
▲
by
srush
10y ago
Better: Train the same model for all twenty language pairs ( http://opennmt.net/Models/#multi-way---fresptitrofresptitro ). Even better: Use OpenNMT to do the OCR too ( https://github.com/opennmt/im2t
98.
▲
by
srush
10y ago
While this type of translation is heinously understudied, the opposite problem of controlling the politeness forms of translation is actually an important area of research. For instance Controlling Politeness in Neural Machine Translation v
99.
▲
by
srush
10y ago
What's interesting about neural machine translation is that the core model is completely language pair independent. So we roughly use the same code for Russian-English, English-Russian, and Chinese-German. That being said the errors in
100.
▲
by
srush
10y ago
Yes, you can actually run them on a phone! We had an earlier demo of very strong system running on android https://github.com/harvardnlp/nmt-android Without tricks the large models takes up about 700 megs, opennmt.net&
101.
▲
by
srush
10y ago
Hi, I'm Alexander Rush (@harvardnlp), one of the project leads on OpenNMT and an assistant prof at Harvard. Feel free to ask me anything about the project.
102.
▲
The Declassification Engine
(wired.com)
1 points
by
srush
13y ago
|
0 comments