3 ms·
"the disfluent but more likely accurate results" -> Human evaluation does not agree that PBSMT is more accurate than NMT. Sure, the fluency/"sounds good" gains
by MiroF 6y ago
"the disfluent but more likely accurate results"
-> Human evaluation does not agree that PBSMT is more accurate than NMT. Sure, the fluency/"sounds good" gains from NMT are much larger than the accuracy gains, but evaluations I've seen put NMT ahead on both metrics.
- 082349872349872 6y agoFair enough, this is just my sentiment, where accurate was not a term of art, but meant "honest where it is weak and where it is confident." Before, I could skim outputs, and get a fair idea of which areas were poorly translated by where it was garbage. Now, I get the impression that (like with GPT or style transfer) something always comes out, whether it's related (or even necessarily has the right negation parity) or not.
- MiroF 6y agoYou're absolutely right with your impression of GPT/these models - they will generally generate plausible text. And it is also worth noting (as you've sort of mentioned) that appearance of fluency can bias translators who are instructed to just evaluate accuracy, or to evaluate accuracy and fluency separately. There is ongoing research into that problem.