3 ms·
How does this differ from the 2018 NeurIPS paper, Blockwise Parallel Decoding for Deep Autoregressive Models? https://arxiv.org/abs/1811.03115 https://arxiv.or
by nsagent 2y ago
How does this differ from the 2018 NeurIPS paper, Blockwise Parallel Decoding for Deep Autoregressive Models?
https://arxiv.org/abs/1811.03115 https://arxiv.org/abs/1811.03115
- tripplyons 2y agoThey use a separate ngram model to generate the proposed sequence instead of extra heads on top of the main model. The process of verifying the proposed sequence appears to be the same.