4 ms·
Agreed that the generated sample is superior to similar outputs from GPT-2. Looking at the additional samples in the publication, my first thought is that the
by JRKrause 7y ago
Agreed that the generated sample is superior to similar outputs from GPT-2.
Looking at the additional samples in the publication, my first thought is that the model cannot easily stray from or modify the context. Once a fact is stored within the compressed memory, it seems the model cannot easily generate sentences contradictory to that fact.
This is problematic because frequent changes to relational information (e.g. the location a character is standing) is fundamental to story telling.
- deleted 7y ago[deleted]