3 ms·
> In theory, both of the accumulators are unbounded, but in practice, we noticed their values remain quite small This empirical fact is used to their advantage
by iraphael 10y ago
> In theory, both of the accumulators are unbounded, but in practice, we noticed their values remain quite small
This empirical fact is used to their advantage, but it is not proven that this is generally the case (tbh, it might be nearly impossible to prove it). This is okay, but it makes the model seem kinda yucky to me.
> All of the models
use 1024 LSTM nodes per encoder and decoder layers.
Ah, as I was reading the paper I kept looking for how they generated a variable-sized attention vector using a simple feed-forward network. It seems that the LSTM is actually capped at 1024 nodes per layer, so the attention vector only can be a fixed size of 1024 and truncated to the number of unrolled steps.