4 ms·
The way the recurrence in this method works -- ie, using last LLM hidden state at previous time step as input token for the next time step -- isn't directly com
by psb217 2y ago
The way the recurrence in this method works -- ie, using last LLM hidden state at previous time step as input token for the next time step -- isn't directly compatible with how recurrence/autoregression is typically handled during LLM training. One of the major strengths of transformers is that they can be trained for recurrence/autoregression (which have sequential dependency) using convolutions (which are embarrasingly parallel). The proposed method requires introducing some sequential dependencies during training that could otherwise be avoided using "causal masking" and convolutions to enforce the correct dependencies between time steps in a sequence. Introducing these sequential dependencies makes training a lot slower.
tldr; the method requires training in a way that loses one of the major benefits of transformers, but maybe in some scenarios that loss is worth it.