3 ms·
People also often forget "orderless autoregression", which was introduced a while back and has been reinvented many times since. See Sec 4 (pg 8) of "Neural Aut
by psb217 2y ago
People also often forget "orderless autoregression", which was introduced a while back and has been reinvented many times since. See Sec 4 (pg 8) of "Neural Autoregressive Distribution Estimation" [https://arxiv.org/abs/1605.02226 https://arxiv.org/abs/1605.02226]. The main difference from current work is that this 2016 paper used MLPs and convnets on fixed-length observations/sequences, so sequence position is matched one-to-one with position in the network's output, rather than conditioning on a position embedding. Of course, Transformers make this type of orderless autoregression more practical for a variety of reasons -- TFs are great!
Key quote from Sec 4: "In this section we describe an order-agnostic training procedure, DeepNADE (Uria et al., 2014), which will address both of the issues above. This procedure trains a single deep neural network that can assign a conditional distribution to any variable given any subset of the others. This network can then provide the conditionals in Equation 1 for any ordering of the input observations. Therefore, the network defines a factorial number of different models with shared parameters, one for each of the D! orderings of the inputs. At test time, given an inference task, the most convenient ordering of variables can be used."