4 ms·
I agree that eventually this could work because on some training examples, two related entries will be in the same bucket. However, I'm not sure this would real
by lapink 7y ago
I agree that eventually this could work because on some training examples, two related entries will be in the same bucket.
However, I'm not sure this would really scale to parsing an entire book all at once like the author suggest. While the algorithmic complexity might scale, the odds for two related items that could be chapters appart to end up in the same bucket seems so close to zero that training time would explode.
In particular, this approach removes all kind of domain knowledge. For images, it means ignoring entirely the prior that neighboring pixels are related, which is typically encoded through the use of convolutions. With a Reformer, not only does the locality behavior need to be learnt from scratch, but on top of that it will only happen after a sufficient number of iterations so that neighboring pixels do end up in the same bucket.
For parsing books, I think it would make much more sense to build a hierarchical model with one part parsing only a paragraph at a time and generating an intermediate embedding that could then be used as a representation of the paragraph in a larger scale Transformer working over entire chapter, and then another level going from chapters to the entire book, rather than putting all the words at once together in a giant Reformer with no domain knowledge at all and praying that with enough training data and epochs, the model will learn everything from scratch.