6 ms·
Reformer, the Efficient Transformer
- occamrazor 7y agoThe demonstration on images is underwhelming at best. It is only marginally better than extending vertically the bottom row of pixels.
- allovernow 7y agoI'd be curious to see how it handles infill. This seems like it has potential.
- anentropic 7y agothe same with the demos for completing a phrase based on Crime & Punishment... having all recently seen the text completion demos of GPT-2 the Reformer examples are decidedly underwhelming I mean, I'm sure it's a great new technique and all
- 4gotunameagain 7y agoWould it be reasonable to add something to the tittle so it's clear it has nothing to do with electronics? Maybe it's just me.
- pests 7y agoThe transformer term in AI has been around for a few years now. I agree it can be confusing but the article also defines it by the 5th sentence.
- bluescrn 7y agoNothing to do with electronics, or with robots in disguise...
- toomim 7y agoYeah, like "Reformer, the Efficient Transformer for Machine Learning".
- The_rationalist 7y agoHow does accuracy compare in Nlp tasks vs XLnet? If we can have XLnet accuracy and fast inference on a single gpu, that would be revolutionary!
- lapink 7y agoThere is no argument for why the LSH would work well, especially at the beginning of training. As the weights are initially random, bucket assignment would be random as well. If predicting at position A requires info from position B, but they are not in the same bucket, there will be no gradient to get the query embedding of A closer to the key embedding of B. The reversible layer trick is neat though.
- gwern 7y agoWhy is that any worse than, say, starting with randomly initialized weights in general?
- MiroF 7y agoI haven't read this paper yet, but to answer your question: because bucket choice is a discrete decision - and discrete decisions are hard to pass gradients through
- gwern 7y agoBut you aren't 'making decision's to pass gradients through them at all. It's just fixed random projections AFAICT (https://openreview.net/pdf?id=rkgNKkHtvB#page=3 https://openreview.net/pdf?id=rkgNKkHtvB#page=3).
- MiroF 7y agoOkay, I've now done a brief perusal of the paper. Perhaps I'm wrong, but it seems to me that deciding on the bucket is the discrete decision. If you have two "words"/"contexts" in a sequence that ought to attend to each other, but they don't get bucketed together early in training, then there is no gradient pushing those two hidden states to be close to each other, because there is no comparison being done between the two contexts. In a standard transformer, on the backprop we can see something like "oh, you would have been quite closer to the correct answer on this sentence if you had matched the context for 'dog' with the context for 'treat' about 20 words back." But, here, if 'dog' doesn't get bucketed with 'treat', then there's no such gradient pressure. Eventually (and with enough hashing+bucketing), the embedding of the more relevant contexts will move closer together, but I'd suspect this might occur more slowly. Here's the authors describing the process: > We don’t differentiate through the hash bucket assignment procedure, or the choice of what order to sort the items into. Rather, these operations take query/key vectors as input where LSH maps nearby vectors to the same bucket with high probability. Therefore, the sorting re-adjusts any time parameter updates to cause relevant vector pairs to have higher dot product, and “unhelpful” vector pairs to have lower dot products. e: And here is a reviewer noting what I suspected about number of gradient updates, > the performance achieved by the proposed method after 140k iterations is achieved by the full attention after ~40k iterations [on imagenet64]
- foota 7y agoI wonder if this could be used for the Wikipedia compression challenge?
- gwern 7y agoProbably, but it seems like it would run into diminishing returns except on the longest articles, because AFAIK the articles in wikitext are provided alphabetically, and consecutive articles may have little or nothing to do with each other, rendering the very wide window pointless.
- foota 7y agoEh? My understanding from the article was that long is anything beyond a couple paragraphs, many (though maybe a minority by count) of the Wikipedia pages are much longer than this.
- gwern 7y agoLast I saw any Wikipedia statistics, the average page was ridiculously small, like a few paragraphs. People just happen to not spend much time reading 'stubs', is all, but you can get an idea by spending some time on Special:Random. So, most WP articles are something that would fit entirely into many architectures' windows: for example, GPT-2 at 1024 context is roughly 3k characters (and you can easily scale GPT-2 way beyond that, see sillysaurus's comment above - we're training a GPT-2 with a context window of 25k right this second). Chonky paragraphs! Reformer's advantage would come only from the subset of articles longer than that, and only from the improvement in prediction from the subset of characters out of window at the beginning of the article in trying to predict toward the end of the article. And then you have the article boundaries which largely 'reset' the memory. Reformer's advantage then would have to come from the chance that there is a relevant article somewhere accidentally alphabetically close enough to be in its window while predicting the current article.
- foota 7y agoMaybe you could do a topological sort over references between articles? There's certainly cycles but those could be broken arbitrarily.
- gwern 7y agoDiscussion & links to various implementations: https://www.reddit.com/r/MachineLearning/comments/eg1wr3/reformer_the_efficient_transformer_anonymous_et/ https://www.reddit.com/r/MachineLearning/comments/eg1wr3/ref...
- silvestrov 7y agoso many smart people and still using fuzzy PNG instead of SVG
- mkolodny 7y agoAny papers / blog posts / GitHub repos you can recommend to learn about using SVG images for neural networks?
- bowmessage 7y agoThe input examples are photographs, how can one take an SVG photograph..?
- the8472 7y agoThis looks like building blocks from cryptography inspiring ML
- darawk 7y agoThis seems like a big deal. An asymptotic reduction in the resource explosion created by larger attention windows should allow the development of substantially more complex models here.
- overlords 7y agoVowpal Wabbit has been doing this 'hashing trick' since the 200s. It also the feature interaction, which are the same thing as a layer in transformers (all against all matrix). So it seems like they are still catching up to where John Langford and crew were over a decade ago. And, the vowpal wabbit approach is extremely fast to train because it's only doing stochastic gradient descent on a linear function - linear regression. Transformers are much slower to train. EDIT: Downvoters, please see my last leaf to see why they're effectively the same. The guy responding here seems unfamiliar with all the functionality of vowpal wabbit.
- marcinzm 7y agoThe Google paper's hashing has, as best I can see, nothing to do with the Vowpal Wabbit's 'hashing trick.' The VW hashing trick is about hashing your input data (ie: words, fields, etc.) into an array to lower storage requirements and deal with novel data at run time. The google paper is about ordering the intermediate states of the neural network (ie: vectors) while preserving distance. This is done so you can chunk the resulting ordered list and perform computations on individual chunks (and their neighbors). The only thing in common I see is the fact they both use the word hashing.
- overlords 7y agoThey are doing the same thing - using less memory by hashing. The hashing trick in VW hashes multiple same words into one integer, not the same as reformer, but similar to how reformer puts similar vectors together. With VW's ngram/skipgram features, you get the same kind of effect - similar strings hash into the same hash. So locality sensitive hashing = (is around about the same thing as) ngram/skipgram on strings plus hashing trick.
- marcinzm 7y agoExcept in Google's paper the hashing does not directly reduce memory usage in any way. It's a lossless operation on the original vectors unlike VW's lossy operation. Google's representation allows for memory reduction down the line but those mechanisms have nothing to do with hashing.
- sillysaurusx 7y agoOne neat trick is that you can extend GPT-2 117M's context window from 1024 up to 30k on a TPU, since TPUs can allocate up to 300GB of memory for backprop. https://twitter.com/gwern/status/1218001309435072513 https://twitter.com/gwern/status/1218001309435072513 It's not quite 1M words, but a 30k context window is big enough for e.g. most midi songs.