7 ms·
1. This is not exact attention, but an approximation of it. Specifically, they use k-nearest neighbors to retrieve the top-k most similar tokens, out of an "unl
by mxwsn 3y ago
1. This is not exact attention, but an approximation of it. Specifically, they use k-nearest neighbors to retrieve the top-k most similar tokens, out of an "unlimited-length input" say of size N, where k << N.
2. This idea is quite similar to retrieval transformers and Hopfield networks which have been known and published for several years now. It's not really that novel.
3. Due to the preceding points, the title can easily mislead people. It's not really a conventional transformer, and it's not a breakthrough.
4. This paper is a preprint and not peer-reviewed.
"I generally don't enjoy seeing preprints like this going to the top of Hacker News. This would be a higher quality submission if the paper was peer-reviewed or put into a greater context, like a blog post discussion or something like that."
Let me retract this and say something a bit nicer :)
I personally think there this specific preprint making it to the top of HN is potentially harmful, because of the hype around LLMs, the diverse audience of readers here, and the specific title that implies a claim of "transformer with unlimited context length", when this is misleading. I don't have anything against preprints in general - a lot of work outside of the peer-review process ends up being very impactful.
- ShamelessC 3y ago> This idea is quite similar to retrieval transformers and Hopfield networks which have been known and published for several years now. It's not really that novel. Is it? I had thought retrieval transformers "merely" used retrieval as a backend of sorts rather than a substitute for the attention itself?
- mxwsn 3y agoYeah, RETRO [0] embeds all an entire question/prompt, and searches for similar text passages with k-NN, then does further processing. This can kind of be understood as attention on paragraphs. This preprint instead does k-NN and calls it attention on single tokens. So not the same. But similar. [0] https://jalammar.github.io/illustrated-retrieval-transformer/ https://jalammar.github.io/illustrated-retrieval-transformer...
- ShamelessC 3y agoAh, I see - thanks for the clarification.
- make3 3y agoretro doesn't attend itself, which is a big difference
- dhruvdh 3y agoI generally don't enjoy something being diminished on account of being "not really that novel". Your comment essentially says - this is not a high quality submission because readers might not actually read it, which is no fault of the work, or submitter.
- MasterScrat 3y ago> Your comment essentially says - this is not a high quality submission because readers might not actually read it I'd argue that on average, most readers won't have a good enough understanding, or read the paper far enough, to understand that the reality is closer to "it's not a breakthrough" rather than "Transformers with Unlimited Length Input". So, I wholeheartedly welcome this type of hype-breaking leading comment.
- jjoonathan 3y agoAgreed 100%. Not only do I appreciate "well actually" comments, I think they are the single most useful aspect of forum discussions. The headline will always be "BATTERY BREAKTHROUGH PROMISES TO ROCKET ELON MUSK TESLA TO THE MOON!!!" and while it's easy to know that some amount of cold water is necessary you need to spend a nontrivial amount of attention and have a nontrivial amount of knowledge to figure out just how much cold water. It's a useful thing to outsource. Did a research group see outperformance in an experiment with 1% probability of translating into production? Or is CATL scaling up a production process? The "well actually" comment will contextualize for you. If there's a "well actually" reply to the "well actually" comment, that tells you something too. Upvotes/downvotes dial in the distributed consensus. It's far from perfect, but I'd challenge detractors to point to a more effective method for large-scale democratic truth seeking.
- zamnos 3y agowell actually (sorry, couldn't resist) the refrain of "correlation is not causation" despite not having read beyond the headline, or "...in mice" when that's mentioned in the abstract, is pretty frustrating when that's the entire substance of the comment. it seems some technical and social countermeasures could be deployed so there was at least token visiting of a link before commenting was allowed, which court raise discourse in this forum, at least. It's a false dichotomy to only consider the two extremes of Peer Reviewed and HN reviewed. In particular, as mentioned up thread, the incentive to do peer review (or replicate an experiment, for that matter) isn't as high as working on your own research, attempting to do some novel work for a shot at a Nobel prize. as every coder who's been the reviewer on a code review knows, it's difficult, and one that often is not well prioritized among other priorities a senior IC might have, leading to poor quality reviewing or other work slipping (only so many hours in a week). Thus, one could imagine a system that gives direct compensation for review work to grad and smart undergrad students who work in the area being discussed vetting and/or contextualizing claims like "this work is/is not novel", rather than hoping someone who does work in that area is procrastinating on HN at just the right time to make the claim and the rebuttal and the rebuttal-rebuttal. If the rubuttal is posted after the thread falls off the front page, is anyone not in that thread even going to know that what they read was wrong?
- chaxor 3y agoThere's nothing really wrong with a preprint making it to the top - there can be genuinely good work that stays in preprint for quite some time. I believe the original ELMo work that spurred the Sesame street gang is still in preprint despite its importance in NLP (:shocked Pikachu face: not a transformer?!). But yes, you're correct in this instance that it's not necessarily 'huge news' since it is highly similar to a long list of the Reformer (LSH-based), Performer (FAVOR**), FNet (Fourier-based), Routing Transformer, Sparse Transformer, Longformer (task specific sparse), blockbert, XLNet/xfmr-xl (slide + relative PE), BP-Transformer (binary partition), BigBird (global and rand attn), RWKV which is..., etc. ** FAVOR actually is innovative and different in this space, but towards similar ends anyway
- visarga 3y agoHow come you know the efficient-transformers family, when I ask questions about transformers in ML interviews nobody has heard of them. Can't figure out why it's not common knowledge. For years all the transformer papers were about reducing O( N^2 )
- chaxor 3y agoThe reason they don't know them is they're not serious researchers or practitioners of ML - it's as simple as that. Anyone in this area should have this exceedingly basic common knowledge.
- f_devd 3y agoTo be fair ML is (used to be?) pretty broad, so unless someone is actively keeping up with the sota in the high-data sequence modeling area it's quite possible to miss. I know ML teams which were entirely made up of OSML practicioners, because that was the most commonly useful until recently.
- Nimitz14 3y agoWhy learn something noone is using.
- 3y ago
- cs702 3y agoAfter a very quick read, that's my understanding too: It's just KNN search with some bells and whistles. So I agree on points 1-3. When something works well, I don't care much about point 4. Personally, I've had only mixed success with KNN search on long sequences. Maybe I haven't done it right? I don't know. In my experience, nothing seems to work quite as well as explicit token-token interactions by some form of attention, which as we all know is too costly for long sequences (O(n²)). Lately I've been playing with https://github.com/hazyresearch/safari https://github.com/hazyresearch/safari , which uses a lot less compute and seems promising, though it reminds me of things like FNet. Otherwise, for long sequences I've yet to find something better than https://github.com/HazyResearch/flash-attention https://github.com/HazyResearch/flash-attention for n×n interactions and https://github.com/glassroom/heinsen_routing https://github.com/glassroom/heinsen_routing for n×m interactions. If anyone has other suggestions, I'd love to hear about them.
- joseph_grobbles 3y ago[dead]
- ftxbro 3y ago> I generally don't enjoy seeing preprints like this going to the top of Hacker News. This would be a higher quality submission if the paper was peer-reviewed or put into a greater context, like a blog post discussion or something like that. This opinion seems totally backwards to me. I'm not sure what you think peer-reviewed means? Also I prefer full preprints than blog posts. But then again, I have no idea why ones like the daily blogposts of Seth Godin (to pick on one randomly, sorry it's not personal) so often go to the top of hacker news. Maybe opinions like yours explains it?
- MacsHeadroom 3y ago> This opinion seems totally backwards to me. I agree. > I'm not sure what you think peer-reviewed means? Posting to HN is a form of peer-review, typically far better than the form of "peer-review" coopted by journal publishers.
- pyth0 3y ago> Posting to HN is a form of peer-review, typically far better than the form of "peer-review" coopted by journal publishers. This is a rather self-aggrandizing view, and I think it speaks to the level of ego that underpins a lot of the discussion on here.
- nullc 3y agoIt would be more charitable and accurate to read it as a statement of the sad state of review at many journals. Plenty are rubber stamps where the most you might expect from reviewers is an insistence to add citations to their own barely relevant papers.
- JoshuaDavid 3y agoDo you think it's factually incorrect that the HN comment section is more likely to find problems which invalidate the conclusions of the paper than the journal-driven peer review process?
- godelski 3y agoHonestly, these complaints (other than 4) apply to the vast majority of papers. #4 is just false. It has already been viewed by other lab members (peers) and open publication is peer reviewing. The "peer review system" (publishing to conferences/journals) is relatively new and I think ML demonstrates all the problems with the system (yay hype). Novelty is especially a joke. ViTs are "just" NLP encoding transformers. T2I models are "just" NLP models connected to generative models. Diffusion models are "just" whitening models. GPT3 is just GPT2 with more layers and more data which is just GPT with more layers and more data. We can go even deeper if we pull from math and physics works. But that doesn't mean these works haven't been highly fruitful and useful. I'm happy all of these have been published. > because of the hype around LLMs I too hate the hype, but it is often bimodal. There are people who are far too critical and people who are far too accepting. The harm is not preprints or people reading papers, the harm is people who have no business/qualifications evaluating works confidently spouting out critiques. It is people not understanding that researchers are just critical of one another's work by default and that doesn't mean it shouldn't have been published. It is well known that reviewers are good at identifying bad papers but not good at identifying good papers[0,1]. Which let's be honest, that means reviewers just have high reject rates in a noisy system. Making publication as a metric for merit a highly noisy one at best. As for the paper: Many LLMs and large models are using attention approximations. Nor is the kNN technique particularly new. My main complaints are a lack of comparisons for Figure 3 and 4, but I'm not a NLP person so I don't even know if there's some other good works that can compare better (BART is a common baseline). But generative models are (unfortunately not notoriously known) extremely difficult to evaluate. Paper seems fine to me. It is useful to the community. I don't like the name either, but their input is limited by computer memory, not the model. I would want to see more on this. Not a NLP person all I can say is that this looks neither like a strong reject nor a strong accept. I'll leave it to the community to determine if they want more experiments for the conference publication or not but the work seems useful. [0] https://inverseprobability.com/talks/notes/the-neurips-experiment-snsf.html https://inverseprobability.com/talks/notes/the-neurips-exper... [1] https://arxiv.org/abs/2109.09774 https://arxiv.org/abs/2109.09774