10 ms·
Unlimiformer: Long-Range Transformers with Unlimited Length Input
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- bighoki2885000 3y ago[dead]
- sva_ 3y agoI think infiniformer would've sounded better. The bench scores seem pretty marginal.
- mirekrusin 3y agoPretty marginal score gains once a week is all you need.
- sdenton4 3y agoOnly so long as a) the gains are real, and not overfitting the test dataset, and b) you don't balloon in complexity, so that stacking approaches becomes impossible to manage. Point (a) is extremely hard to discern, especially when people are chasing third-significant-digit gains on common benchmarks; it's essentially multiple-testing false discovery in action. I've seen whole families of methods fail to transfer to new domains... Point (b) is also a real issue. As you increase the number of bells and whistles, each with their own hyperparameters with non-linear impacts on model quality, it becomes impossible to say what's working or not. In practice, i think we see some cycles of baroque incremental improvements, followed by someone spending a year stripping away the bullshit and getting something simple that outperforms the pack, essentially because it's easier to do hyperparam search over simpler models once you figure out the bits that actually matter.
- smusamashah 3y agoWhat does it mean for ChatGPT and likes? Can they employ this method to virtually get rid of context tokens limit?
- Kranar 3y agoYes it looks like it can use this method. This method is a preprocessor and post-processor that can be used on an existing GPT model to augment it to handle unlimited tokens.
- vintermann 3y agoAnd that makes it pretty notable compared to all the linear attention/retrieval schemes that didn't pan out. Not saying this will pan out, but we'll know more without waiting six months for the model to train.
- XorNot 3y agoHang on, how unlimited is unlimited here? Surely the immediate thing you'd do with this is just never delete any prior inputs so it becomes defacto long term memory for the model?
- shishy 3y agoLast paragraph touches on that: The length of inputs is theoretically bounded by the memory limitations of the computer used. More practically, using a CPU datastore is many times slower than a GPU datastore because of slower search and the need to transfer retrieved embed- dings to the GPU... (continues)
- 0xDEF 3y agoThe limit is RAM but GPU RAM is much faster than computer RAM.
- davrosthedalek 3y agoIs that really the limit? There is no real restriction that everything is in memory at the same time, right? You could maybe stream from SSD?
- capableweb 3y agoCreate a swapfile and you essentially trade disk space for memory space.
- space_fountain 3y agoAs I understand it the approach here is to use an approximate nearest neighbor database to retrieve highly relevant tokens from across large documents using the existing attention heads. So each attention head retrieves context from entire document. They say this can work without fine tuning, but performance improves with it. This is apparently extending this piece of prior work, but they've managed to re-range the linear algebra of attention so they only need one database for all attention heads across all layers of the model. I'm a bit confused how attention would here for layers below the top and a bit confused about how position is encoded for tokens across a long document like this.
- im3w1l 3y agoI don't understand how this could work. Like if you select a small fixed number of tokens from a large document won't you necessarily lose a lot of important data?
- MagicMoonlight 3y agoWhen you read a book do you remember every word? Even every chapter? You only need the important concepts, not individual words
- Xelynega 3y agoI feel that's a bit different. When skimming or reading a book, even though we don't remember all the information typically a lot of the information is dependent on prior information, so we form a sparse knowledge graph from the entire text. If I'm understanding the article, this approach would not use the skipped words to influence the output, so I thinks ita necessarily different.
- ztratar 3y agoGiven the model performance is thus affected by a k-nearest neighbor, but those algorithms are proving not great for baseline vector search, how well will this actually work? It seems mostly like a vertically integrated vector DB + existing LLM call, but correct me if I'm wrong. There are of course some performance gains with that, but the holy grail of "understanding" at unlimited length still seems unsolved.
- mrbungie 3y agoIsn't the performance (as in the capacity of retrieval, not performance as compute/memory usage) of kNN mostly given by the quality of the vectors/embeddings themselves? Most vector DBs use (at least) some kind of KNN anyways.
- j2kun 3y agoAppendix C of the PDF has an example book summary of The Brothers Karamazov, with noted minor factual errors https://arxiv.org/pdf/2305.01625.pdf https://arxiv.org/pdf/2305.01625.pdf > Unlimiformer summary: > The first part of the novel focuses on the question of whether or not the Russian nobleman, Dmitri Fyodorovitch, has killed his father. In the town of Ivanovna, the lieutenant-colonel of the Mushenkhanovitch is accused of the murder of his brother Ivanovitch. The lieutenant-incommand, Vasilyevitch, takes the form of a dog, and the two men–the two men and the woman who are questioned by the court-martial–murphy. The two men cry out to the God of Russia for help in their quest to save the town. The man, afraid of the wrath of the God, hands the dog a bunch of letters that are supposed to be proof of his love for his brother. The old man–the one who had killed his mother, and then found the letter–arrives. He reads it–asked the old man to forgive him for the murder and then takes the dog away. The other men, all of whom are prisoners, demand that the man confess his crime to the court. The first and most important thing they tell the court is that they love the man. The court acquits the man and sentences the man to death. The second man–an old officer of the town, Alekandrovitch–askes to tell them the same thing. The third man–in the process of confessing his crime–is Vashenka, a drunk man who has been sent to the town to kill his father, for reasons which are not entirely clear to the people. The woman’s servant, Evgenyevna, is also the one who has told the court the story of the Medvedevitch’s murder, for the good old man’s and the young man’s love. The three men, who are separated for the first time, are laughing at the man’s attempt to seduce Mitya. The young man, in the meantime, is conscripted into the town-side. He tells the court that he loves her, but he has yet to tell her the true story. The men, in this room, demand a man to kill her, and she will not betray them. The women, in their own country, are rebelling against the man who had sent them three thousand roubles, and they will not allow the man of the people to see them. They will not let the man in the town be allowed to see the man–or Dmitriovitch; he will have her husband killed him. He will not tell the people who love him. The next man, named Vashenovitch, arrives, and takes the man away. They all begin to laugh at the fact that he has succeeded in seducing and entrusting his brother Dmitri. He is then taken away to the old woman’s house, where the governor-side-of-the-world, and his sister, Arkadin, is being punished. The priestesses and the baron are shocked, for they have been so virtuous and well-suited. The only thing they will be able to do is kill the priest. They threaten to burn the priestess to death, for she has been so wicked and libidinous that she has not yet seen the priest, for her husband. The priests–ostensibly convinced that she is a woman who loves the priest and has been punished for her love and for allowing the priest to marry her. The last man, Yakivitch, arrives at the house, and, after a long day of drinking and then some of the men–is killed. He and the priest are ordered to leave the town so that the priest can finally be reunited with the people of the old lady. The final man, the commander of the St. Petersburg town of Arkadina, is sentenced to death for the crime of having killed and then the lieutenant of the governor, for taking the money. The commander, the former lieutenant-delegation of the People’s Army, is summarily executed, and all the men, except for the commander, have been summarily punished for their crime. The entire town is shocked and, in a very dramatic way, the priestesses plead for the forgiveness of the man, for allowing them to kill and imprison Ivan. They plead for their brother to be restored as well, for all the people they have loved, and for the priestor to tell the story
- mxwsn 3y ago1. This is not exact attention, but an approximation of it. Specifically, they use k-nearest neighbors to retrieve the top-k most similar tokens, out of an "unlimited-length input" say of size N, where k << N. 2. This idea is quite similar to retrieval transformers and Hopfield networks which have been known and published for several years now. It's not really that novel. 3. Due to the preceding points, the title can easily mislead people. It's not really a conventional transformer, and it's not a breakthrough. 4. This paper is a preprint and not peer-reviewed. "I generally don't enjoy seeing preprints like this going to the top of Hacker News. This would be a higher quality submission if the paper was peer-reviewed or put into a greater context, like a blog post discussion or something like that." Let me retract this and say something a bit nicer :) I personally think there this specific preprint making it to the top of HN is potentially harmful, because of the hype around LLMs, the diverse audience of readers here, and the specific title that implies a claim of "transformer with unlimited context length", when this is misleading. I don't have anything against preprints in general - a lot of work outside of the peer-review process ends up being very impactful.
- ShamelessC 3y ago> This idea is quite similar to retrieval transformers and Hopfield networks which have been known and published for several years now. It's not really that novel. Is it? I had thought retrieval transformers "merely" used retrieval as a backend of sorts rather than a substitute for the attention itself?
- mxwsn 3y agoYeah, RETRO [0] embeds all an entire question/prompt, and searches for similar text passages with k-NN, then does further processing. This can kind of be understood as attention on paragraphs. This preprint instead does k-NN and calls it attention on single tokens. So not the same. But similar. [0] https://jalammar.github.io/illustrated-retrieval-transformer/ https://jalammar.github.io/illustrated-retrieval-transformer...
- ShamelessC 3y ago
- ftxbro 3y agoOther times this was put on hacker news: https://news.ycombinator.com/item?id=35823039 https://news.ycombinator.com/item?id=35823039 https://news.ycombinator.com/item?id=35803470 https://news.ycombinator.com/item?id=35803470
- swores 3y agoWhile I appreciate your intent and effort - I don't think it's actually useful to link to other submissions unless either they have comments (ideally only if there's at least one interesting comment, but at least more than no comments at all), or if it's a submission of the same subject but to a different source link - in which case it's probably more useful to just link the alternative source, if it's worth reading, rather than potentially split the discussion into separate comment threads if the other is empty. Linking to a different submission of the same link with 0 comments doesn't add anything.
- ftxbro 3y agoI must have submitted it at the wrong time of day.
- swores 3y agoSure or just random luck, maybe this submission just happened to take place when the only few people who care about this subject happened to come online, or vice versa for bad luck before etc. But unlike sites like Reddit, with the exception of self / ask HN / etc posts, nobody really pays attention to who the submitter is, so enjoy the conversation finally breaking out on it as consolation for not getting karma points, but skip linking to dead submissions :) FYI, if you ever submit something that fails to get any traction / upvotes, then I've seen mods say before (@dang will hopefully correct me if I'm wrong) that a) it's OK to try submitting a second time maybe after a day or so (but not keep submitting over and over) or b) send the mods an email with a brief reason why it's a link that should interest HN readers for it to be potentially added to a "second chance pool". Though in the case of this link, between three of you it was posted two days ago, one day ago, and today which has finally got a bit more notice, so worked out alright in the end :)
- adamnemecek 3y agoThe attention mechanism corresponds to the Hopf algebraic convolution, a generalization of the commonly known convolution. I'm in the process of implementing a framework based on this idea. I have written a paper on this recently, https://arxiv.org/abs/2302.01834 https://arxiv.org/abs/2302.01834 I have a discord channel https://discord.cofunctional.ai https://discord.cofunctional.ai.
- capableweb 3y agoYou never just work on something until it's being ready to be shared, and then share it once? It has to be shared before it's even a little bit usable, with just some vague words about what it might be?
- adamnemecek 3y agoI'm gauging interest and looking for potential users. Steve Blank and all that.
- verdverm 3y agoThe first step to crossing the chasm is finding those innovators and learning if you are solving a problem!
- adamnemecek 3y agoI have and I am. Next.
- TeMPOraL 3y agoIs this how Kagi's "universal summarizer" works? They wrote a lot of copy about how it's able to summarize websites and documents of arbitrary length, while not revealing how on Earth this actually works. It does seem to work, though.
- KaoruAoiShiho 3y agoCould that not just be some kind of langchain like system?
- GistNoesis 3y agoI've read the paper quickly, the main idea is simple and interesting, but maybe a little dubious (it's kind of an accuracy for memory trade-off). In the transformer architecture one has to compute QKT. QKT=(hd * Wq * WkT)heT (equation (2) page 3 in the paper). Where hd is the hidden state of the decoder, and he is the hidden state of the encoder, and Wq and Wd are some parameters matrices, and T denotes the transposition operation. By grouping the calculation this way, in a transformer encoder-decoder architecture, they can build and use only a single index (you index the he vectors using a vector database) for all the decoder layers queries. Instead of having to build 2 * L * H indices (with L the number of layers of the decoder and H the number of head in the decoder). But what makes it a little dubious, is that this transformation mean you make your near neighbor queries in a space of dimension "dimension of the hidden state", instead of "dimension of a head" that is H times smaller. So if you had to build 2 * L * H indices each index would be H times smaller. So you only gain a factor 2 * L. But the trade-off is that you are doing a near neighbor search in higher dimension where you are then subjected to the curse of dimensionality (the higher the dimension the more similar all points are to each other). Whereas the whole point of projections in transformer is to lower the dimension so that the knn search make more sense. So to get the same accuracy, your near-neighbor search engine will have to work a lot harder. Also as an approximation of the transformer, because it's using some knn search, it comes with the problems associated with it (for example harder to train because more sparse, and a tendency to hyperfocus), but it can be complemented with low-rank linearization of the attention to also have the neural net act on the gist rather than the closest neighbors.
- numeri 3y agoThis technique can be added on to any encoder–decoder Transformer model post-training, so the added training difficulties you mention don't apply. It honestly is a very interesting approach to me – the main issue I see (which they discuss in the paper) is in pure latency. If you're using a large enough vector database, it will be on the CPU, and transferring hidden states from GPU to CPU and then the embeddings back from CPU to GPU is going to eat up a ton of time.
- jerrygenser 3y agoThis is a nitpick but also it's been a few years since I was taking academic ML and linear algebra courses. However regarding this part. > So you only gain a factor 2 * L. But the trade-off is that you are doing a near neighbor search in higher dimension where you are then subjected to the curse of dimensionality (the higher the dimension the more similar all points are to each other). I thought that the curse of dimensionality meant that in higher dimension, points got farther apart
- nephanth 3y agoBtw, why do transformers have a limit input size in the first place? I'm pretty sure the self-attention mechanisms scale (although with bad complexity) to arbitrary sizes
- MacsHeadroom 3y ago>(although with bad complexity) Because of exactly that. Also the attention mechanism is baked in during pretraining. So whatever max context length you want increases the compute cost of training by at least a function of said "bad complexity." Even just 4096 tokens of max context is much more expensive to train than 2048. So if we want models with 8k, 32k, or more context then the training costs get out of hand quickly.
- rewq4321 3y ago> Also the attention mechanism is baked in during pretraining IIUC, this is no longer necessarily true with positional encodings like ALiBi: https://github.com/ofirpress/attention_with_linear_biases https://github.com/ofirpress/attention_with_linear_biases
- chrgy 3y agoIn the age of transformers , lets ask a transformer to summarize this paper: The Unlimiformer paper is about a new way to make computer programs that can summarize really long pieces of text. Normally, when you ask a computer program to summarize something, it can only handle a certain amount of text at once. But with Unlimiformer, the program can handle as much text as you want! The way Unlimiformer works is by using a special technique called a "k-nearest-neighbor index" to help the program pay attention to the most important parts of the text. This makes it possible for the program to summarize even really long documents without losing important information. Overall, Unlimiformer is an exciting new development in natural language processing that could make it easier for computers to understand and summarize large amounts of text.
- moffkalast 3y agoSaid transformer as it handled the article's length anyway: sensible chuckle
- bighoki2885000 3y ago[dead]
- szundi 3y agoInput should be the Internet then.
- quickthrower2 3y agoPricing: $0.1 per nano token.
- opportune 3y agoThis seems like a definite attention optimization but I think the fundamental problem with attention is that it doesn’t handle state in a way that scales well. Personally I think the RNN/LSTM state handling approach is going to be something we revisit when trying to advance past transformers. It handles state in a way that generalizes and scales better (it should in theory learn an attention-like mechanism anyway, and state is independent of input size). It may be harder to train, and require further improvements, but it really seems more like an engineering or cost problem than a theoretical one. But I’m only an amateur and not an expert. Maybe continued improvement on attention will approach generalized state handling in a way that efficiently trains better than improvements on more generalized stateful approaches improve training.
- logophobia 3y agoAn alternative which I've used with some succes are structured state space models: https://srush.github.io/annotated-s4/ https://srush.github.io/annotated-s4/. A very different approach that works well for quite a few types of problems.
- intalentive 3y agoThe ML community keeps rediscovering the work of Steve Grossberg. This is very similar to his decades-old ART model.
- vardhanw 3y agoCould you explain in simple terms (if possible) what is the similarity? For context, I worked for a brief time with ART before Y2K in BU CNS, and took a few courses there, but had to leave it for 'reasons'.
- intalentive 3y agoSome ART concepts 1) Normalize input (batch norm, 2015) 2) Competitive dynamics / lateral inhibition (softmax in attention layers, 2017) 3) Cluster best matching activation vectors (top-k keys, 2023)
- jfisher4024 3y agoNeubig is the real deal. I’d take this paper seriously.