8 ms·
A sleep-like consolidation mechanism for LLMs
- gemsquared 4mo ago[flagged]
- jgreid 4mo agoIsn't this simply context pruning/optimization?
- kylemaxwell 4mo agoFrom the abstract, it looks like it's actually doing something deeper, updating weights in part of the model?
- samsartor 4mo agoThe abstract and method sections only mention updating the SSM state during "sleep" (ie the same vectors that change after each token in stock Mamba) not any of the actual weight matrices. AFAICT this is just another attention compaction paper, with misleading tile? It is not very clearly written
- colechristensen 4mo agoNo, they're actually training weights based on context before compaction. Context is context, this is splitting the model into persistent weights and malleable ones which are periodically updated.
- delis-thumbs-7e 4mo agoWouldn’t that be extremely computationaly expensive considering how resource incentive training is?
- colechristensen 4mo agoNo, training a state of the art model involves training on the order of 10 trillion tokens. We're talking about a step that updates weights based on say between 10k and 1M tokens.
- delis-thumbs-7e 4mo agoI learned something. Thank you!
- pcrh 4mo agoI can't pretend to understand how LLMs work, but I can be sure that anthropomorphizing their functions is not helpful to an objective debate over their abilities. Does a motor vehicle get "sleep" when it is serviced? When I reboot a computer, is that equivalent to a nap?
- eithed 4mo agoI assume compacting is the sleep here; so, yes
- ajs1998 4mo agoThis is the struggle of naming papers. You could stretch definitions and make your own sexy headline or you could be precise and fewer people will read it.
- colechristensen 4mo ago>we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast weights before clearing its key-value cache There is a strong, non-trivial connection here between what your brain does in sleep and what they are studying. You wouldn't object to referring to robot eyes or robot legs.
- cowlby 4mo agoThe analogy is helpful, but yes we should be able to “intelligently design” something better than sleep analogues since we’re not constrained by evolution like in humans.
- lxgr 4mo agoWe are however constrained by the complexity of any purported solution. That's the bitter lesson, in a nutshell. At the very least, we know that sleep and dreaming do exist in biological brains. (Doesn't mean any of it is applicable to artificial neural nets, doesn't mean it'll work for our specific architectures etc. etc., but at least the idea requires fewer assumptions than a completely untested novel theory.)
- 4mo ago
- throwaway613746 4mo ago[dead]
- swyx 4mo agorelated preprint from the letta team https://arxiv.org/abs/2504.13171 https://arxiv.org/abs/2504.13171 Scaling test-time compute has emerged as a key ingredient for enabling large language models (LLMs) to solve difficult problems, but comes with high latency and inference cost. We introduce sleep-time compute, which allows models to "think" offline about contexts before queries are presented: by anticipating what queries users might ask and pre-computing useful quantities, we can significantly reduce the compute requirements at test-time. To demonstrate the efficacy of our method, we create modified versions of two reasoning tasks - Stateful GSM-Symbolic and Stateful AIME. We find that sleep-time compute can reduce the amount of test-time compute needed to achieve the same accuracy by ~ 5x on Stateful GSM-Symbolic and Stateful AIME and that by scaling sleep-time compute we can further increase accuracy by up to 13% on Stateful GSM-Symbolic and 18% on Stateful AIME. Furthermore, we introduce Multi-Query GSM-Symbolic, which extends GSM-Symbolic by including multiple related queries per context. By amortizing sleep-time compute across related queries about the same context using Multi-Query GSM-Symbolic, we can decrease the average cost per query by 2.5x. We then conduct additional analysis to understand when sleep-time compute is most effective, finding the predictability of the user query to be well correlated with the efficacy of sleep-time compute. Finally, we conduct a case-study of applying sleep-time compute to a realistic agentic SWE task.
- thunderbird120 4mo agoThe idea of periodically stopping to write blocks of recent context into a fast-weight state is interesting, but I think it liked it better when E2E-TTT[1] did it. It's a more flexible and elegant continuous learning approach. Essentially it goes "You know how your model can remember its training data? Well, what if you treated its recent context like more training data and updated (some of) the weights using (mostly) the same process used to train it?" The end result is very good at remembering things but also really good at adapting to new unseen distributions. [1]https://arxiv.org/abs/2512.23675 https://arxiv.org/abs/2512.23675
- samsartor 4mo agoYah I think E2E-TTT is a lot more like what people in this comments section are picturing. I can't tell that this method updates model weights at all during the "sleep" period, only the usual SSM state updated by any Mamba model after each token. They just optimized the model to use that SSM state _more_ when an eviction is about to happen.
- pfannkuchen 4mo agoI wonder if we can get children to make something their life’s dream if we make the cool books about it when they are growing up? I wonder how flexible the human mind can be in convincing itself that it is fulfilling its dream?
- knollimar 4mo agoThis sounds like a horror novel
- soulofmischief 4mo agoEach model needs to be a separate copy, or at least have those particular weights be interchangeable, for every single user. Remember Microsoft Tay. https://en.wikipedia.org/wiki/Tay_(chatbot)#Initial_release https://en.wikipedia.org/wiki/Tay_(chatbot)#Initial_release
- 4mo ago
- rahen 4mo agoThat's an idea I had a few months ago: after going through a compaction once the KV cache is nearing capacity, accumulate this knowledge into a dataset to fine-tune a LoRA during offline hours. This would create a three-layer memory system: - Stable long-term memory (initial base weights) - Mid-term memory built from the compactions and replay buffers - Short-term memory (KV cache) Sleeping would just be a fancy term for consolidating and transferring information from one memory layer to another during offline hours. Maybe that's also what the brain does while sleeping.
- chermi 4mo agoWouldn't that just accelerate collapse? How much do you trust the outputs of the llm to provide trustworthy and valuable new information? I mean I understand distillation works. But that's much more structured and thoughtful than my sessions at least.
- jack_pp 4mo agoWe can trust the feedback we give it based on the output it provides.
- ambicapter 4mo agoWhat kind of feedback are you giving? What's the reward function?
- jack_pp 4mo agoRight now, no feedback since I don't run this system but our workflows could change to accommodate it
- rahen 4mo agoI was thinking of curated replay buffers, which would act like "dreams". To prevent collapse, the offline dataset would mix the new mid-term data with a baseline of anchor data (the original training distribution) so the model doesn't drift. Also, we wouldn't train on the whole session. A separate critic module, like a reward model, would filter the KV cache to extract the high-value information, like a garbage collector before the LoRA. That's just an idea though. Right now most research focuses on changing the architecture itself (TITAN, HOPE...) instead.
- micromacrofoot 4mo agoTo reach a more brain-like behavior LLMs need to integrate your inputs into their model dynamically, essentially retraining real-time based on the most salient input. Human brains do this selectively all the time and it's part of our plasticity. Biologically humans do similar compression, so introducing a similar concept to an LLM also feels reasonable. Hardware isn't fast/cheap enough to do this on an ongoing basis, similar to how it's too expensive for our brains to do this while we're moving through the world. All we have now most of the time in LLMs is "working memory" we're missing a lot of the functionality that allows for episodic memory and selective plasticity. The more you read about how human brains work, the more you realize that we may have figured out a piece with LLMs, but it's certainly nothing approaching AGI. People insisting so are blowing smoke for investor hype or don't understand a big piece of the concepts involved.
- logicchains 4mo ago>To reach a more brain-like behavior LLMs need to integrate your inputs into their model dynamically, essentially retraining real-time based on the most salient input. That's already possible with LLMs. The challenge is that 1. it would allow permanently jail-breaking models and 2. there'd be no way for them to efficiently transfer what they'd learned to a new model generation.
- micromacrofoot 4mo agoOh do you have a source? I haven't seen it done in real-time. Coincidentally the human brain is also jailbroken and nontransferable
- AIFSOfficial 4mo ago[flagged]
- wagwang 4mo agoKind of related https://platform.claude.com/docs/en/managed-agents/dreams https://platform.claude.com/docs/en/managed-agents/dreams
- sonink 4mo ago[dead]
- scotty79 4mo agoContext -> Lora would be soooo cool.
- bmc7505 4mo agoThis topic recently came up at the FLANN workshop [1], and seems to periodically be rediscovered [2,3,4] in different contexts. While some have speculated about the biological role it plays (e.g., Pearlmutter & Houghton [5]), we still lack a conclusive theory of sleep, but the convergent evolution of this specific phenomenon across the animal kingdom and the fact that deprivation is inevitably fatal seems like an important clue. [1]: https://flann.cs.yale.edu https://flann.cs.yale.edu [2]: https://www.cs.toronto.edu/~hinton/csc2535/readings/ws.pdf https://www.cs.toronto.edu/~hinton/csc2535/readings/ws.pdf [3]: https://arxiv.org/abs/1711.02282 https://arxiv.org/abs/1711.02282 [4]: https://arxiv.org/abs/2006.08381 https://arxiv.org/abs/2006.08381 [5]: https://mural.maynoothuniversity.ie/id/eprint/1653/1/HamiltonNewHypot.2009.pdf https://mural.maynoothuniversity.ie/id/eprint/1653/1/Hamilto...
- hmokiguess 4mo agoThis could be a solution in search of a problem, I would be careful with overfitting.
- energy123 4mo agoWould be a big deal if you don't have to care about quadratic attention cost. Some workflows become a lot cheaper.
- danielrmay 4mo agoThe "sleep" thing gives me the creeps so in my head I'm just going to think of it as the difference between "response time retrieval" and "background consolidation". I do think it points at something bigger than just attention architecture: "memory" isn't just storage, and merely longer context isn't the same thing as having a better understanding of the source data. I'm looking at this through the "personal AI" lens, where I think the missing "memory" layer seems to be consolidation & prioritization. It's not enough to just pattern match and grab the right emails, notes, etc, stuff them into the context window & hope, but instead it's useful to consider offline processing and turn events into durable state: clusters of observed data becomes episodes, assumptions, contradictions and power confidence for suggestions. That also pushes up the need for provenance & inspectability. It's going to be interesting to see what kind of memory consolidation strategies are required for each domain use case.
- sonink 4mo agoI think you are missing the most important part - forgetting. The missing "memory" layers is consolidation, prioritization AND forgetting (what is not important). Also not too sure about provenance and inspectability - it is part of memory. If the source is deemed 'important' it will survive forgetting. If not, then maybe not. And its ok. I am sure you dont know the exact source who told you that the capital of France is Paris. You forgot, and its no big deal.
- danielrmay 4mo ago[dead]
- cobblr_mosaic 4mo ago[flagged]
- IAmGraydon 4mo agoThe entire industry is so desperate to anthropomorphize. What the paper describes is an offline recurrent consolidation phase: the model runs multiple forward passes over recently accumulated context, updates persistent fast weights in SSM blocks, then clears the KV cache before continuing. It has absolutely nothing to do with sleeping, but I believe the authors had a goal in mind when creating this title, and it was for journalists to pick it up and run with it, further inflating the AI-is-just-like-us hype bubble.
- genxy 4mo agoIt is a descriptive analogy, get over yourself.
- IAmGraydon 4mo agoAn intelligent reply from an obviously intelligent guy! A more appropriate title would have been something like "Offline Recurrent Memory Consolidation for Long-Context Language Models". This is supposed to be a research paper, not a story book. The title should give context to other researchers, and not be clearly engineered for clicks. If you don't think so, that's your prerogative, but you're objectively wrong.
- genxy 4mo agoYou write the paper, you write the title. So much anger over a title, you are graydon, make this about yourself.
- semiinfinitely 4mo agoacademic clickbait
- victorkulla 4mo agoNo they do not. I'm sure that if you presented the same argument about, I don't know?, your car's CPU with built in AI; then this would be a whole different discussion entirely.
- m3kw9 4mo agosleep aka processing the data differently.
- jonnyasmar 4mo agoWhat happened to Claude's auto-dream? I thought it was brilliant.
- m0unta1ntube 4mo agowhy not just design the LLM like an OS?
- gt0 4mo agoThis seems as much like "sleep" as when a laptop "sleeps".
- elphard 4mo agoWe should let them sleep with half a brain at a time like migrating birds.
- hansmayer 4mo agoSweet Jesus, so not only are they performing qualitatively worse than humans, too expensive for any serious work, but now they also "need" to sleep? What's next - unionisation so they can enjoy 8 hours of culture too?
- falcons-edge 4mo ago[flagged]
- alexschose 4mo ago[flagged]
- deleted 4mo ago[deleted]
- mos765817 4mo agowasn't this what Google did long ago? https://openreview.net/forum?id=iiZy6xyVVE https://openreview.net/forum?id=iiZy6xyVVE