3 ms·
The underlying LLMs do, but we choose not to use the capability because it's expensive and doesn't quite work as well as we'd like it to, or quite in the way th
by jephs 1mo ago
The underlying LLMs do, but we choose not to use the capability because it's expensive and doesn't quite work as well as we'd like it to, or quite in the way that we'd like it to.
We are perfectly capable of running LLMs in a way that does a backward pass to update some or all of its weights after every user message. But, naively implemented, you only get partial, fragmentary absorption of the info in those messages, it costs three times as much compute, and you lose out on the ability to implement a ton of optimizations that making modern LLM serving economical.
If you want to do it, though, ask your friendly neighborhood robot to get it working with a tiny model (whose full precision weights fit several-times-over on your machine's resources).
- brookst 1mo agoWhile I agree it’s theoretically possible, do you believe this capability exists in the LLMs we use today?
- wongarsu 1mo agoDepends on whether you define "the LLMs we use" as the collection of weights or if your definition contains the software stack that runs it Technologically the LLMs we use today don't implement this behavior, but you could take the weights of Sol and add a couple (very large) patches to vllm (or whatever OpenAI has today) and have a version of Sol that does have "memory"
- c-hendricks 1mo ago"Dave constructs a homemade megaphone using only some string, a squirrel, and a megaphone"
- jephs 1mo agoThere's a ton of experimentation on it, the field is called continual learning. It's not something you need to believe in like Jesus, you can just go read about the current state of things.