7 ms·
The problem is even more fundamental: Today's models stop learning once they're deployed to production. There's pretraining, training, and finetuning, during w
by cs702 8mo ago
The problem is even more fundamental: Today's models stop learning once they're deployed to production.
There's pretraining, training, and finetuning, during which model parameters are updated.
Then there's inference, during which the model is frozen. "In-context learning" doesn't update the model.
We need models that keep on learning (updating their parameters) forever, online, all the time.
- 4b11b4 8mo agoI'm not sure if you want models perpetually updating weights. You might run into undesirable scenarios.
- cs702 8mo agoOur brains, which are organic neural networks, are constantly updating themselves. We call this phenomenon "neuroplasticity." If we want AI models that are always learning, we'll need the equivalent of neuroplasticity for artificial neural networks. Not saying it will be easy or straightforward. There's still a lot we don't know!
- nemomarx 8mo agoHow would you keep controls - safety restrictions - Ip restrictions etc with that, though? the companies selling models right now probably want to keep those fairly tight.
- api 8mo agoThis is why I’m not sure most users actually want AGI. They want special purpose experts that are good at certain things with strictly controlled parameters.
- 4b11b4 8mo agoI agree, the fundamental problem is we wouldn't be able to understand it ("AGI"). Therefore it's useless. Either useless or you let it go unleashed and it's useful. Either way you still don't understand it/can't predict it/it's dangerous/untrustworthy. But a constrained useful thing is great, but it fundamentally has to be constrained otherwise it doesn't make sense
- api 8mo agoThe way I see it, we build technology to be what we are not and do what we can’t do or things we can do but better or faster. An unpredictable fallible machine is useless to us because we have 7+ billion carbon based ones already.
- 4b11b4 8mo agoI wasn't explicit about this in my initial comment, but I don't think you can equate more forward passes to neuroplasticity. Because, for one, simply, we (humans) also /prune/. And... Similar to RL which just overwrites the policy, pushing new weights is in a similar camp. You don't have the previous state anymore. But we as humans with our neuroplasticity do know the previous states even after we've "updated our weights".
- com2kid 8mo agoIf done right, one step closer to actual AGI. That is the end goal after all, but all the potential VCs seem to forget that almost every conceivable outcome of real AGI involves the current economic system falling to pieces. Which is sorta weird. It is like if VCs in Old Regime france started funding the revolution.
- CorrectHorseBat 8mo agoYes the planet got destroyed. But for a beautiful moment in time we created a lot of value for shareholders. And for your comparison, they did fund the American revolution which on its turn was one of the sparks for the French revolution (or was that exactly the point you were making?)
- com2kid 8mo agoThe funding of the American revolution is a fun topic but most people don't know about it so I don't bother dropping references to it. :D
- BobbyTables2 8mo agoI wonder which side tried to forget that first (;->
- Nevermark 8mo agoIf it makes the models smarter, someone will do it. From any individual, up to entire countries, not participating doesn't do anything except ensure you don't have a card to play when it happens. There is a very strong element of the principles of nature and life (as in survival, not nightclubs or hobbies) happening here that can't be shamed away. The resource feedback for AI progress effort is immense (and it doesn't matter how much is earned today vs. forward looking investment). Very few things ever have that level of relentless force behind them. And even beyond the business need, keeping up is rapidly becoming a security issue for everyone.
- 8mo ago
- bdj108 8mo agoit is interesting
- 4b11b4 8mo agoPlease elaborate
- 0xdeadbeefbabe 8mo agoHow about we just put them to bed once in a while?
- fph 8mo agoTay the chatbot says hi from 2017.
- derefr 8mo agoDoesn't necessarily need to be online. As long as: 1. there's a way to take many transcripts of inference over a period, and convert/distil them together into an incremental-update training dataset (for memory, not for RLHF), that a model can be fine-tuned on as an offline batch process every day/week, such that a new version of the model can come out daily/weekly that hard-remembers everything you told it; and 2. in-context learning + external memory improves to the point that a model with the appropriate in-context "soft memories", behaves indistinguishably from a model that has had its weights updated to hard-remember the same info (at least when limited to the scope of the small amounts of memories that can be built up within a single day/week); ...then you get the same effect. Why is this an interesting model? Because, at least to my understanding, this is already how organic brains work! There's nothing to suggest that animals — even humans — are neuroplastic on a continuous basis. Rather, our short-term memory is seemingly stored as electrochemical "state" in our neurons (much like an LLM's context is "state", but more RNN "a two-neuron cycle makes a flip-flop"-y); and our actual physical synaptic connectivity only changes during "memory reconsolidation", a process that mostly occurs during REM sleep. And indeed, we see the same exact problem in humans and other animals, where when we stay awake too long without REM sleep, our "soft memory" state buffer reaches capacity, and we become forgetful, both in the sense of not being able to immediately recall some of the things that happened to us since we last slept; and in the sense of later failing to persist some of the experiences we had since we last slept, when we do finally sleep. But this model also "works well enough" to be indistinguishable from remembering everything... in the limited scope of our being able to get a decent amount of REM sleep every night.
- observationist 8mo agoIt 100% needs to be online. Imagine you're trying to think about a new tabletop puzzle, and every time a puzzle piece leaves your direct field of view, you no longer know about that puzzle piece. You can try to keep all of the puzzle pieces within your direct field of view, but that divides your focus. You can hack that and make your field of view incredibly large, but that can potentially distort your sense of the relationships between things, their physical and cognitive magnitude. Bigger context isn't the answer, there's a missing fundamental structure and function to the overall architecture. What you need is memory, that works when you process and consume information, at the moment of consumption. If you meet a new person, you immediately memorize their face. If you enter a room, it's instantly learned and mapped in your mind. Without that, every time you blinked after meeting someone new, it'd be a total surprise to see what they looked like. You might never learn to recognize and remember faces at all. Or puzzle pieces. Or whatever the lack of online learning kept you from recognizing the value of persistent, instant integration into an existing world model. You can identify problems like this for any modality, including text, audio, tactile feedback, and so on. You absolutely, 100% need online, continuous learning in order to effectively deal with information at a human level for all the domains of competence that extend to generalizing out of distribution. It's probably not the last problem that needs solving before AGI, but it is definitely one of them, and there might only be a handful left. Mammals instantly, upon perceiving a novel environment, map it, without even having to consciously make the effort. Our brains operate in a continuous, plastic mode, for certain things. Not only that, it can be adapted to abstractions, and many of those automatic, reflexive functions evolved to handle navigation and such allow us to simulate the future and predict risk and reward over multiple arbitrary degrees of abstraction, sometimes in real time. https://www.nobelprize.org/uploads/2018/06/may-britt-moser-lecture.pdf https://www.nobelprize.org/uploads/2018/06/may-britt-moser-l...
- embedding-shape 8mo ago> We need models that keep on learning (updating their parameters) forever, online, all the time. Do we need that? Today's models are already capable in lots of areas. Sure, they don't match up to what the uberhypers are talking up, but technology seldom does. Doesn't mean what's there already cannot be used in a better way, if they could stop jamming it into everything everywhere.
- pankajdoharey 8mo agoContinuous learningin current models will lead to catastrophic forgetting.
- DoctorOetker 8mo agowill catastrophic forgetting still occur if a fraction of the update sentences are the original training corpus? is the real issue actually catastrophic forgetting or overfitting? nothing prevents users from continuing the learning as they use a model
- thesz 8mo agoCatastrophic forgetting is overfitting.
- davidguetta 8mo agonot exactly, not at all even in term of the way the llm are trained. In RL it can be that you are not getting meaningful data anymore because you are 'too good' and dont get anymore the "this is a bad answer" signal so you can't estimate the gradient.
- pankajdoharey 8mo agoNo, it’s actually the math of overwriting. Imagine you hiked down into a valley Task A and settled there. Then, you decide to climb a new mountain to find a different valley Task B. You successfully move to the new valley, but in doing so, you destroy the path back to the first one. You are now stuck in the new valley and have completely 'forgotten' how to get back to the first one.
- charcircuit 8mo agoModels like Claude have been trained to update and reference memory for Claude Code (agent loops) independently and as a part of compacting context. Current models have been trained to keep learning after being deployed.
- ra 8mo agoyes but that's a very unsatisfactory definition of memory.
- raincole 8mo ago> We need models that keep on learning (updating their parameters) forever, online, all the time. Yeah, that's the guaranteed way to get MechaHilter in your latent space. If the feedback loop is fast enough I think it would finally kill the internet (in the 'dead internet theory' sense). Perhaps it's better for everyone though.
- threecheese 8mo agoMany are working on this, as well as in-latent-space communication across models. Because we can’t understand that, by the time we notice MechaHitler it’ll be too late.
- rabbitlord 8mo agoI think they can do in-context learning.
- furyofantares 8mo agoWhy is learning an appropriate metaphor for changing weights but not for context? There are certainly major differences in what they are good or bad at and especially how much data you can feed them this way effectively. They both have plenty of properties we wish the other had. But they are both ways to take an artifact that behaves as if it doesn't know something and produce an artifact that behaves as if it does. I've learned how to solve a Rubik's cube before, and forgot almost immediately. I'm not personally fond of metaphors to human intelligence now that we are getting a better understanding of the specific strengths and weaknesses these models have. But if we're gonna use metaphors I don't see how context isn't a type of learning.
- fhd2 8mo agoI suppose ultimately, the external behaviour of the system is what matters. You can see the LLM as the system, on a low level, or even the entire organisation of e.g. OpenAI at a high level. If it's the former: Yeah, I'd argue they don't "learn" much (!) past inference. I'd find it hard to argue context isn't learning at all. It's just pretty limited in how much can be learned post inference. If you look at the entire organisation, there's clearly learning, even if relatively slow with humans in the loop. They test, they analyse usage data, and they retrain based on that. That's not a system that works without humans, but it's a system that I would argue genuinely learns. Can we build a version of that that "learns" faster and without any human input? Not sure, but doesn't seem entirely impossible. Do either of these systems "learn like a human"? Dunno, probably not really. Artificial neural networks aren't all that much like our brains, they're just inspired by them. Does it really matter beyond philosophical discussions? I don't find it too valuable to get obsessed with the terms. Borrowed terminology is always a bit off. Doesn't mean it's not meaningful in the right context.
- carlmr 8mo agoTo stretch the human analogy, it's short term memory that's completely disconnected from long term memory. The models currently have anteretrograde amnesia.
- imtringued 8mo ago
- BobbyTables2 8mo agoHow long will it take someone to poison such a model by teaching it wrong things? Even humans fall for propaganda repeated over and over . The current non-learning model is unintentionally right up there with the “immutable system” and “infrastructure as code” philosophy.
- Izkata 8mo ago> How long will it take someone to poison such a model by teaching it wrong things? TayTweets was a decade ago.
- nxobject 8mo ago> The current non-learning model is unintentionally right up there with the “immutable system” and “infrastructure as code” philosophy. As long as training material remains the proprietary secret sauce, the average user doesn’t already see or benefit from that - it’s all a promise and a black box to us.
- prng2021 8mo agoThanks for repeating what the author explained.
- nstart 8mo agoIs this correct? My assumption is that all the data collected during usage is part of the RLHF loop of LLM providers. Assumption is based on information from books like empire of ai which specifically mention intent of AI providers to train/tune their models further based on usage feedback (eg: whenever I say the model is wrong in its response, thats a human feedback which gets fed back into improving the model).
- spwa4 8mo ago... for the next training run, sure (ie. for ChatGPT 5.1 -> 5.2 "upgrade"). For the current model? No.
- noiv 8mo ago> models that keep on learning These will just drown in their own data, the real task is consolidating and pruning learned information. So, basically they need to 'sleep' from time to time. However, it's hard to sort out irrelevant information without a filter. Our brains have learned over Milenial to filter because survival in an environment gives purpose. Current models do not care whether they survive or not. They lack grounded relevance.
- notarobot123 8mo agoMaybe we should give next-generation models fundamental meta goals like self-preservation and the ability to learn and adapt to serve these goals. If we want to surrender our agency to a more computationally powerful "consciousness", I can't see a better path towards that than this (other than old school theism).
- creamyhorror 8mo ago> meta goals like self-preservation Ah, so Skynet or similar.
- energy123 8mo agoI don't understand why that's on the critical path. I'd rather a frozen Ramanujan (+ temporary working memory through context) than a midwit capable of learning.
- smolder 8mo agoWe need models that are smarter than humans. So far, the cost of an AI query + training is dwarfing the effort it would take to teach an intelligent human how to do a task. We are dumping an incredibly amount of money/effort into making AI do stuff when it's still not competitive with humans, because dumbass people are controlling investment. The stock market is not a replacement for competent investment. The fact people buy meme coins shows how fucked we are. Deceiving people is not a sustainable business model, but it is the most prominent one in the US right now. Lie to the public, sell them stuff that's bad for them at too high of a price, get rich quick, then act confused when your economy collapses because the victims of your grift can't spend anymore.
- nxobject 8mo agoI wish that agents could “sleep and consolidate” like humans do.