4 ms·
When I infer I also train.
by willis936 1mo ago
When I infer I also train.
- deleted 1mo ago[deleted]
- oblio 1mo agoBrilliant. We keep pretending LLMs learn. No, they're smart idiots/stupid geniuses. They're turn based intelligence in a real time world. 1. Each time someone talks to me they don't have to repeat the entire conversation from the beginning with each reply. 2. If my boss/partner/whoever gives me some mandates/orders (basically), I don't just forget about them because they were at the beginning of the conversation. 3. If during the conversation I access external data sources to get new info or refresh stale info (a presentation, a book, whatever), I don't instantly forget about it after the conversation ends and forget to incorporate this information if 10 000 other people ask me again. 4. I verify new inputs/lessons against my core principles. 5. I protect myself/ignore requests if new inputs/lessons contradict my core principles. 6. Etc, etc.
- scotty79 1mo ago> If my boss/partner/whoever gives me some mandates/orders (basically), I don't just forget about them because they were at the beginning of the conversation. Your context window is 80 years. You are forgetting plenty before you reach the end of it.
- 10xDev 1mo agoYou generally forget things you don't retrieve. It is not really a bug but a way to declutter for efficiency. That's not the same as it not fitting your context window because it was at the beginning.
- scotty79 1mo agoIt's not about not fitting in context window. LLMs also can "forget" the things from their early context window that they do not restate later. It's also a form of decluttering. You can't (and shouldn't) remember (pay attention to) old stuff with the same priority as new, more relevant stuff.
- bcrosby95 1mo agoYeah, I already don't remember what I had for breakfast two days ago.
- oblio 1mo agoI forget that, too. But if my partner tells me they're allergic to shellfish, I'm not going to order oysters for them tonight. See the difference?
- cheschire 1mo agoIsn’t the stateless nature of chat models a contrived method for scalability? I’m pretty sure that’s why so many in-the-know people have been saying we have achieved AGI already. Not just sama’s contract-breaking tactics of late. I’m referring to all the really intelligent folks who have been crying doomsday scenarios for modern society for the last couple years. What I’m getting at is that the toolset we get exposed to is not what’s available in the labs. This stateless method of managing chat context is just how we are allowed to interact with it.
- Bilal_io 1mo ago> the toolset we get exposed to is not what’s available in the labs Do you have a source for this?
- ozgung 1mo agoThis is called Test Time Learning and some research architectures can do that. Current Mainstream models may not do that because their design is mostly about scalability. They have to serve millions of people with low latency. Alternatively they could design and run a single super-intelligent model, with no scalability constraints. Probably whey are already doing that as well.