8 ms·
What's a large language model doing when it's not being queried? Am I correct that they only compute information when dealing with a prompt? If so, that seems
by optimalsolver 3y ago
What's a large language model doing when it's not being queried?
Am I correct that they only compute information when dealing with a prompt? If so, that seems like a fundamental flaw. An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output.
- iinnPP 3y agoHooking an LLM up to a loop would solve that. Then you can find a way to include a described video feed and method of movement into the mix.
- TheRoque 3y agoNot really, because it would hit a moment where it runs out of context. It can't really learn anything for now.
- gmerc 3y agoNot really, hook up external sensors to keep shoveling data , feedback data into continuous lora and performance comparison,
- danaris 3y agoHow? What kind of data? You think you can just hook up a firehose of visual or auditory data to ChatGPT and have that produce anything meaningful? That's not even the basic format that ChatGPT operates in. Furthermore, isn't part of the point of the dataset an LLM is trained on that it has to be at least moderately structured and tagged, or you end up with garbage in the output?
- deleted 3y ago[deleted]
- Zambyte 3y agoAlso, based on its continuous experience, it would be able to prompt you (send multiple messages in a row after not getting a response) or it should be able to wait for you to send multiple messages before responding.
- Aerroon 3y ago>it would be able to prompt you (send multiple messages in a row after not getting a response) This already happens. When you play around with local LLMs you'll run into the situation where the model answers your prompt. After that it will generate a new query as a response to it and then reply to that too.
- visarga 3y agoNo, unfortunately it doesn't learn online. It forgets everything after each interaction. They can collect the data and retrain later, but the hard part is doing all the fine-tuning steps all over again and ensuring the new model has no regressions. GPT3.5 and 4 are years out of date, an unfortunate situation when generating code or asking about recent events. And now they removed the search plugin, probably they got sued for copyright leaks from the search engine results into the generated text. Using copyrighted data in the prompt is not necessarily legal. So we have to deal with out-of-date AI that updates once every couple of years.
- jasfi 3y agoThe cost to compute these language models should eventually lower. Will OpenAI then release more frequently, or release larger LLMs? They'll probably to try achieve both goals in some proportion.
- ben_w 3y agoTraining on web pages created after ChatGPT (and equivalently Stable Diffusion) was released, has a strong risk of trying to create the next AI (of whatever type) on the barely-even-proofread output of the current models.
- rtkwe 3y agoA big issue with constant retraining is going to be the self referential consumption of it's own generated material as training material. There's already been the studies that these models quickly break down when they're fed their own generated data as training data. It won't immediately degrade them as it's a small percentage but there's already a lot of people using these models to generated spam junk out there and it'll only get worse with time.
- visarga 3y ago> A big issue with constant retraining is going to be the self referential consumption of it's own generated material as training material That's just a hypothetical scenario. Usually LLMs don't generate text used indiscriminately to train new models. In creative tasks, the generated texts are the result of multiple explorations and corrections, filtered by the user, then posted online and commented on by other people. So there is plenty of correction and feedback mixed into the process. A child model could learn something from this feedback. Then there is code and math, where it is possible to run tests or double check by a different method. These tasks could be iteratively improved by LLMs by incorporating the feedback. Another field ripe with feedback signals is using LLMs in games and simulations, where you can specify a goal easily but finding a solution requires search and learning. And a related application was RLHF, where a preference model was trained on human judgments and used to fine-tune a LLM. That preference model can be used to filter bad training data as well. To generalize, LLMs alone can't do it, but LLMs with external systems can get feedback and improve. The garbage-in-garbage-out scenario can be avoided with a bit of external feedback.
- danaris 3y agoThank you. This is one of the questions far too few people seem to be paying attention to. "Thinking" in any way that we truly understand the term requires consciousness, and consciousness requires much more continuity than LLMs have. It would need continuity of input as well as continuity of learning in order to even be able to begin to approach something we might recognize as consciousness.
- hackinthebochs 3y agoWhy?
- danaris 3y agoWell, "consciousness", at least as we recognize it, requires a mechanism by which the entity being measured can continuously form new "thoughts" and "memories" (which requires continuity of learning, and at the very least continuity of input being fed back from its own output), and some form of continuous external input of information about the world to at least be available, even if it is not always on. A standard LLM is a static bundle of trained data that sits, inert, on a drive, with a process waiting for discrete input. When that input arrives, it does nothing to modify the trained data—the LLM's "memory"—it simply triggers a computational process that reads both the input and the trained data and produces an output based on them. This does not resemble in any way the structure of something that could be reasonably described as a conscious mind.
- hackinthebochs 3y agoContinuous processing may be putting undue weight on accidental features of humans to be a necessary feature. Consider that a conscious mind can't represent the gaps in its processing, and so has an appearance of continuity. But this appearance probably doesn't map onto a continuous reality. An example is anesthesia patients that report a seemingly uninterrupted time from counting down pre-surgery to waking up. So interruptions, gaps, discontinuities, and so on don't necessarily eliminate the possibility for consciousness. It may be the case that LLMs are conscious when they are engaged in active inference. While I generally favor a requirement for recurrent processing, I have low but non-zero credence for certain feedforward networks being conscious. The point of recurrence is to allow information about itself to influence its processing. But It seems plausible that feedforward constructs can represent meta-information in a way that is computed as part of constructing the output.
- ben_w 3y ago> What's a large language model doing when it's not being queried? Nothing. > Am I correct that they only compute information when dealing with a prompt? Yes. > If so, that seems like a fundamental flaw. Flawed in what way? It clearly doesn't need to be like us to be useful, because it's useful and definitely not like us. > An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output. This might be good, but it's not clear if, or to what extent, we really do that ourselves — the differences between working/short/long term memories, between episodic and skill, even linguistically between knowledge of phonemes, words, grammar, and the connection between those and the things they represent all being impaired independently of each other by localised brain damage[0]. Then there's how much this changes with some stages of sleep, and meditation to clear your mind. Given the number of users (what is it, 100 million?), having it always on, continuously integrating, would still be inhuman even if the architecture was a perfect mirror of the human brain. Also, if the AI is structured to be a "thinking machine", does that make it murder to switch it off? [0] Cognitive Psychology for Dummies, currently listening to it as an audiobook.
- danaris 3y ago> It clearly doesn't need to be like us to be useful, because it's useful and definitely not like us. For many people, the question is not—and has never been—"is it useful?", but "is it conscious?", or "when will it bring the singularity?" Yes, it is useful as it is. But it is not conscious, and if it is a step closer to the still-very-hypothetical singularity, then it is not a particularly large one, because of these very real fundamental limitations. But engaging with people who are clearly talking about whether AI is conscious by responding as if they're talking about whether it is useful could very easily be seen as engaging in bad faith.
- ben_w 3y ago> But it is not conscious How can we tell? We don't have a single working definition of what "consciousness" is, rather we have at least half a dozen wildly different definitions, some of which aren't testable[0], some of which include VHS players[1], some of which exclude many humans most of the time[2]. If consciousness exists purely in working memory and not short/long term, then LLMs may have it within any given context window, even if they're effectively amnesiac between contexts. Possibly. Like I said, we don't know what consciousness is. > But engaging with people who are clearly talking about whether AI is conscious by responding as if they're talking about whether it is useful could very easily be seen as engaging in bad faith. It may seem clear to you, but it didn't to me; furthermore, the singularity doesn't seem to me to require any of the many definitions of consciousness that I'm aware of, merely the capabilities for problem solving that LLMs increasingly demonstrate regardless of if they're conscious or not. [0] what even is qualia [1] can record sensory experiences in memory and change responses based on this [2] rational/logical reasoning
- gmerc 3y agoOur brain is made of interconnected systems but somehow expect LLM architecture to encompass the whole spectrum. Nothing stops you from running a loop that involves other systems such as long term memory (vector /dev storage), visual pre-processor (CNN), auto lora, and more. That’s the fundamental flaw with most of the criticism - the tech is out only a few short months in the hands of everyone. The disruption will come from plugging it into feedback systems.