3 ms·
I see these takes a lot, but I struggle to take them seriously. As a trivial example: LLMs don't learn new skills on the fly. A key aspect of human behavior--i
by md_ 3y ago
I see these takes a lot, but I struggle to take them seriously.
As a trivial example: LLMs don't learn new skills on the fly. A key aspect of human behavior--implicit memory--simply doesn't exist for these models.
That seems like a pretty huge gap! Like, if ChatGPT is mediocre at your job today, it's going to be mediocre at it tomorrow and the next day, until a new model is trained!
- pixl97 3y agoWith current computing power limits training a new model takes a long time. Seems like months at this point. But lets say Nvidia pulls a rabbit out of its magic hat, and comes out with a training chip that is a million times more powerful. Now instead of months training that drops down to 24 hours. Would you still say that is a pretty huge gap? And I ask this because this is Nvidia's goal within a decade. Where, for you, does a big gap turn into 'that is dangerously close'?
- md_ 3y agoI don't think the difference between human memory and retraining a DNN is purely about the latency to retraining. They seem to be fundamentally different things, no?
- pixl97 3y agoWe would both agree that birds and planes are fundamentally different things, but we would both agree that they fly, correct? Both humans and LLMs have 'short term' and 'long term' memory. The process of turning short term to long term is significantly different. In LLMs this is taking the history data given to it and putting it in the training corpus then recalculating the weights. In humans this involves falling to sleep for some number of hours. I think in the long term here AI will actually have the benefit of forming long term memories. You have to sleep to form them, it just has to run reweighting on another cluster while the primary cluster goes about its business.
- md_ 3y agoWell, my point was that explicit memory and "skill development" (as one kind of what is known as "implicit memory") are fairly different. No matter how many times you retrain GPT4, it isn't going to learn to drive a car.
- pixl97 3y agoThere are a few problems with your argument here.... You're using just text of GPT-4 and have not used any of the multi-modal capabilities. Trying to assess the total capability of a system when exposed to only one mode of a system is difficult if not impossible. I would assume that total compute capacity and cost of running are the biggest inhibitors of this being pushed out to more (any?) users. Now lets imagine GPT-4 with image input, and command output. Tell it that it has an accelerator and a brake. Now feed it an image of a wide open road with no obstructions. My assumption would be that it would reason that it can accelerate. If you fed it a second image of a wide open road, would it assume that it can maintain its speed? Then a third image that shows a child in the road some distance in front of it. The system that already exists already has the reasoning capability that it should stop or it would hit a child. Now you have the problem of "Is this learning to drive a car?". Via just image and text alone it's never going to get very good as we need some more sensory and environmental data passed to the system. But here's the thing about LLMs... Just about any structured information can be considered a language. Passing simulated driving data to an LLM could very likely teach it how to drive.