6 ms·
Why would you assume that years of exposure isn't terabytes of data? At least we can train language models in hours or days rather than years. If you view it t
by mpoteat 6y ago
Why would you assume that years of exposure isn't terabytes of data?
At least we can train language models in hours or days rather than years. If you view it that way we've made quite a bit of progress!
- dlkf 6y agoYour first point is spot-on. Neural nets aren't unrealistic in virtue of having too much data, the problem is they have too little. Your second point doesn't make sense though. Language models don't understand language in the way that humans do. They just maximize the probability of the next word given some corpus. This is completely different than what humans do.
- andrepd 6y agoGPT-3 certainly was fed far more data than any human needs to start talking.
- dlkf 6y agoSo your claim is that if a human child was given a subset of the text GPT-3 was fed — with no audio, video, or reinforcement feedback from its environment — the child could learn English? The important point to observe here is that what a kid does is fundamentally different than what GPT-3 does: - GPT-3 learns the word that minimizes perplexity given some context - A kid learns the word that helps them accomplish some task in some environment. In the learning process, children get feedback from their environment - including the responses of other agents in the environment (eg Mom, Dad, friends, etc). This is going to be far more memory-intensive than some text files. Also important to observe is that the child's experience can't be reduced to the audio, video and haptic streams they get, because the audio and video stream depend on their own actions. So you need the conditional statements that were embedded in the environment the kid was learning in. All of which is to say this is a little heavier than ascii. Edit: From your other comments, it seems we agree that bigger neural networks with more training data are not going to yield some quantum leap where GPT-3 starts talking. I'm not at all suggesting that throwing more data at the existing architectures is a fruitful direction for understanding cognition. (In general I think NLP is utterly pointless as regards understanding cognition, and if that's your goal you should focus on vision and reinforcement learning.) But my point is that we can't expect that a smarter architecture on a smaller corpus will get us anywhere. If we ever develop machines that are intelligent in some robust sense of the word, their "training data" will most likely be a physical environment with other agents in it.
- earthboundkid 6y agoIf the problem with GPT-3 is that it’s just being fed non-interactive data, why are robots still so primitive?
- deleted 6y ago[deleted]
- dlkf 6y ago> If the problem with GPT-3 is that it’s just being fed non-interactive data, why are robots still so primitive? I never said we understand the right architectures / algorithms that are necessary for robust machine intelligence. In fact I explicitly said that we don't (which answers your straw-man question).
- wodenokoto 6y agoI think the argument is that gpt-3 has read more text than any human possibly could in their entire life, and as parent said, children needs way less input that an entire life worth of sentence in order to start creating novel sentences that makes sense.
- andrepd 6y agoGPT-3 was trained over months using a huge cluster of computers far more powerful than a child's brain (if such a comparison is even meaningful). It is still laughably bad at what it does, which is completing simple sentences.