4 ms·
Right, as a human, I would presumably learn French for a variety of reasons, including overlap of vocabulary, a familiarity with French loan-words in English, a
by pwinnski 4y ago
Right, as a human, I would presumably learn French for a variety of reasons, including overlap of vocabulary, a familiarity with French loan-words in English, and a very human general inability to memorize or even think in numbers well.
This is the mistake people keep making, literally anthropomorphizing LLMs! Because humans would find it easier to just learn the meanings of French words than regurgitate from a gigantic table of memoized tokens, these results must reflect understanding!
But of course, it's far, far easier for a computer to index a large number of tokens without any understanding at all than it would be for a person. Computers are not people. They have different strengths and weaknesses, and it is a mistake to forget that. (Ironically, one of our human weaknesses is also a strength: our pattern-matching skills are so strong we sometimes see Jesus in toast, or a face on Mars, when neither are actually there!)
What I described is literally how LLMs works. There is open source code to examine, and what it does is essentially what I said: tokenize words and store multi-word strings and identify their prevalence in training data, where the training data is an incredibly large corpus.
The experiment is about finding one missing word in a sentence, but the same logic is true even for long responses, or as I suggested, conversations.
To over-simplify: humans think in words, computers think in math. But it turns out that if you assigned a number to each word and parse enough of them, you can fake an understanding of words surprisingly well, even though it's still math.
I've given example elsewhere in this thread and on previous ones about the March 14 Chat-GPT giving incontrovertibly false information, even repeatedly in the face of correction, or seeming to correctly identify a popular logic puzzle while breaking the rules of the puzzle three times, then seeming to claim it had followed the rules of the puzzle.
> So Consider that as an existence proof that even if you "train to predict" you can still end up with actual understanding.
As a human, that is often true. Humans have trouble with arbitrary data, and it helps us to connect new data to existing data. In addition, we generally have a curiosity that prompts us to pursue new knowledge and meaning wherever possible, even when it doesn't actually exist. But if you're trying to suggest that because it is possible that humans can develop understanding while trying to only "train to predict," therefore it is probable that computers will develop understanding along the same lines, I suggest that's about three leaps too far.
- dilap 4y agoI think with sufficiently good technology the human brain, too, could be modeled with just math. So far, it seems like basically everything in the universe can be described as just math! ChatGPT is a neural network, a design inspired (loosely) by the human brain. How exactly it works is still not known. Like of course the raw details of how the computations work is known, but why the particular trained weights are what they are, and how the thinking actually happens is not known (so far as I know). But I think it must be more than just something like looking up a conditional probability table -- calculate out the size of a conditional probabiltty of 8000 tokens or whatever and it'll be way way way more than the number of parameters of the model. And no doubt as it stands ChatGPT is still far inferior to human intellect. But I would note that making mistakes doesn't mean something like thinking isn't happening -- humans make mistakes all the time, often very dumb ones!