4 ms·
> it's simply not possible to "identify the patterns" without understanding the patterns I don't think that's true at all. To explain my thinking, let's assume
by pwinnski 4y ago
> it's simply not possible to "identify the patterns" without understanding the patterns
I don't think that's true at all. To explain my thinking, let's assume that I am an American, and therefore according to popular understanding, completely incapable of learning any language other than English. I despair! But then I think no, wait, I happen to have a gift for numbers. Like, my mental arithmetic stunned teachers from a young age, and I currently hold the world record for reciting digits of pi from memory. So what I will do, I think, is assign a number to every word I read or hear and where possible, visualize where that chain of numbers exists among the decimal digits of pi.
At some point the person would come in and utter some words I couldn't comprehend, but I would think, aha! That was almost 498153217845632754, except it was missing the second seven, which mapped to "contraire." And I'd do that at least 20% of the time, escaping the room with no understanding of French at all.
If I could map between memorized numbers and words quickly enough, I'd even be able to converse in French given suitable trigger words without a shred of understanding.
- dilap 4y agoBut if you actually did the experiment, that's not what would happen -- what would happen is you would learn French! Because that's easy for a human, and the mathematical approach you describe is something like impossible. Even for a machine, the mathematical approach is going to be impossible in practical terms as soon as input gets longer than a few sentences. To build any sort of probability table conditional on more than just a small number of tokens becomes astronomically big. The intuition I'm trying (not very successfully) to get across is if you did this with a human, the human would end up learning French. So Consider that as an existence proof that even if you "train to predict" you can still end up with actual understanding. Beyond that, (and while this is less certain), probably to be able to predict well w/ any accuracy for texts of non-trivial length, you have to be able to understand; probably this is the only way to "compress" the insanely large conditional probabilities.
- pwinnski 4y agoRight, as a human, I would presumably learn French for a variety of reasons, including overlap of vocabulary, a familiarity with French loan-words in English, and a very human general inability to memorize or even think in numbers well. This is the mistake people keep making, literally anthropomorphizing LLMs! Because humans would find it easier to just learn the meanings of French words than regurgitate from a gigantic table of memoized tokens, these results must reflect understanding! But of course, it's far, far easier for a computer to index a large number of tokens without any understanding at all than it would be for a person. Computers are not people. They have different strengths and weaknesses, and it is a mistake to forget that. (Ironically, one of our human weaknesses is also a strength: our pattern-matching skills are so strong we sometimes see Jesus in toast, or a face on Mars, when neither are actually there!) What I described is literally how LLMs works. There is open source code to examine, and what it does is essentially what I said: tokenize words and store multi-word strings and identify their prevalence in training data, where the training data is an incredibly large corpus. The experiment is about finding one missing word in a sentence, but the same logic is true even for long responses, or as I suggested, conversations. To over-simplify: humans think in words, computers think in math. But it turns out that if you assigned a number to each word and parse enough of them, you can fake an understanding of words surprisingly well, even though it's still math. I've given example elsewhere in this thread and on previous ones about the March 14 Chat-GPT giving incontrovertibly false information, even repeatedly in the face of correction, or seeming to correctly identify a popular logic puzzle while breaking the rules of the puzzle three times, then seeming to claim it had followed the rules of the puzzle. > So Consider that as an existence proof that even if you "train to predict" you can still end up with actual understanding. As a human, that is often true. Humans have trouble with arbitrary data, and it helps us to connect new data to existing data. In addition, we generally have a curiosity that prompts us to pursue new knowledge and meaning wherever possible, even when it doesn't actually exist. But if you're trying to suggest that because it is possible that humans can develop understanding while trying to only "train to predict," therefore it is probable that computers will develop understanding along the same lines, I suggest that's about three leaps too far.
- dilap 4y agoI think with sufficiently good technology the human brain, too, could be modeled with just math. So far, it seems like basically everything in the universe can be described as just math! ChatGPT is a neural network, a design inspired (loosely) by the human brain. How exactly it works is still not known. Like of course the raw details of how the computations work is known, but why the particular trained weights are what they are, and how the thinking actually happens is not known (so far as I know). But I think it must be more than just something like looking up a conditional probability table -- calculate out the size of a conditional probabiltty of 8000 tokens or whatever and it'll be way way way more than the number of parameters of the model. And no doubt as it stands ChatGPT is still far inferior to human intellect. But I would note that making mistakes doesn't mean something like thinking isn't happening -- humans make mistakes all the time, often very dumb ones!