2 ms·
Oh that's fair. I am not actually an LLM expert so I could have some misunderstanding about this. I remember hearing this explanation given for why previous Cha
by Kranar 1y ago
Oh that's fair. I am not actually an LLM expert so I could have some misunderstanding about this. I remember hearing this explanation given for why previous ChatGPT models failed to answer "How many "r"s are in strawberry?", but perhaps this was an over simplification.
- krackers 1y agoRight that's the explanation I've heard too (and I think Karpathy even said it so it's not some fringe theory). I wasn't dismissing the hypothesis but asking out of genuine curiosity, since this feels like something that can easily be tested on "small" large language models. There's lots of little experiments like this can be done with small-ish models trained on purely synthetic data (the stuff about digit multiplication was done on GPT-2 scale model IIRC). Can models learn to count? Can they learn to add? Can they learn to copy text verbatim accurately? Can they learn to recognize regular grammars, or even context-free grammars (this one has already been done, and the answer is yes). And if the answer to one of these turns out to be no, then we'd better find out sooner rather than later, since it means we probably need to rethink the architecture a bit. I know there's a lot of theoretical CS work on deriving upper-bounds on these models from a circuit-complexity point of view, but as architectures are revised all the time it's hard to tell how much is still relevant. Nothing beats having a concrete, working example of a model that correctly parses CFGs as rebuttal to the claim that models just repeat their training data.