9 ms·
Bridging empirical-theoretical gap in neural network formal language learning
- schiffern 3y ago[edit: title changed]
- tines 3y ago> However, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum. The paper is about how the "default" training doesn't work, not how neural networks "can't learn" something.
- richrichie 3y agoIndeed. I read it as what loss function works better for this use case.
- gibsonf1 3y agoThe issue is that language comes after we form the concepts mentally that the words refer to from our space-time experience. So the language itself is just the token used to symbolize a concept or entity etc that is mentally modeled by the speaker. So the tokens themselves give you no actual knowledge of the world without being able to convert those tokens into the mental objects they represent and then think about them. This is why LLM tech will never improve as its just statistics on serialized symbols, rather than about the world model those symbols represent.
- tines 3y agoThis isn't what the paper gives evidence for. They found that if you use a different function in the training phase, the networks learn formal languages perfectly well. It's about the capabilities of a particular training process, not about the capabilities of the neural networks themselves.
- gibsonf1 3y agoI'm not sure what you mean by the system learning "perfectly well" as it literally knows nothing about what those words mean. This is why LLMs can give statistical word results that can be true of false or in between and those systems have no way of knowing which status the result has as the system has no understanding - that is, it does not model what the words mean as we humans do to understand if something is true or false etc.
- yorwba 3y agoThe words in the article look like #aaaaaaabbbbbbb#, they do not mean anything.
- naasking 3y ago> as it literally knows nothing about what those words mean Speculation. We don't know what it means, mechanistically, to know what something means. Knowing what something means could consist of having a model of how a word relates to other things, in which case the system does indeed know something about what those words mean. It doesn't have precisely the same meaning as humans, because humans also relate words to other sense data, but there's still meaning.
- gibsonf1 3y agoSure we do. When you read a sentence, and are able to model that in your mind and understand the space-time being referred to, you understand it. LLM is like a calculator in the sense that given an input it gives an output. There is no modeling of space-time in the mind to know if the sentence makes sense etc. That's why if you read a language you don't understand, you get nothing as you can't link those words with the concepts in your mind to understand it. For the LLM, all languages are like that as it understands nothing.
- naasking 3y ago> For the LLM, all languages are like that as it understands nothing. Again, that's too far. An LLM will not assert that "married men are bachelors" because it's been trained that "bachelor" is not associated with "married". Even if it doesn't understand what a bachelor is or marriage is, it understands that these two concepts are not positively associated with each other, and in fact, are negatively correlated. Logical associations like this are semantic content, which thus exhibits some degree of understanding. That's why LLMs can make sense, even when they're lacking the full understanding of a human.
- mcyc 3y agoIf you are interested in this field (exploring the limits of neural models using formal language theory), I help run a weekly seminar on it, Formal Languages and Neural Networks: https://flann.super.site/ https://flann.super.site/ We have had many great speakers (most of them are recorded and available on YouTube) and have a welcoming Discord.
- 33a 3y agoAnd yet, ChatGPT can generate these strings. Somehow despite using the wrong loss function it still seems to work by simply absorbing more training data. https://chat.openai.com/share/82509815-d418-43bb-95a3-348bd5c700a4 https://chat.openai.com/share/82509815-d418-43bb-95a3-348bd5... It can also recognize them, albeit it tends to cheat by shelling out to python (which makes sense, since it tends to lose count on large strings just like a human...) https://chat.openai.com/share/b106ca5f-409a-43db-bc02-21da86ab40ce https://chat.openai.com/share/b106ca5f-409a-43db-bc02-21da86...
- toxik 3y agoArguably the skill to generate a program to do this is a higher-order skill and much more impressive.
- yorwba 3y agoChatGPT incorrectly put a space between the as and bs, but if we let that slide, there's still the issue that the best trained model in the article got 77.3% of the first 1500 strings correct, i.e. even if ChatGPT performed exactly the same, you'd expect it to get a single example correct more often than not.
- canjobear 3y agoThe difference is that you told it the language you wanted to recognize. In the paper, they are trying to learn the language from example strings alone.
- Majromax 3y agoTo summarize the article: using backpropagation to train a small LSTM to accept strings of the form (a^n b^n), regularization with both L1 and L2 losses result in loss spaces where the loss-optimal solution is not the generally-correct one. The resulting networks fail out-of-sample. The LSTM architecture does permit a correct, 'general' solution to exist, and the authors show that the general solution is an optimum when using a minimum-description-length error function. The MDL error function is effectively entropy(weights) + relative entropy(training data | model). The researcher's problem is that entropy(weights) is not well-defined. To prevent the network from "smuggling" information through highly precise but otherwise random-looking weights, they define an entropy (coding length) that rewards simple rational fractions (1/2, 3/4, 7/11) and penalizes complex ones (715/937). With this loss function, their hand-crafted optimal solution also lies at a loss-optimum. Unfortunately, the resulting loss space is non-differentiable, making it useless for any gradient-descent-like training. To muse about this, the authors' problem of weight entropy here is similar to the problem faced by autoencoders, whereby they want to have a minimally-structured latent space to avoid the same information-smuggling problem. Perhaps we could evolve a differentiable version of the model-entropy loss function by treating the model weights as random variables, drawn from a distribution of learned mean and variance with the same 'reparameterization trick' that makes variational autoencoders work. The model entropy loss term is then the same as the latent-space regularization term in VAE training.
- eli_gottlieb 3y ago>The LSTM architecture does permit a correct, 'general' solution to exist, and the authors show that the general solution is an optimum when using a minimum-description-length error function. The MDL error function is effectively entropy(weights) + relative entropy(training data | model). >The researcher's problem is that entropy(weights) is not well-defined. I never thought I would say this, but it sounds like a job for Bayesian neural nets. The prior or posterior entropy of the weights should be well-defined in that setting.
- deleted 3y ago[deleted]
- gwern 3y agoThe RNN in question is over-parameterized, so most loss functions wouldn't be expected to coincide with MDL (unless maybe you added in heavier regularization than just L1/L2). There is no incentive towards simplicity. But it is also a single-task network, and it is increasingly rare to train any of those. A multi-task network must share its weights across all the tasks, and so the more tasks it does for a particular parameterization, the more incentive it has towards simplicity in order to do the tasks at all. So I wonder if scaling up task diversity would increasingly approximate MDL and presumably help it learn the true algorithmic cores?