4 ms·
I was hoping for this take—-of course LLMs have an internal model, and we can externally verify how accurate it is statistically. And the more you train it the
by vladf 4y ago
I was hoping for this take—-of course LLMs have an internal model, and we can externally verify how accurate it is statistically. And the more you train it the better it gets. And for discrete spaces you might even get a perfect model eventually with nonzero probability.
Since you mentioned different layers, one question to ask is if your poetry example is “just a combination of existing rare patterns,” one that’d be addressed by a bigger model with more data.
And as you say the integration of new information does not seem to rely on any fundamentally new principles, just more engineering work.
I think then the question to ask would be, well, how much more data to x% more faithful of the rules of Othello?
This rate of learning matters, and can be the difference between “in our lifetime” and “never in hundreds of years”. As ML practitioners our jobs are often to improve these rates through better formulation of the learning problem.
I think there’s a connection here that amounts to “how quickly you learn”, and it has implications on causality: https://vladfeinberg.com/2019/12/01/metaphysics-of-causality.html https://vladfeinberg.com/2019/12/01/metaphysics-of-causality...