4 ms·
Yeah it's an interesting concept. In the ML community recently there has been some informal chatter along similar lines. Essentially, you can prove that by mode
by thunderbird120 5y ago
Yeah it's an interesting concept. In the ML community recently there has been some informal chatter along similar lines. Essentially, you can prove that by modeling the probability of every element in a dataset given every other set of elements, what you are learning is the core underlying structure which defines the data you are modeling. A perfect model of this structure represents the most efficient possible representation of the information. This is kind of like how if you want to represent an arbitrarily long "game" in Conway's Game of Life you only need to give the starting position because we know exactly how the game state will look at every step because we know the rules of the game.
This basically suggests that generalization in a ML model is a function of compression efficiency. ML models memorizing data isn't actually an issue. It's memorizing data INEFFICIENTLY that a problem. Models which overfit have learned inefficient representations of the underlying relationships.