4 ms·
>Notably, the neural network component behaves like a compressed index of the training data. The "Compression" part is indeed about the model finding order in
by cpldcpu 20d ago
>Notably, the neural network component behaves like a compressed index of the training data.
The "Compression" part is indeed about the model finding order in the training data. But this is not about applying some predefined compression algorithm on it, but by actually learning an algorithmic representation of the data.
This has nothing to do with "consulting the training data", because the training data cannot be reconstructed based on this information.
Roughly speaking, there are two types of information being stored in the NN: Memorization and generalization, or, shannon entropy and kolmogorov complexity.
This paper is highly underrated: https://arxiv.org/pdf/2505.24832 https://arxiv.org/pdf/2505.24832
- 112233 20d agogreat distinction. can we use it to freely produce and share mp3/h264 copies of any media? because it is provably impossible to reconstruct original based on that information.
- cpldcpu 20d agoWell, that is what video models do. But keep in mind: So far there is no way to train the models while completely avoiding memorization and only including generalization. That would be a great way to avoid any copyright issues, but all attempts I have seen so far were fairly limited.
- 112233 18d agoreally? so what is the bright red line? amount of compute used? use of linear algebra? Is https://github.com/google-deepmind/c3_neural_compression https://github.com/google-deepmind/c3_neural_compression AI enough? </snrk> Because I really want to see difference between approximating input signal using smooth orthogonal basis functions, and approximating those using quantization with rate/distortion feedback as "violates copyright, compression algorithm", and then approximating input signal using perceptron, with backprop loss based off original media as "not compression algorithm, skip jail"