3 ms·
This might give a decent estimate of the entropy of the file, but it is a poor approximation of the Kolmogorov complexity in general. All major compression lib
by jsenn 3y ago
This might give a decent estimate of the entropy of the file, but it is a poor approximation of the Kolmogorov complexity in general.
All major compression libraries that I know of essentially work by "factoring out" statistical regularities in the input string--i.e. repeating patterns. But many strings have regularity that can't be captured statistically. For example, the sequence 1, 2, 3, 4, ... has an entropy rate of 1, because all digits and all subsequences of digits occur with equal frequency, so most modern compression libraries are going to be unable to compress it to any significant degree [1]. However, it obviously has a very compact representation as a program in your favourite language.
If you want to go down a bit of a rabbit hole, this video from a researcher in this field gives an overview, as well as some proposed methods for approximating Kolmogorov Complexity in a reasonable way (even though in a technical sense approximating it with known, fixed bounds is I believe impossible): https://www.youtube.com/watch?v=HXM3BUXsY4g https://www.youtube.com/watch?v=HXM3BUXsY4g. This paper also discusses a theory that decomposes KC into a statistical part (Shannon entropy) and a "computational part": https://csc.ucdavis.edu/~cmg/compmech/pubs/CalcEmergTitlePage.htm https://csc.ucdavis.edu/~cmg/compmech/pubs/CalcEmergTitlePag....
[1]: Note that if you encode it as ASCII or even N-bit integers there will be significant redundancy in the encoding. To properly test it you'd have to encode the numbers in a packed binary encoding with no padding.
- Panzer04 3y agoThus is how some video encoding algorithms ans the like sort of work, iirc - you encode interface differences under the assumption that any particular part of the screen isn't likely to change much. Encoding algorithms end up getting specialised for particular types of data because different data exhibits different patterns.