4 ms·
“Sequence looks random” is not nonsensical. The authors say your sequence should be indistinguishable from e.g. 12 coin tosses. This would have a uniform distri
by moconnor 4y ago
“Sequence looks random” is not nonsensical. The authors say your sequence should be indistinguishable from e.g. 12 coin tosses. This would have a uniform distribution.
One approach would be estimate the probability distribution from the input sequence and calculate the KL-divergence [1] of that to the uniform distribution. This gives one objective measure of randomness. There are many others!
TL;DR: There are definitions of randomness that can be tested against.
[1] https://en.m.wikipedia.org/wiki/Kullback–Leibler_divergence https://en.m.wikipedia.org/wiki/Kullback–Leibler_divergence
- moasda 4y agoExactly, it's impotant to chose a definition of randomness first.
- moconnor 4y agoThe test specifies the random distribution very clearly in each case, though.
- icambron 4y agoYou can define terms however you want, and thus devise any measure you want. But that’s not a definition of randomness I recognize. Perhaps entropy. Randomness is not a property of the result; it’s a property of the process used to generate it. Let’s imagine applying your measure to a whole bunch of sequences, each of length N, with the elements of each actually drawn from a uniform distribution. You measure the randomness of each one. You’ll get a range of DLK results, distributed from, in your interpretation, “very random” to “not that random”. All-H or mostly-H will come up sometimes, the distribution estimator will return a skewed result, and it will get a divergent “score”. But everything came from the same distribution. So we’re now measuring the output of an actually random process and saying “we’ll, it’s usually random but not quite always” In contrast, let’s try your method on a different population of sequences, where instead of pulling the sequence elements from a distribution, every sequence is a hardcoded copy of HTHTHT… That gives “perfectly random”, even though it was very far from a random process. That’s close to what the study authors are doing here, except with a different definition of “random looking”. It’s measuring a property of a sequence, but it isn’t whether it was randomly generated. We could debate about whether this is a good measure of “random looking” and there could be lots of alternatives with no objectively best. But that is my point: if I ask “make me something random-looking”, I am only asking “how closely does your measure of post-facto randomness match mine?” An actual random sequence would be, well, random.