3 ms·
I have a question about the statement "you can't truly measure randomness", if it's alright (since you seem knowledgeable about this). > For example a "random
by asdf_snar 5y ago
I have a question about the statement "you can't truly measure randomness", if it's alright (since you seem knowledgeable about this).
> For example a "random number generator" that just has a large hardcoded list of random-looking values and loops over the list may pass a test if the list is long enough that the loop can't be detected, even though the values are decidedly not random.
My understanding is that's entirely fine from a statistical point of view. One can establish properties of this sequence up to some probabilistic error, perhaps also depending on the total number of values from the list that end up being used.
Isn't this in some sense analogous to measurement error? As far as I know, nothing can be precisely "measured" either. Is there a deeper sense in which the quality of the random sequence can't be analogously assessed?
- kevincox 5y ago> since you seem knowledgeable about this I should probably state to be clear that I am not a cryptographer. I just have an interest in the subject. It also depends a lot what you mean by random. I generally take "random value" to mean "a value which can not be predicted". Of course in cryptography we often use a CSPRNG which can be perfectly predicted, it is deterministic, not random at all! However in that case what we actually care about is "a value which can not be predicted without knowing the secret state". For many use cases this is sufficiently close to true random. (Or the predictability can actually be useful, for example using a CSPRNG as a stream cypher.) The reason why I said that my large hardcoded list was not random is that it is predictable. It is easily predictable if you know the hardcoded list of course. But even if you declare that list as "secret state" it is still predictable once it starts repeating and spews out the exact sequence of bits over and over again. > My understanding is that's entirely fine from a statistical point of view. Sure, but statistics are different than cryptography. If you just want your simulation to work well it is probably fine (unless the repeating causes funny artifacts in the simulation). I would argue that for most non-cryptographic use cases you don't even need anything close to "true random", you just need a distribution that matches your desired distribution with little bias and ideally doesn't repeat in a way that aligns poorly with how you use the numbers. > Isn't this in some sense analogous to measurement error? I'm not sure I completely follow here but I think the answer is "no". The difference is that randomness is a property of the generation process, not the output. https://xkcd.com/221/ https://xkcd.com/221/ is a great way to highlight this. In this case 4 was a very good quality random number. No once could have predicted it with greater that 1/6 probability. However once you check it into source code it isn't "random" any more. This is because the die produces "random numbers" but there is nothing special about the number itself. There is no test that will tell you that "4" is random or non-random, the question doesn't really make sense. What the tests do is they take a list of values and try to detect "problems" (situations in the output that would be unlikely if the source was truly random). For example if you have a weighed die that rolls 6 80% of the time a good test will tell you that your random source is likely biased based on a sample of output values. Of course the more data you feed to the test the more likely that it is that failures are true failures. If you roll a die 10 times and don't get a 6 it isn't proof that the die is unfair or predictable in some way. If you roll that same die 1000 times without a six it is looking much more suspicious, but still impossible to prove just by looking at the outputs.