3 ms·
People don't pick words from a fixed distribution in the manner which you wrote. I scanned over 500,000 words I wrote and I never mentioned the word mentioned
by ececconi 10y ago
People don't pick words from a fixed distribution in the manner which you wrote. I scanned over 500,000 words I wrote and I never mentioned the word mentioned on the headline of this article. I'm not sure if any kind of meaningful interpretation can be made from the headline, but I think it's an interesting observation nonetheless.
- hrayr 10y ago> I scanned over 500,000 words I wrote and I never mentioned the word mentioned on the headline of this article. Funny how you're trying to avoid breaking your streak of not mentioning nazis.
- Someone 10y agoPeople do not do that, but at least it is a model that one can improve. For example, as other commenters indicate, one can argument that occurrences within threads will be correlated (that will happen in discussions about World War Two, for example) but others indicated that "Godwin's law" occurrences may follow the model better. In contrast, the OP just presented a bare number, implying that it was high, but without any evidence for that claim. And yes, there's lots of room for improvement. The model is very simple, the guesses at the Reddit thread sizes are 100% guesses (how long is the average comment? What's the distribution of thread lengths? Etc), I did some serious rounding here and there, trusted Google to do the math right (it is not that hard, but IEEE may be insufficient to get results that are somewhat correct, and I don't know what Google does), etc. As to your "I never wrote nasi goreng in 500,000 words" argument: let's take my simple/flawed model. It predicts (1-1/125000)^500000 ~= 0.018 as the probability that 500,000 words do not contain "taboo". So, that simple/rough/flawed model predicts that, out of 1000 people writing 500,000 words, 18 wouldn't mention "Voldemort". You could easily be one of them. It's about as likely as somebody flipping 6 heads in a row. So, I don't think that single data point invalidates the model (no, I haven't done the proper statistics)