4 ms·
Sample size is the anyway|anyways|yeah|yea instances. I get those as they get posted, so that's why it fluctuates with time. Since they're so commonly used, it
by ddod 14y ago
Sample size is the anyway|anyways|yeah|yea instances. I get those as they get posted, so that's why it fluctuates with time. Since they're so commonly used, it should incidentally give you an idea of all of Twitter's load. I'm also grabbing some other words that I haven't implemented on the clientside yet, but I don't include them in the sample size.
HN comments tend to be more varied and sparse, so I think you're right that measuring yea:yeah, etc. wouldn't be too enlightening. That said, I've noticed a defined qualitative shift in HN comments over the past few months, and I'd like to develop ways of measuring that before they reach Reddit/Twitter levels.
As for your last point, I could track individuals but a relational comparison based off of /all/ data would be pretty difficult due to the number of comments vs. the few number of any individual's comments. Also, ARI isn't a great metric (hence me putting it in a tiny graph) because it measures chars instead of syllables. For example, "FFFFFUUUUUUUUU" has the same score as "constructivism".