4 ms·
Thanks, I changed that sentence to be more accurate. As for the usefulness of the measure, I've found it to be extremely handy for teasing out relationships tha
by eigenvalue 3y ago
Thanks, I changed that sentence to be more accurate. As for the usefulness of the measure, I've found it to be extremely handy for teasing out relationships that get missed using regular correlation measures.
- kqr 3y agoThis sounds a little like p-hacking but I guess I'm missing some nuance?
- eigenvalue 3y agoReally depends on the context and what you’re trying to do. If you’re trying to come up with an explanatory or causal theory of the relationship between some sequence and thousands of other sequences, then maybe that starts to turn into excessive “data mining.” If you’re using it more as a form of search (information retrieval), then I think there’s no harm in using it. For example, for ranking relevant embedding vectors. Apparently it works quite well for finding similar genes (I guess you replace base pairs with integers or something like that). Sometimes you just need a good place to look and then you can confirm things independently.
- tomrod 3y agoI'm personally a fan of mutual information and flavors, like transfer entropy.
- eigenvalue 3y agoMutal information is definitely a good measure, but it can struggle with complex and highly non-linear associations-- particularly when you are dealing in high dimensional spaces (because of the curse of dimensionality). Mutual information can have some serious bias/variance issues, especially when you don't have a huge amount of data to work with. Hoeffding's basically sidesteps all of these problems. The main downside of it is that it's so computationally intensive, much more than mutual information.
- CrazyStat 3y agoYou repeat the same error about negative association in a couple other places: > The final formula for Hoeffding's D combines D_1, D_2, and D_3, along with normalization factors, to produce a statistic that ranges from -0.5 to 1. This range allows for interpretation of the degree of association between the sequences, with values near 0 indicating no association, values closer to 1 indicating a strong positive association, and values near -0.5 indicating a strong negative association. > And a score near -0.5 suggests they're moving in opposite directions, perhaps clashing rather than complementing each other.
- eigenvalue 3y agoThanks again, fixed them.