4 ms·
I used my own metric that is based on jaccard similarity. Which in turn is based on "users who posted to X also posted to Y" metric. That said, there are a few
by anvaka 7y ago
I used my own metric that is based on jaccard similarity. Which in turn is based on "users who posted to X also posted to Y" metric.
That said, there are a few subreddits that are too popular and similarity results were too saturated (/r/videos, /r/funny, etc.) so I did a manual override by looking into most commonly mentioned other subreddits, and sometimes into `about` blurb of subreddit).
Please don't consider these recommendation as source of truth! It's just a fun way to discover other subreddits :).
I'm also very open to change this metric to something else - please let me know if you have any recommendations!
[1]: https://github.com/anvaka/sayit#the-data https://github.com/anvaka/sayit#the-data - describes the data, indexing scripts are here: https://github.com/anvaka/sayit/tree/master/scripts https://github.com/anvaka/sayit/tree/master/scripts
[2]: Manual overrides can be found here https://github.com/anvaka/sayit-data#sayit---recommendation-data https://github.com/anvaka/sayit-data#sayit---recommendation-...
- ketralnis 7y agoI work for reddit and here I’ve used a similar technique to your jaccard distance but with one twist: divide by the size of the smaller subreddit (in your case, the number of unique posters that you’ve recorded). That gives you a directional relatedness, that is programming->python but not necessarily python->programming. Used this way you account for the giant subreddit problem automatically but now the results are less “amitheasshole is related to askreddit” and more like “linguisticshumor is a more niche version of linguistics”. The great thing is that it’s actually more actionable as far as recommendations go! Everybody has already heard of the bigger version of this subreddit, but they probably haven’t heard of the smaller versions. And it’s self-correcting. As a subreddit gets bigger we are less likely to recommend it (which is great because it needs our help less)
- jcims 7y agoCan people still make their upvotes/downvotes public? I found that useful for sussing out related subreddits when it was possible back in the day.
- anvaka 7y agoThis is super awesome, thank you for sharing! If you guys are interested in seeing how your recommendation work for the entire reddit, I'd be happy to build you a spaceship similar to this one https://github.com/anvaka/word2vec-graph https://github.com/anvaka/word2vec-graph . I couldn't find an easy way to download the entire recommendation graph, but it would be awesome if we could make it work. My email is the same as this account at gmail, and twitter is all open: https://twitter.com/anvaka https://twitter.com/anvaka
- yantrams 7y agoIn case you haven't come across it already, here is a very exhaustive list of distance measures for dealing with problems of this kind - http://www.iiisci.org/journal/CV$/sci/pdfs/GS315JG.pdf http://www.iiisci.org/journal/CV$/sci/pdfs/GS315JG.pdf I fooled around a bit with lastfm data for band recommendation and found this sheet quite helpful. If you are interested in learning more about asymmetrical similarity, here is a great primer by Tversky - http://www.cogsci.ucsd.edu/~coulson/203/tversky-features.pdf http://www.cogsci.ucsd.edu/~coulson/203/tversky-features.pdf
- anvaka 7y agoThis is absolute treasure trove. Thank you so much!
- bobosha 7y agoNormalizing by the user count should help. My suggestion would be something much simpler ie. to try to compare content itself e.g. top 1000 posts from each subreddit, and estimate (say) cosine sim + Tfidf. Wouldn't that be a better indicator? Also, instead of pairwise comparison, you could try clustering (HDBSCAN for example) to reduce computational complexity. But great work, love your visualizations!
- ethn 7y agoI’ve also rediscovered this in similar work I do. But upon later research I’ve learned this is called containment.
- mikk14 7y agoIf you're interested, in the scientific literature this problem is known as "network backboning". Basically you have all nodes connected to practically all other nodes with weighted edges, and you want to know which are the edges with statistically significant weights. I wrote on this topic [1]. My method [2] basically uses simple counts on edge weights, and then estimates the expected edge weight and its variance using Bayesian priors. It then attaches a t-score or p-value to each edge, and then you can filter out edges with too low t-score. The idea is that weak edges can still be statistically significant if they connect "small" nodes. In any case, the library I wrote includes the implementation of a few other methods, in case they work better for your data type. [1] https://arxiv.org/abs/1701.07336 https://arxiv.org/abs/1701.07336 [2] http://www.michelecoscia.com/?page_id=287 http://www.michelecoscia.com/?page_id=287
- zmix 7y agoA lot of Reddits have "Recommended" or "Related" boxes at the side. Why don't you implement these as first, and then use the algorithm(s)?
- anvaka 7y agoThank you for the suggestion! I think when I created this tool there was no recommendations on reddit. When it was introduced later on reddit I was contemplating about using reddit's own recommendations, but at that time it was missing a few smaller subreddits, so I just put it of onto the shelf of projects to try.
- dannytatom 7y agoThis has a similar issue to the way Reddit recommends subs itself. A sub will be considered similar if it's the exact opposite due to users of one sub going to another to shitpost or brigade or etc. I'm not sure of a fix for that, but would it be possible/helpful to weigh it by average upvote/downvote of the comments left from users of said sub? Meaning if sub A is about how much baseball sucks, and sub B is about how amazing baseball is, while determining if the 2 are similar you'd find out most posts from sub A to sub B are heavily downvoted and so probably not similar.
- anvaka 7y agoIt seems though relationship can be considered to have a sign: positive, when people mostly align in their upvote intents, and negative when people do the brigading, etc. I'd still see ability to determine absolute value of relationship as a valuable property of a recommender
- grawprog 7y ago>I used my own metric that is based on jaccard similarity. Which in turn is based on "users who posted to X also posted to Y" metric. Is that really a good metric of similarity? Just myself, I post in several unrelated subreddits semi-regularly from programming to video games to music, art and even stone masonary, i've posted in subreddits for TV shows i've watched, or just completely random things. I use reddit as a place where I can learn about and interact with people on nearly any subject or topic and I take advantage of that when I can. I'm sure my posting habits aren't that unusual. I'm just not sure that's really an accurate way to gauge similarity.
- anvaka 7y agoRight, I don't think it's 100% accurate either. It gives just some hints what might be related, but like you said, it's not necessary the best possible measure of similarity. Jaccard similarity does not count only the number of people who posted to A and B, it checks how many people posted to A, how many people posted to B, and how many of those people posted TOGETHER to A and B. That togetherness gives us hints what is related, and after it is computed, we can divide by the total number of poster to both A and B (independently), which brings the value to something that we can use to compare against other subreddits. If that value is close to 1, it means that almost all users who posted to A have also posted to B. If it is close to 0, then the overlap is much smaller.