6 ms·
Show HN: Analyzing top HN posts with language models
Hi HN,
I spent a few weeks looking at the top HN posts of all time. This included exploration, clustering, creating visualizations, and zooming in on what (to me personally) seems like some of the best discussions on here.
Three things in this post:
1- The interesting groups of HN posts
2- The interactive visualizations that you can explore in your browser
3- The data from this exploration -- this includes CSV of the titles as well as the text embeddings of 3,000 Ask HN articles.
Blog post about this whole process here: [1]
============
1- The interesting groups of HN posts
From the exploration, Ask HN proved the most interesting. These are the top four groups of topics I found insightful. Each group contains about 400 posts.
- Life experiences and advice threads [2]
- Technical and personal development [3]
- Software career insights, advice, and discussions [4]
- General content recommendations (blogs/podcasts) [5]
============
2- The interactive visualizations that you can explore in your browser
- Top 10,000 Hacker News articles of all time [6]
- Top 3,000 posts in Ask HN [7]
============
3- The data from this exploration
CSV file of top 3K Ask HN posts: [8]
The sentence embeddings of the titles of those posts: [9]
This is a colab notebook containing the code examples (including loading these two data files): [10]
============
If you've ever wanted to get into language models, this is a good place to start. Happy to answer any questions
- jayalammar 4y ago[1] https://txt.cohere.ai/combing-for-insight-in-10-000-hacker-news-posts-with-text-clustering/ https://txt.cohere.ai/combing-for-insight-in-10-000-hacker-n... [2] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_6a.html https://assets.cohere.ai/blog/text-clustering/askhn_cluster_... [3] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_7a.html https://assets.cohere.ai/blog/text-clustering/askhn_cluster_... [4] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_5a.html https://assets.cohere.ai/blog/text-clustering/askhn_cluster_... [5] https://assets.cohere.ai/blog/text-clustering/askhn_cluster_3a.html https://assets.cohere.ai/blog/text-clustering/askhn_cluster_... [6] https://assets.cohere.ai/blog/text-clustering/hn10k_clustered.html https://assets.cohere.ai/blog/text-clustering/hn10k_clustere... [7] https://assets.cohere.ai/blog/text-clustering/askhn-3k.html https://assets.cohere.ai/blog/text-clustering/askhn-3k.html [8] https://storage.googleapis.com/cohere-assets/blog/text-clustering/data/askhn3k_df.csv https://storage.googleapis.com/cohere-assets/blog/text-clust... [9] https://storage.googleapis.com/cohere-assets/blog/text-clustering/data/askhn3k_embeds.npy https://storage.googleapis.com/cohere-assets/blog/text-clust... [10] https://colab.research.google.com/github/cohere-ai/notebooks/blob/main/notebooks/Analyzing_Hacker_News_with_Six_Language_Understanding_Methods.ipynb https://colab.research.google.com/github/cohere-ai/notebooks...
- jayalammar 4y agoDisclosure: These were made by Cohere's embeddings, a company where I work. The process should work on text embeddings from other sources.
- toppy 4y agoI don't know how HN score metrics work but after some short review of the datafile [1] I've noticed a lot of the posts has the form of a simple questions and as such seems to be naturally biased when comes to user engagement. Have you considered to add additional metrics to remove that bias and re-analyze? [1] https://storage.googleapis.com/cohere-assets/blog/text-clustering/data/askhn3k_df.csv https://storage.googleapis.com/cohere-assets/blog/text-clust...
- jayalammar 4y agoWhat do you mean by naturally biased? That people seem to favor them?
- toppy 4y agoPeople seem to be more engaged in discussions arising from questions rather than statements, no?
- jayalammar 4y agoI think that's part of the expectations out of "ask HN". I don't know that the same effect happens outside of Ask HN.
- uniqueuid 4y agoInteresting, but it doesn't seem like the dimensionality reduction produces a good separation of topics. The UMAP projection looks pretty dense. Did you consider pruning or using something other than embeddings?
- jayalammar 4y agoSo it really depends on what you use for clustering. In this case, I'm clustering by the original embeddings so the UMAP results are different. I've also seen: 1- Clustering by UMAP. Here the plot would show clean separation of topics. But the clustering algorithm would be working on highly compressed data (from the 1024 dimensions of the embedding down to the 2 of UMAP). 2- BERTopic's approach of doing UMAP down to 5 dimensions, using this dimensionality for clustering, then UMAP again from 5 to 2. Which is an interesting approach. I've heard people having good results with all three. It's kinda hard to objectively compare, but my leaning was to give the clustering algorithm the representation containing the most information about the text.
- PaulHoule 4y agoTry t-SNE. I used to scoff at cluster plots until I saw those but with t-SNE… wow, those clusters are actually separated!
- uniqueuid 4y agoAre you sure t-SNE and UMAP actually perform very differently? Last I looked, they were somewhat comparable. [edit]: Seems they are similar for some purposes: https://blog.bioturing.com/2022/01/14/umap-vs-t-sne-single-cell-rna-seq-data-visualization/ https://blog.bioturing.com/2022/01/14/umap-vs-t-sne-single-c... Also interesting: Rapidsai has a cuda accelerated version of umap that is very fast (hdbscan as well BTW).
- nighthawk454 4y agoAll these dimension reduction methods are extremely similar. The math essentially just preserves nearest neighbors, with a setting for how 'tight' you want the clusters to be. Check out this image [1] and accompanying paper [2] for further reference [1] https://www.semanticscholar.org/paper/A-Unifying-Perspective-on-Neighbor-Embeddings-along-B%C3%B6hm-Berens/9411cea70d31834f6cb4d2e9bee205e02b56d938/figure/0 https://www.semanticscholar.org/paper/A-Unifying-Perspective... [2] https://arxiv.org/abs/2007.08902 https://arxiv.org/abs/2007.08902
- bryanrasmussen 4y agoas people upvote other things than your list of relevant links it becomes difficult to find the relevant links. although I guess people can find it by your name. on edit: so it seems some are upvoting the links to keep them on the top in opposition to those upvoting discussion points.
- hoerzu 4y agoI grouped posts by topic: https://n3ws.ploomber.io https://n3ws.ploomber.io Explanation: https://ploomber.io/blog/hn_classifier/ https://ploomber.io/blog/hn_classifier/
- chrisMyzel 4y agosorry for the OT post, wanted to cheer for ploomber - love it for writing pipelines, didnt know your EU based (guessing by ur username?).
- wpietri 4y agoThe conflict of interest here concerns me. I don't object to content marketing, but I'd rather a) you were clear from the start that you work for this company and are promoting its product, and b) that this "revolves around [...] using Cohere’s Embed endpoint", so that people can judge how much they want to "get into language models" with pay-per-character pricing, as opposed to something more open.
- jayalammar 4y agoThanks. I just added a disclosure to the comment (can't edit the parent anymore). The full embeddings are freely provided here without the need to use the service.
- wpietri 4y agoThanks! I appreciate it.
- rmbyrro 4y agoDo you advocate for the disclaimer only because 1) the sample uses their product or 2) just because they sell a product correlated to the topic? I see a lot of articles that fall into #2 being published here without a disclaimer. And I think a disclaimer isn't necessary for #2. Even for #1 I wouldn't bother, but I understand the expectation. Many advocate a lot against ads, targeting, etc. If we also advocate against promotional content, what would companies do to get attention and traffic?
- notatoad 4y agomost content marketing goes to a post on the company's own website, which is a sort of inherent disclosure. i don't think asking people to disclose that they work for the company whose product they're promoting is "advocating against promotional content".
- rmbyrro 4y agoThat's true, it's not advocating against. Bad way of expressing. But I do think it may undercut the power of content marketing by introducing an unconscious and unjustified negative bias. If a stranger shares it on HN, should I judge it in a different way? What is the true value of highlighting that the author works for the company?
- victorianoi 4y agoNice idea and analysis! I reproduced it as well with https://graphext.com https://graphext.com and got similar clusters https://drive.google.com/file/d/1-kXsKezu2_S07rQn-0bjbHuUXHEZUqg4/view?usp=sharing https://drive.google.com/file/d/1-kXsKezu2_S07rQn-0bjbHuUXHE...
- victorianoi 4y agoBTW there is an implicit recency bias in the dataset, since 2017 the number of top 3K post became more frequently and the avg score is larger year after year as the community in HN grows: - Number of top 3K per month of publishing - https://drive.google.com/file/d/1beAPP9ijruMUs5DN5wOVsBArvxPzYhmU/view?usp=sharing https://drive.google.com/file/d/1beAPP9ijruMUs5DN5wOVsBArvxP... - Avg score of top 3K per month of publishing - https://drive.google.com/file/d/10nSIgH1a6DN6XrDU2DyMJTCgsIgdqyX8/view?usp=sharing https://drive.google.com/file/d/10nSIgH1a6DN6XrDU2DyMJTCgsIg...
- victorianoi 4y agoand it also looks like most topics are constant over the years - https://drive.google.com/file/d/1ilYn9cnEZwiH1FioUtU9ummhmvngP2LI/view?usp=sharing https://drive.google.com/file/d/1ilYn9cnEZwiH1FioUtU9ummhmvn...
- natch 4y agoThere's too much fixation with "top" in our industry. Top voted tends to mostly be a function of early posting. Later posts don't get votes because they simply were not seen. There seems to be a misreading on a mass scale of what "top" really indicates though; people think it means "quality" when it does not. Study after study, website after website, policy after policy, our online world is built on this fundamental misunderstanding of what is really going on. How do you avoid piling on to this misunderstanding?
- tra3 4y agoIs there a name for this phenomenon so I can google further? Intuitively makes sense, because I’ve seen this before.
- ultra_nick 4y agoPreferential Attachment,Power law distribution, and the 80-20 rule are all the same. https://en.m.wikipedia.org/wiki/Preferential_attachment https://en.m.wikipedia.org/wiki/Preferential_attachment
- mateo1 4y ago"Already covered"
- earthboundkid 4y agoThe scissor statement but it’s Go vs Rust.
- arolihas 4y agoLove your blog posts and visualizations Jay, thanks for sharing!
- renewiltord 4y agoI love this shit, dude. Not for any great insight. I just enjoy this sort of thing. Good stuff.