6 ms·
Hacker News BigQuery Dataset
- cannabisfarmer 8y agoBigQuery keeps adding useless data. What we truly need is common crawl data then we can check specific site on our own. Or wait, BigQuery simply can't handle common crawl size dataset in their public service! Otherwise there is no reason to not add it, maybe it puts their search engine/ad business in geoparady. Is there any other Google public dataset BigQuery like platform? Where their direct search engine/ad platform interests don't get in way of Common Crawl like data searching/indexing?
- minimaxir 8y agoLooks like it stopped updating as of February 2nd, but otherwise it's pretty reliable, and as noted in the description, it's free. (you probably won't hit the 1TB limit working with this dataset). Here's a few queries I've done recently to answer ad-hoc questions to get an exact answer: Top posts about bootstrapping (https://news.ycombinator.com/item?id=19258249 https://news.ycombinator.com/item?id=19258249): #standardSQL SELECT * FROM `bigquery-public-data.hacker_news.full` WHERE REGEXP_CONTAINS(title, '[Bb]ootstrap') ORDER BY score DESC LIMIT 100 Count of YC startup posts over time by month (https://news.ycombinator.com/item?id=19185946 https://news.ycombinator.com/item?id=19185946): #standardSQL SELECT TIMESTAMP_TRUNC(timestamp, MONTH) as month_posted, COUNT(*) as num_posts_gte_5 FROM `bigquery-public-data.hacker_news.full` WHERE REGEXP_CONTAINS(title, 'YC [S|W][0-9]{2}') AND score >= 5 AND timestamp >= '2015-01-01' GROUP BY 1 ORDER BY 1
- avian 8y agoOne thing that I was missing last time I checked was comment ranking data. Neither score nor rank was there for comments posted in recent years. I understand that upvote counts are not available in the API, but ranking should be (as in, the order the comments appear on the page).
- shanecglass 8y agoHey all, I manage the BigQuery Public Datasets Program here at Google. You're right, the dataset last updated February 2nd, but we intend to continuing updating it. We had an issue on our end that disrupted our update feed, but we're working to repair it now and get the latest data uploaded to BigQuery.
- minimaxir 8y agoGreat to hear! :)
- tobr 8y agoA dataset like this is going to have a bunch of personal information in it. When it’s distributed like this, how does that jive with regulations like GDPR? If a HN user would like to delete all their comments, how would that request be forwarded to every user of this dataset?
- damajor 8y agoI support this question. Any comments ? NB: That's easy to downvote without commenting...
- espeed 8y agoThe Internet is written in ink. You should assume that any and all public posts you make have already been replicated and archived by countless parties in countless ways by the time you hit delete. HN public postings are no different. The HN API [1] has been around in various forms for years and includes the same public data that's used to generate the public pages on the HN site, but rather than returning HTML pages designed for human consumption, the API returns the data in a JSON serialized form [2] designed for machine consumption [3]. When the HN API went live, it reduced the overhead and redundant work from all the programmers having to independently crawl and parse site. The HN BigQuery dataset is the same data returned by the HN API, Google just took the next step and did the work of loading it into BigQuery. [1] https://github.com/HackerNews/API https://github.com/HackerNews/API [2] https://en.wikipedia.org/wiki/Category:Data_serialization_formats https://en.wikipedia.org/wiki/Category:Data_serialization_fo... [3] https://en.wikipedia.org/wiki/Machine_to_machine https://en.wikipedia.org/wiki/Machine_to_machine
- kolinko 8y agoThe data is already publicly available for use and reuse, just in a different form. Why would it be any different than the rules regarding the public/api display of the information?
- tobr 8y agoMaybe I’m misunderstanding what this is, is it not possible to query, download and process at a completely different scale than the API? If not, I suppose you might ask the same thing about the API.
- cobookman 8y agoTop Commentors of all time. tptacek is at 1st place with 33839 comments. Hacker news is 12 years old. That's an average of 7 comments per day since inception. Wow #standardSQL SELECT author, count(DISTINCT id) as `num_comments` FROM `bigquery-public-data.hacker_news.comments` WHERE id IS NOT NULL GROUP BY author ORDER BY num_comments DESC LIMIT 100;
- minimaxir 8y agoDon't use the `comments` table: it was last updated December 2017. On the full table: #standardSQL SELECT `by`, COUNT(DISTINCT id) as `num_comments` FROM `bigquery-public-data.hacker_news.full` WHERE id IS NOT NULL AND `by` != '' AND type='comment' GROUP BY 1 ORDER BY num_comments DESC LIMIT 100 tptacek is in first place with 47283 comments.
- fhoffa 8y agoHi, Felipe Hoffa at Google here. We're aware the dataset hasn't been updated since a month ago, and we are working to fix it. You can track the issue here: - https://issuetracker.google.com/issues/127132286 https://issuetracker.google.com/issues/127132286 In the meantime you can still play with the dataset, and dig into the full history of Hacker News - less this last month. I left some interesting queries to get you started here: - https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-daily-updates-so-what-are-the-top-domains-963d3c68b2e2 https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-d...
- espeed 8y agoWow, what timing. Late last night I had a conversation with someone explaining that Hacker News is not your typical message board -- it's owned and operated by YC and sits atop algorithms developed by some of the pioneers in spam and anomaly detection [1] [2], and it's is also an open dataset -- analyzed and scrutinized -- used by hackers worldwide to train and test bespoke AI. HN is a live MNIST [3] for anomaly detection. [1] http://www.paulgraham.com/spam.html http://www.paulgraham.com/spam.html [2] http://googlesystem.blogspot.com/2007/07/paul-buchheit-man-behind-gmail.html http://googlesystem.blogspot.com/2007/07/paul-buchheit-man-b... [3] http://yann.lecun.com/exdb/mnist/ http://yann.lecun.com/exdb/mnist/
- wodenokoto 8y agoHow do you run the import? Love to read more about how you consume the data
- physcab 8y agoI would also like to know where the data comes from!
- naveen99 8y agoAny particular reason you don’t include user profiles in the dataset ? I ended up pulling them myself using the api...
- danielecook 8y agoI've been using the HN API to maintain a bigquery table of all posts, comments, and URLs on HN and putting it on BigQuery for a while now. I use it to put this site together: https://hntrending.com/ https://hntrending.com/. BQ is awesome. It's a side project so may have some issues!
- vinnyglennon 8y agohttps://hnify.com/leaderboard.html https://hnify.com/leaderboard.html using the dataset tool too, amazing to have so much data freely available to play with.
- lettergram 8y agoI’m actually fairly excited to learn about this. I painstakingly scrapped HN to build: https://hnprofile.com/ https://hnprofile.com/ I’m excited about this alternative
- refrigerator 8y agoLast year I built a domain leaderboard based on this dataset: https://hnleaderboard.com https://hnleaderboard.com — planning to update for 2019 soon!
- sbr464 8y agoI added a simple api endpoint to access favorites on HN, since they weren’t available on the normal api. https://github.com/reactual/hacker-news-favorites-api https://github.com/reactual/hacker-news-favorites-api
- fsiefken 8y agois there a way to download the dataset and query it locally from for example postgresql or sqlite? How big is the database, 4G compressed?