7 ms·
Hacking Hacker News
- pflanze 13y agoPrevious discussion: https://news.ycombinator.com/item?id=3602407 https://news.ycombinator.com/item?id=3602407
- minimaxir 13y agoFYI, the new Hacker News API allows easy programmatic access of story/comments and infinite chronological paging. You could download every Hacker News story in less than 2 hours without breaking the API request limit. https://hn.algolia.com/api https://hn.algolia.com/api
- jrussino 13y ago"new Hacker News API" -> do you mean there's now an official API? I must have missed this news. Can you provide a link / additional info?
- sillysaurus3 13y agothe new Hacker News API allows easy programmatic access of story/comments and infinite chronological paging. You could download every Hacker News story in less than 2 hours without breaking the API request limit. Which API? https://www.google.com/search?q=hacker+news+api https://www.google.com/search?q=hacker+news+api
- jared314 13y agoWith hnsearch.com being shutdown, I believe he is referring to: http://hn.algolia.com/api http://hn.algolia.com/api
- sillysaurus3 13y agoUsing that API, you could download 48,000 hacker news stories in 2 days, so if there have been less than 48,000 submissions, then what minimaxir said is true. But first you'd need to generate a list of all 48,000 story ids, and there seems to be no way to actually do that.
- minimaxir 13y agoNo need to generate all IDs beforehand. The search_by_date endpoint is fine. You have to paginate using the created_at_i parameter, not the page parameter. Also, you can set hitsPerPage = 1000. ;)
- sillysaurus3 13y agoWhen trying to access page 2 via that endpoint: http://hn.algolia.com/api/v1/search_by_date?tags=story&hitsPerPage=1000&page=2 http://hn.algolia.com/api/v1/search_by_date?tags=story&hitsP... "you can only fetch the 1000 hits for this query, contact us to increase the limit" It was a nice try, but it did seem too good to be true.
- minimaxir 13y agoYou can paginate using the created_at_i parameter (edited OP) Just pass created_at_i<X, where X is the time stamp of the earliest submission. I was able to download 500k stories (i.e. about half of HN's 1.26M stories) before I ran into memory issues; I've fixed them and am downloading the rest.
- sillysaurus3 13y agoThat's incredible. Will you upload the raw database somewhere, please? If you make a torrent, I'll help seed it. Would you email me at sillysaurus3@gmail.com whenever it's ready?
- deleted 13y ago[deleted]
- deft 13y agoThis is old so what you're suggesting wouldn't be possible unless he updated it (which I don't think he has any reason to do).
- e15ctr0n 13y agoThere's a list of all the apps that have been built based on this API: http://hn.algolia.com/cool_apps http://hn.algolia.com/cool_apps
- pak 13y agoYup, I've switched over my Chrome extension [1] to use Algolia's API instead of HNSearch (which is shutting down), and so far, it seems to be working peachy. [1]: https://chrome.google.com/webstore/detail/hacker-news-sidebar/ngljhffenbmdjobakjplnlbfkeabbpma?hl=en-US https://chrome.google.com/webstore/detail/hacker-news-sideba...
- ehsanu1 13y agoWas this there when you posted?: RATE LIMITS We are limiting the number of API requests from a single IP to 1000 per hour. If you or your application has been blacklisted and you think there has been an error, please contact us.
- minimaxir 13y agoYes, that was there, and the math is still correct. You can query 1,000 stories per request, and do 1,000 requests per hour. That's 1M requests per hour. There are 1.26M Hacker News stories indexed by the API. :) EDIT: Finished downloading all the entries (and can confirm that 1.26M is indeed all of them). Took 3 hours due to a conservative wait period between each request to make sure I stayed within the limits.
- tripzilch 13y agoHow big is that dataset? Can you share it somehow?
- karangoeluw 13y agoI have written an API for HN: - Python module: https://github.com/karan/HackerNewsAPI https://github.com/karan/HackerNewsAPI - REST API: https://github.com/karan/HNify https://github.com/karan/HNify
- himal 13y agoGithub link: https://github.com/joelgrus/hackernews https://github.com/joelgrus/hackernews
- AznHisoka 13y agoIt seems you prefer to read articles from: - WashingtonPost - BusinessWeek - MarginalRevolution - NY Times [1] Based on BuzzSumo's social data: http://app.buzzsumo.com/#/influencers?q=@joelgrus&type=influencers&result_type=relevancy&blogger&influencer&company&journalist®ular_people&ignore_broadcasters=false&page=1 http://app.buzzsumo.com/#/influencers?q=@joelgrus&type=influ... (Press View Links Shared, Analyze Links Tab]
- joelgrus 13y agoOh jeez, who submitted this again? I learned my lesson a couple of years ago, everyone hates this. :)
- joelgrus 13y agoAlso, FYI, I don't even use this anymore, these days I just read the HN frontpage. :)
- tlarkworthy 13y agoIts exactly the kind of thing I would build and abandon. So the maybe more interesting things is why you don't use it? I presume its a UI thing, or classifier is unreliable, or something? I would love to hear why vanilla HN is better now.
- davidw 13y ago> I would love to hear why vanilla HN is better now. Probably because it has more politics and less 'hacker news' these days.
- deleted 13y ago[deleted]
- joelgrus 13y agoI don't know that vanilla HN is better now. I abandoned it for two main reasons: 1. The Hacker News API I was using was very unreliable and would go down for days / weeks at a time, which made the whole pipeline unreliable. 2. I was consuming this as an RSS feed, but when Google Reader shut down I abandoned my RSS habit cold turkey, so now I pretty much only read sites that I visit directly, or things people link to on FB / Twitter.
- icebraining 13y agoWhy did you use the API instead of consuming the HN RSS feed itself? They even offer a big feed for such usage: http://ycombinator.com/newsnews.html http://ycombinator.com/newsnews.html
- rickdale 13y agoThis is totally cool and kudos to you. But be aware there are limits of personalized hacker news. I think Bill Maher put it best when ranting just last night about facebooks customized news feeds: Newspapers may be old-fashioned, but here's what we're losing if you never see one. They are trying to tell you what's actually important, not just what's important to you. You may not read the whole paper, but you at least see headlines, making you aware that something's going on outside of your microtargeted world of fashion or music or Wiccans or zombies or whatever you're into. Replace 'newspapers' with hacker news and you get the point. https://www.youtube.com/watch?v=WohtmZDZCGM https://www.youtube.com/watch?v=WohtmZDZCGM
- ThomPete 13y agoNewspapers are trying to show you what they think is actually important. But of course omnibus papers will never really tell you what is important only what is current.
- hnriot 13y agototal nonsense, the classifier is just pushed upstream, from what I want to see, to what some ad-motivated editor wants me to see. newspapers reported stories to sell you something, be it papers, ads or someone's agenda, not because they believed we'd all be more rounded citizens.
- ketralnis 13y agoBut in the case of filtering Hacker News you're taking that pre-filtered list (filtered by geeks and startup folk and whathaveyou) and filtering it even more. Whether that bubble is a subset of some other bubble, it's still a bubble.
- hayksaakian 13y agoideally you could train the data set from stories i've upvoted on HN https://news.ycombinator.com/saved?id={{username} https://news.ycombinator.com/saved?id={{username}}
- j2kun 13y agoIt seems there is a small but very strong subculture of HackerNews readers who enjoy reading and discussing mathematical things. I would love to have a separate feed of those stories (and then after I'm done I could browse the HN front page), and I have often thought about the possibility of writing a program to do that. DataTau (the HN for data mining) seems to have failed, so I imagine a filter is the way to go rather than make a new website.
- jt2190 13y agoI think one of the challenges of hosting a "sub-HN" is that the hosting costs are hard to justify. This raises the question: How does YC justify hosting costs? My completely-off-the-cuff-assumption-take-this-with-a-huge-grain-of-salt is that YC benefits by having a huge audience to make announcements to, like job postings at YC funded companies, various pg essays, or just investing in overall goodwill from the HN audience. Probably the most likely reason is to increase deal-flow to YCombinator itself, though.
- icebraining 13y agoWhy does it need to have some justification beyond being a fun hobby? HN is hosted on a single server, and it probably uses less than 1TB/month, so it's not that expensive for someone with a Bay Area tech salary, let alone the whole YC.
- piracyde25 13y agoWait, this is 2 years ago?
- siculars 13y agoIdeas this the first time around. Cool, but still have the same problem I had then. Confirmation bias.
- seizethecheese 13y agoFrom the first paragraph: "people vote [links] up or down." Um... can't links only be voted up?
- sethaurus 13y agoPast a certain karmic threshold, both are allowed.
- ColinWright 13y agoAre you sure? I have over 60k karma and still can't downvote links.
- benaiah 13y agoNope, only comments. The only actions you can take on a link are upvoting or flagging it.
- arnorhs 13y agoIf this is a solution to the "not enough links that I personally like" - kudos to the author. Nice to find a fun project to work on that will also solve a problem for them. I personally despise recommendation / personalization algorithms of any kind. I still have never found one that's actually better than myself at distinguishing articles that I'd like to read, music that I want to listen to, tweets I'd like to see, etc. When reading HN, I'm constantly surprised by links that would not normally be on my radar for things I'm interested in. I think personalization algos, in general, are good at filtering those away. Since the author mentioned HN being too much of a firehose and this then also being a solution to the "too many links to keep up to date on" problem, the solution might be a bit simpler than the author suggested. HN already has the "best" links at https://news.ycombinator.com/best https://news.ycombinator.com/best It's hard to find - it's in the 'Lists' section in the footer, but it's still there and I use it all the time, when I haven't been actively reading HN for a while.
- ingend88 13y agoInteresting. This will go into today's top5HN Newsletter. Signup at top5hn.launchrock.co
- smoyer 13y ago"The model can only get better with more training data, which requires me to judge whether I like stories or not. I do this occasionally when there’s nothing interesting on Facebook. Right now this is just the above command-line tool, but maybe I’ll come up with something better in the future." If you let your program log into HN using your account, it should be able to tell which of the stories you've up-voted there. If you use that as the input to your classifier, as you read stories on HN, simply mark those that interest you by up-voting them. I'm also curious to know whether the stories are weighted by age to account for changes in what you find interesting.
- ninjakeyboard 13y agoCool story bro. You may want to check out digitalocean for hosting - their cheapest option is only $5 a month - about equivalent to the smaller $40/month aws option. It's very simple as well.
- kra 13y agoI just use the 50 or 100 point minimum feed in my reader, and skip articles that don't look interesting based on how much time I want to spend. Sometimes I only read articles if they're a day old and the first comment makes them look interesting.