3 ms·
To clarify I'm not asking about HN itself but articles linked from HN. As you said the HN api is great and there are at least 2 existing published crawls of it
by agencies 4y ago
To clarify I'm not asking about HN itself but articles linked from HN.
As you said the HN api is great and there are at least 2 existing published crawls of it that help a lot.
- krapp 4y agoThe fastest way to get that would probably still be through HN's API, you just have to take the URL field for stories and ignore everything else.
- tedunangst 4y agoAnd how do you get the content once you have the URL?
- krapp 4y agoUse IA more responsibly, perhaps. Instead of scraping it, convert the list of links from HN to point to IA? You still have to work with whatever limits the site puts up in any case.
- arinlen 4y ago> And how do you get the content once you have the URL? I don't understand your question. If you have the URL, you just GET it, like any regular URL? Is there something that I'm missing?
- agencies 4y agoMany domains have expired or content is no longer available.
- agencies 4y agoIf a HN story is a link to Wikipedia, the HN api serves the content of the Wikipedia page??
- arinlen 4y ago> To clarify I'm not asking about HN itself but articles linked from HN. I might not have a clear picture of what you're looking for, but items of type "story" returned by the HN API do have a URL field, which I believe correspond to submitted links. You can scrape the text field of comment items, but that takes a bit more work.
- lcnPylGDnU4H9OF 4y agoHopefull this will help: you're talking about a submission to HN, e.g. a link to a WSJ article complete with comments section, and OP is talking about the specific WSJ article.