6 ms·
> You can accomplish 90% of what you want with RSS just by parsing the metadata on an HTML page. I am not sure you can. Metadata is quite inconsistent and many
by capdeck 8y ago
> You can accomplish 90% of what you want with RSS just by parsing the metadata on an HTML page.
I am not sure you can. Metadata is quite inconsistent and many websites that I have the RSS feed for block crawlers and scrapers. Advantages of RSS are common standard and design to be consumed by a bot. What we've got instead are curated fb and twitter feeds that people get hooked on and never look back. Openness is the aspect of RSS that I'll miss the most.
- irrational 8y agoHow do they block crawlers and scrapers? If I grab the page using headless chrome, how can it tell what I'm doing with the data on my end, especially if I have the script grab the pages at a rate that would be more consistent with a human?
- tubbs 8y ago> at a rate that would be more consistent with a human? I think that's the hard part. Some content providers put out content way quicker than a human could ever keep up with. I wrote a web crawler once to download high resolution pictures of cars from a Russian website at a quick but modest pace and I got IP banned within a couple hours.
- kbenson 8y agoIt depends on the site. If the Russian site wasn't used to the amount of traffic you generated through your downloads, it's easy to see. If a site is updating content so fast that crawling it all would get you blocked, you use multiple crawlers. It's trivial and cheap to get 100 different ips for crawling for your headless chromes, instances. At that point you could get something new once a minute and not reuse an ip for close to two hours. A more aggressive reuse schedule would make that sufficient for almost any site that it's worth crawling.
- shabbyrobe 8y agoHow do you get 100 different IPs trivially and cheaply? I suppose we might have different definitions of cheap; 100 EC2 t3.nanos would be a bit out of my price range, but that might be small change for the next person.
- gsich 8y agoVPN provider with different locations?
- kbenson 8y agoAWS is far from a cheap provider. DigitalOcean (as well as many others) have 1GB instances with 25GB disks for $5/mo. Some allow additional IPs for a few dollars. Deploy a bunch of minimal instances that all they do is run a socks proxy and are firewalled off from everyone except your source IP address, and use them as the configured proxy for your actual scanning machine(s). For chrome, you can configure a proxy on the command line. Most quality http libraries in most languages support proxies (otherwise I wouldn't consider them quality libraries...). For a business (which what I'm thinking in terms of), $500/mo for 100 IP addresses which you can trivially cycle by destroying and creating new virts is well within cost. For an individual, I assume you wouldn't be crawling 24/7, so bring up as many as needed for short bursts. At $0.007/hr, you could bring up 100 from a snapshot/template, use them for a few hours to crawl whatever you need, and it will cost you just over $2.[1] For longer term but smaller needs, just spin up 5-10 for $25-$50/mo cost. 1: 100 * $0.007/hr * 3 hours = $2.10
- jjeaff 8y agoYou also have to factor in the data transfer which will add up quickly. Aws also charges an extra fee for cycling through IP addresses. I think maybe 25 makes it under the free tier, but more than that and it will start to add up as well. DO data is cheaper, but not free, and I'm not sure if they charge for changing IP addresses a lot. Not to mention that in many cases, admins start blocking while up ranges. And due to scraping, lots of sites block popular ip ranges like that of AWS. There are actually forums where people keep updated ranges for the different cloud platforms.
- aftbit 8y agoHeadless Chrome sets a different User Agent than normal Chrome. Once you get past that, there are trickier ways to distinguish bots from humans. These range from access pattern analysis (who visits hundreds of URLs per day but never clicks anything?) to differences in the JavaScript and rendering behavior of headless Chrome from normal Chrome. Developing a truly human-like bot that can pass modern anti-bot logic is very hard, and that's without even dealing with IP reputation or ReCaptcha.
- jjeaff 8y agoEven if you manually browse lots of pages in Google or Amazon, you will eventually be presented with a captcha because it thinks you are a bot.
- pteraspidomorph 8y agoVery large websites aside, the RSS offer on most of the web is pretty appaling. Standards are not followed or followed poorly. Things like publication dates that are actually the upload date and completely unrelated to the date the content went online, or images that are loaded by html embedded in the description field in nonstandard ways, or sometimes in the image field, or not at all... Or it's the wrong image, because the feed is provided automatically and doesn't have the intelligence to correctly parse the content... That said, if Firefox hadn't hidden live bookmarks away several years ago they might see more use. The small subset of users who bothered to look for them and bring them back to the toolbar might even overlike the set of users who disable telemetry, too.
- PaulHoule 8y agoI have been through this Wikipedia vs DBpedia, any kind of case where there is the primary product (which is part of quality control loops; bugs in the HTML get fixed) vs secondary products (RDF/RSS bugs don't get reported never mind fixed.) Part of the solution is this sort of A.I. http://ontology2.com/essays/ClassifyingHackerNewsArticles/ http://ontology2.com/essays/ClassifyingHackerNewsArticles/ but trained by you and not some other person who is training it to control your behavior.
- ehnto 8y agoI'm not sure it would be any different for any other feed specification unless the spec maintainer supplies all the software. It's a human problem rather than a software/spec problem. But I agree of course, it's a bit of a mess and definitely not something you can rely on to use in automated systems without a lot of extra validation.
- Sophira 8y agoThere's an article from 14 years ago that I remember to this day, mainly for the immortal line "RSS 2.0 is incompatible with itself": http://www.diveintomark.link/2004/the-myth-of-rss-compatibility http://www.diveintomark.link/2004/the-myth-of-rss-compatibil... It goes into some detail about the various RSS incompatibilities and is well worth the read, because RSS never really got any better even afterwards.
- ChuckMcM 8y agoThis captures it. Metadata isn't a controlled spec so there is no leverage to say "you are doing it wrong". As a result when you only care about a couple of "big players" it is feasible to just follow along with their changes as they change but a dozen different players? That is impractical.
- greglindahl 8y agoWebsites tend to use Facebook opengraph metadata correctly, because they can see the results on Facebook. I think that's the kind of thing that burtonator was referring to.
- eli 8y agoDunno if RSS is the example I’d hold up of a controlled spec success story. Worked fine if you only ever had links, titles and, sometimes, some body text.