5 ms·
>A major internet site had a URL that went something like somedomain/group?id=xxxxx. It turns out that a simple scraper, that called id=1, id=2, id=3, ect, ect,
by odorousrex 9y ago
>A major internet site had a URL that went something like somedomain/group?id=xxxxx. It turns out that a simple scraper, that called id=1, id=2, id=3, ect, ect, caused a major problem!
This is a failure on the part of the developers at that "major internet site". Using a guid instead of consecutive IDs, a rate limiter, hell even just a cache...or all of the above. There are lots of solutions here.
You have to take robot scraping and indexing into consideration, and assume people will ignore robots.txt. (Certain bots, i.e. msnbot/bingbot are quite aggressive!)
- dec0dedab0de 9y agoNo, that is a failure of the developer of the scraper. I am definitely pro scraping, but you have to be a good neighbor.
- jstarfish 9y agoHow the hell is the scraper dev supposed to anticipate how poorly-written these particular views are with no backend knowledge? If not an automated scraper, a thundering herd from content gone viral would trigger the same result.
- bottled_poe 9y agoScraping is not an intended purpose for most websites. Unless the website specifically states that this is an intended function, it is not reasonable to assume so. In fact it may be in violation of the terms and conditions of the given website.
- kevin_thibedeau 9y agoI don't intend people named Steve to access my open site so I can sue all Steve's for their felonious behavior?
- nostrademons 9y agoIf the law assumed that only intended functions are permissible, innovation would be a crime. By definition, innovation is finding new and unforeseen uses for resources.
- blowski 9y agoYou both make good points. If you make the law too strict you punish reasonable uses of the website, like scraping a few publicly available pages to help users. If you make it too lenient you permit DOS attacks. It’s not easy to craft a law that will punish bad behaviour without blocking innovation.
- davvolun 9y agoI've done some scraping work -- one of my rules of thumb is to always assume the worst of their site and try to be as gentle as possible.
- r3bl 9y agoOh come on, you're trying to scrape the data out of a black box. You have no idea what their infrastructure is like, and for your purposes, you don't really care. Of course, some sense is more than welcome, but if my scraper makes one request every 2 sec knocks down your server, it's your fault, not mine.
- carterehsmith 9y ago>> This is a failure on the part of the developers at that "major internet site". Using a guid instead of consecutive IDs, a rate limiter, hell even just a cache...or all of the above. There are lots of solutions here. You are right, but few organizations are sophisticated.. or wealthy enough to employ all of that. I mean, a couple years ago there was a thing that Google's Docs could be enumerated. And that's Google, they can obiously afford to get competent people working on that, yet they made a mistake (and who doesn't?).
- harpiaharpyja 9y agoFair enough. It still shouldn't become a criminal issue.
- tomarr 9y ago>You have to take robot scraping and indexing into consideration, and assume people will ignore robots.txt. (Certain bots, i.e. msnbot/bingbot are quite aggressive!) Who owns LinkedIn again?