10 ms·
Having spent a week battling a particularly inconsiderate scraping attempt, I’m quite unsurprised by the juvenile tone and fairly glib approach to the ethics of
by ebbp 5y ago
Having spent a week battling a particularly inconsiderate scraping attempt, I’m quite unsurprised by the juvenile tone and fairly glib approach to the ethics of bots/scraping presented by the piece.
For the site I work for, about 20-30% of our monthly hosting costs go towards servicing bot/scraping traffic. We’ve generally priced this into the cost of doing business, as we’ve prioritised making our site as freely accessible as possible.
But after this week, where some amateur did real damage to us with a ham-fisted attempt to scrape too much too quickly, we’re forced to degrade the experience for ALL users by introducing captchas and other techniques we’d really rather not.
- jtdev 5y agoConsidering the demand for your content, why haven’t you created and provided an API? Maybe you could monetize?
- chewmieser 5y agoLike everyone and their brother has a web spider. And some of them are VERY badly designed. We block them when they use too many resources, although we'd rather just let them be. Can't speak for the op but we have APIs and move the ones scraping and reselling our content to APIs. The majority are just a worthless suck on resources though.
- ebbp 5y agoWe do offer an API - the scrapers are trying to circumvent using that, presumably.
- purerandomness 5y agoWhy do you think are they trying to circumvent it? Does your API provide all the information that can be found on the site, or are they scraping because the API is incomplete? We've once had to scrape Amazon product pages because they have a lot of API endpoints, but those didn't contain the data we needed.
- scarygliders 5y agoWhy would Amazon wish to provide you with easy to access data on their products and prices when you could either be a competitor wishing to undercut those prices, or be a scraper company hired by such a competitor? In what universe is providing such a straightforward way of helping a competitor considered sane business practice?
- jtdev 5y agoIn the end, they still get the data, just in a much less desirable way for both you and the customer.
- matheusmoreira 5y agoBecause they will get the data regardless of what you do and if you don't make an API it will cost you more due to overhead.
- weird-eye-issue 5y agoMan your comment is hilarious because in fact Amazon DOES provide an API for exactly that
- scarygliders 5y agoAnd yet... > We've once had to scrape Amazon product pages because they have a lot of API endpoints, but those didn't contain the data we needed. ...only a couple of comments up.
- weird-eye-issue 5y agoYou don't know what data they needed. Maybe they needed reviews or product descriptions. The API doesn't cover everything but it does cover the exact use case I was replying to.
- manquer 5y agoMost sellers who are on Amazon platform give Amazon that information and a lot more, knowing full well Amazon will use their sales data to launch an Amazon Basics competitior. It is a sane business approach when you are a pragmatic business who knows the limits that constrain your business. Either the content company is going to build a simple API (could be just a static CSV file hosted on S3 or whatever) with useful information or try to monetize/hide this information and force scapers to use the website . A bot is always going to win unless you want to make users also a lot of friction. In the era of deepfakes and fairly robust AI tooling the difference between bot action and humann action is not all that much. If you are going to be agressive with captcha , IP blocks and other fingerprinting, users who get identified false positive.or annpyed would leave. When the cost of losing those users is more than allowing access to scrapers,you would absolutely setup the API.
- 1cvmask 5y agoWhat is your site may I ask? Just curious about the difference in value from using your API and web scraping as there is a cost to web scraping as well.
- bryanrasmussen 5y agoIf you make your scraper well, and it counterfeits being a real user believably, you end up with a solution that can be tweaked as needed to handle whatever traps people put in to try to defeat your scrapers. If you make your api client well, you don't have the problems of a scraper - but if the api owner decides to change rules for api and you can't do what your business is based on being able to do (think of api owner as Twitter) then you need to make a scraper.
- gmanis 5y agoIs it not viable to put majority of your data behind a login and so the bots only get a very limited snapshot while legitimate users get it through a free login? I’m asking this because I’m going through very similar situation and would love to see other opinions around this.
- weird-eye-issue 5y agoYou are defining legitimate users as those that have a valid session cookie? Good luck
- halfmatthalfcat 5y agoMaybe the API terms/cost are prohibitive? I'm sure there's some equilibrium where they would rather pay you than go through the trouble of scraping.
- kulikalov 5y agoMaybe docs or infra are unbearable
- aninteger 5y agoWait, why wouldn't you have rate limiting on your API? Providers like Cloudflare offer this although I guess you could roll your own too since our industry loves to reinvent the wheel.
- throwaway2993 5y agoI wrote a scraper a couple of years ago to get a single data point from a website where my client was already a paying customer. This website had an API, which they were also paying for, but the API didn't cover that data point, so at the time they had one of their admin people populating that missing piece of data manually, which was taking them around ten minutes a day. I asked them if my customer could pay to access this data point via their API and they quoted 3600 EUR/month! Enter the scraper...
- deleted 5y ago[deleted]
- taytus 5y ago>where some amateur did real damage to us If an amateur can do damage to you, then I have some bad news for you...
- Goronmon 5y agoIf an amateur can do damage to you, then I have some bad news for you... I believe the point wasn't surprise that damage occurred at all, but frustration that damage can occur just out laziness/ignorance rather than malice.
- scarygliders 5y agoIndeed, that was precisely their point, and "bad news for you" is disingenuous as there are many techniques used by incompetent, or just downright unethical and greedy scraper companies which, no matter how robust the target is, can still give it a major headache. I've witnessed a site being basically DOS'ed due to particularly greedy and aggressive mass scraping attempts.
- ebbp 5y agoPrecisely this, thank you.
- convolutionart 5y agoThis is nonsense. It's always easier to destroy than to build/mantain. If you got any real advice, by all means...
- ebbp 5y agoTo be clear, they did “damage” was to our bottom line. Most sites don’t capacity plan for random cliff walls of 2-10x traffic (clearly we should!). We’re scalable enough to handle that traffic after a period, but a) it caused intermittent periods of low availability (costing us money because we didn’t generate income the way we normally do) and b) cost us money from scaling all our services up. It’s just selfish. If you’re going to take the product of other people’s work in a manner they don’t consent to, at least do it in a way that doesn’t cost them twice over.
- marginalia_nu 5y agoBots are one of those things that are easy to build and hard to get right, and there's really no way of preparing for the chaotic reality of real web pages other than fixing the problems as they show up. Weird and unexpected interactions are going to happen. Crawling the real web involves navigating a fractal of unexpected, undocumented and non-standard corner cases. Nobody gets that right on the first try. Because of that I do think we need to be a bit patient with bots. At the same time, even as someone who runs a web crawler, I have zero qualms about blocking misbehaving bots.
- chillfox 5y agoI kinda feel like rate limiting your request to individual domains and IP addresses is an easy thing that goes a long way towards getting it right.
- marginalia_nu 5y agoThere are still snags with that. Stuff like redirect resolution is very easy to overlook. You may think you're fetching 1 URL per second, but if you are using the wrong tool and you're on a server that has you bouncing around like in a pinball machine and takes you through a dozen redirects for every request, the reality may be closer to 10 requests per second. On top of that, sometimes the same server has multiple domains. Sometimes the same IP-address serves a large number of servers (maybe it's a CDN).
- RobSm 5y agoIf you build your site in a way that multiplies each request 10x, well then that's what you get. Don't do that and you won't have issue with requests. Or handle those requests properly. There are solutions to that. You know how many requests your local google CDN gets? They know how to manage load.
- marginalia_nu 5y agoMost pages have at least a http->https redirect, many contain a lot of old links to http content. Usually it's error pages that really drive the large redirect chains. They often have a vibe of like some forgotten stopgap put in place to help with some migration to a version of the site that is no longer in existence. Of course you don't know it's an error page until you reach the end of the redirect chain.
- deleted 5y ago[deleted]
- kulikalov 5y agoWhy not create api endpoint and charge mild cost for that data? You’ll make money instead of spending it.
- scarygliders 5y agoDo you honestly believe all site scraper people/companies are ethical enough to go to whoever pays /them/ to scrape data from a competitor's site and say "oh they offer an API to access this data let's pay for that", instead of "why pay for that data when we can scrape it right off their site"? Also, not all types of company will provide API endpoints. It all depends on the type of site - for example, an online shop might not wish to provide easily accessible data on offered products and prices, to their competitors who may wish to undercut them. Why would an online shop do that?
- kulikalov 5y agoEthical - of course not. Practical. Valuable public data is going to be scraped - this is inevitable. Even paywalled or signup protected valuable data is going to be scraped. Why not sell valuable data for reasonable price then.
- zivkovicp 5y agoWell, you don't need an api, just a CSV file with a catalog. The scraping company WILL use the API/CSV file... they will probably also still charge their customer for scraping, so it's a win-win :D You can think of it this way, the prices and product data are publicly visible already on the website, there are no real secrets, none of it is password protected. You can be principled and insist on blocking bots and spend a lot of time and money on tools, people, and ultimately hosting because the bots will always win; or you can offer the data for free/minimal fee and serve it with almost zero cost and cache it so you can do that with a micro sized server. You can always lie about some of the prices if you want, but you will just encourage bots again. Ethics are nice, but let's be honest, very lacking. Sometimes it's better to be pragmatic.
- scarygliders 5y ago
- scarygliders 5y agoRight with you there. I had a particularly bad time not so long ago, when a customer's site - a shop - was brought to its knees because someone, probably a competitor, hired some scraper-company of some sort to scrape every product and price. The scraper would systematically go through every single product page. And by scraper, I mean - 100's of them. All at the same time, using the old trick of 1 scraper requesting 3 or 4 product pages at a time then pausing for a while. They used umpteen different IP address blocks from all over the globe - but mainly using OVH vps IP address blocks from France. Now, maybe if they'd just thrown, say, 5 or 10 of the scraper "units" at the site, no one would have noticed in amongst Googlebot (which they wanted to use anyway because they are using Google Shopping to try to bring in more sales). But no. This shower of arseholes threw 100's of scraper "tasks" at the site. They got greedy. Now, the site was robust enough to handle this load - barely - which was massive, however, having to do that /and/ also handle normal day-to-day traffic? Nah. The bastards got greedy and like you I spent a few days unfucking the damage they were causing. Seriously, I hate scrapers. I hate the people who make scrapers. I hate their lack of ethics. Fuck those guys.
- _lqaf 5y agoIn a past life, we were consulting with a startup that offered a subscription data service. They were very sensitive about scrapers, especially on the time limited try-before-you-buy accounts, which competitors were abusing. At their request, we built a method to flag accounts for data poisoning. Once flagged, those accounts would start getting plausible-ish looking garbage data. It was pretty effective. One competitor went offline for a few days about a week after that started, and had a more limited offering when they came back up.
- scarygliders 5y agoThat's a good way of going about dealing with this kind of abuse indeed. Wish I'd thought of doing that at the time, but due to the nature of this shop you didn't need a user account to browse the products/prices. I'm now making an entirely new shop for them - I shall bear this in mind. Thanks for that!
- 5y ago
- paco3346 5y agoI'm right there with you. I'm the lead engineer for an automotive SaaS provider (with ~6000 customers and ~4 billion requests per month) and we recently started moving all our services to Cloudflare's WAF to take advantage of their bot protection. We were getting scrapes from botnets in the 100000+ per minute range that was affecting performance. We chose to switch to the JS challenge screen as it requires no human interaction. We now block 75% (estimated to the best of our knowledge) of bot traffic but some customers are livid over the challenge screen.
- deleted 5y ago[deleted]
- EdwardDiego 5y agoWhat were they scraping, if I can ask? Was it targeted or just wget -r style?
- paco3346 5y agoIt was a hybrid of low-effort vulnerability scanning and targeted inventory scraping. Many dealerships in the automotive space will pay gray-hat third parties to scrape and compile data on their competitors. The irony for us as a provider is that it's one of our customers (party A) paying a third party to scrape data from another one of our customers (party B) which in turn affects the performance of party A's site. We've started blocking these third parties and directing them to paid APIs that we offer.
- devwastaken 5y agoIf an amateur can do that to your service by scraping, imagine what someone can do if they actually intend to do you harm. With cloud pricing models someone could find a little misconfiguration or oversight and put you in the hole in operating costs. Anti-abuse is a necessary design when your service is exposed to the internet. Not saying that doesn't suck - it does, it's why many ideas don't work in practice as an online service.
- krzyk 5y agoAs a programmer that just sometimes wants to check if given item is available in store I would like to be able to use API for that. But if it is not available one has to scrape.