5 ms·
"How do I stop all these dinner guests from eating this lovely pie I set out on the table?" I remember working hard on a project for a year, then releasing the
by monodeldiablo 9y ago
"How do I stop all these dinner guests from eating this lovely pie I set out on the table?"
I remember working hard on a project for a year, then releasing the data and visualizations online. I was very proud. It was very cool. Almost immediately, we saw grad students and research assistants across the globe scraping our site. I started brainstorming clever ways to fend off the scrapers with a colleague when my boss interrupted.
Him: "WTF are you doing?"
Me: "We're trying to figure out how to prevent people from scraping our data."
Him: "WTF do you want to do that for?"
Me: "Uh... to prevent them from stealing our data."
Him: "But we put it on the public Web..."
Me: "Yeah, but that data took thousands of compute hours to grind out. They're getting a valuable product for free!"
Him: "So then pull it from Web."
Me: "But then we won't get any sales from people who see that we published this new and exciting-- Oh. I see what you mean."
Him: "Yeah, just get a list of the top 20 IP addresses, figure out who's scraping, and hand it off to our sales guys. Scraping ain't free, and our prices aren't high. This is a sales tool, and it's working. Now get back to building shit to make our customers lives easier, not shittier."
Sure enough, most of the scrapers chose to pay rather than babysit web crawlers once we pointed out that our price was lower than their time cost. If your data is valuable enough to scrape, it's valuable enough to sell.
The only technological way to prevent someone crawling your website is to not put it on a publicly-facing property in the first place. If you're concerned about DoS or bandwidth charges, throttle all users. Otherwise, any attempts to restrict bots is just pissing into the wind, IMHO.
Spend your energies on generating real value. Don't engage in an arms racw you're destined to lose.
- huffmsa 9y agoExactly. If you have data so valuable that people want to take it, it's going to be a lot easier to figure out how to sell it to them (they're probably not professional web scrapers, it's just a means to an end.) Than to waste yours, theirs and everyone else's time going tit-for-tat keeping them away from your data.
- lazyjones 9y ago> Sure enough, most of the scrapers chose to pay rather than babysit web crawlers once we pointed out that our price was lower than their time cost. If your data is valuable enough to scrape, it's valuable enough to sell. Cool story, bro. Some scrapers will buy your data if it's good enough, but most can't be identified (good luck with those Tor/AWS/dynamic IP users) and some will just resell your data at a lower price. So as a general strategy against scraping, this is useless.
- huffmsa 9y agoPeople with subscriptions are going to GIVE your data away for free. Ever heard of KAZAA? Napster? BitTorrent? All sites where people who have usually purchased data make it freely available to others.
- huffmsa 9y agoSide bar, if only Twitter would realize this and turn the firehose back on and charge $x per month for API access. They'd be profitable in a month.
- blowski 9y agoGiven the number of times people assert this, I think Twitter must have explored this option and decided it wasn't a viable option. I don't know why, but I'll assume that they're smart enough to have decided against it for a good reason.
- laughfactory 9y agoI would not make the assumption. I've been working for corporate America long enough to know that there's an astonishing level of incompetence many places. I'm sure Twitter is no exception. I'm sure _someone_ at Twitter has considered this, but perhaps they work for someone who doesn't "get it." Among many other possibilities.
- danpalmer 9y agoTwitter do charge for the firehose, and I hear it’s a reasonable amount of revenue. A person can’t go and buy it on a credit card, but there are authorised sellers, enterprise sales people, etc. Bear in mind that it’s a large technical feat to be able to ingest the firehose effectively, so it’s not really suitable for consumers.
- jorgec 9y agoYou are comparing two different things. To have a free website is not the same than to unlimited grants such as a guy to taking the whole pie. Or worst, lets say that you writes a free essay and everybody could read it and share. However, somebody takes it, deletes your name and put his name instead.
- monodeldiablo 9y agoThen you, my friend, have a copyright issue. There are lots of legal tools for addressing this, and they're all cheaper than wasting your energy trying to counter the bots. If your free content is the sole source of your online revenue, then your reputation is your business. Nefarious crawlers who rebrand your content cannot, by definition, beat you to market. So you stand to gain much more by writing great content, driving readers to it as soon as it's posted, and filing the occasional DMCA takedown than trying to compete with plagiarists in a game of whack-a-mole.
- blowski 9y agoI run a property website that lists properties for sale, rent, etc. A big part of my job is importing feeds, scraping sites (with permission) - and preventing others scraping our site. I know that some people will scrape, but I make I try to make it unprofitable for them to do so. We do a bunch of other stuff, like adding fake properties so we can check who is scraping our content, and using tarpits. Developers always argue with me that it's pissing in the wind, that "content wants to be free", that you shouldn't bother even trying to prevent it since it's inevitable. And yet it has helped. We did a split A/B test on preventing scrapers, and it turns out that it's quite effective.
- pharrington 9y agoYou are deliberately devoting energy and resources towards removing value that already existed in your product. If you want to charge rent (via a subscription service etc), then do that, and be clear that you're in the business of charging rent. Don't conflate selling with renting - that just leads to a product gimping death spiral.
- blowski 9y agoValue to whom? If a competitor scrapes all my content and then gets the leads instead of me, they've removed value from my bank account, and that's the main value I'm concerned about. This is not a hypothetical concern - it happens a lot.
- pharrington 9y agoThe competitor did not remove anything from your bank account. If your assertion is that you will lose in economic competition if you make certain information publicly accessible, why are you publicly providing that information while expecting to make money from it in the first place? To me, it sounds like your business lacks some fundamentals.
- jklein11 9y agoIf your secret sauce is the data that you have why not sell it to your competitor?
- crispytx 9y agoLOL, your boss sounds badass.
- calafrax 9y agoThat is good advice from a technical standpoint but from a legal standpoint creating security features that prevent scraping gives you a clearer cause of action against scrapers so if someone starts making a lot of money off your content you get leverage to force them to pay for it.
- monodeldiablo 9y agoYou can achieve the same thing, from a legal perspective, with a well-placed statement of IP ownership.
- calafrax 9y agoIt is not that open-and-shut. Unauthorized access to a computer system is a different legal category than IP licensing.
- jerkstate 9y agoYou’re making a valid point for many cases, but there are definitely negative–value scrapers out there; let’s say you run a publishing platform and you see scrapers scraping your users content and then see that your site’s content has been rehosted for ad clicks. You can’t really license the content for this purpose and it’s bad for your brand and bad for your users.
- Waterluvian 9y agoThat's a great point. I'm distracted a bit by how you built a project for a year without understanding what the business strategy was going to be. Some big time comms breakdown there eh?