43 ms·
This tends to be a very unpopular opinion around here, but in almost all cases I find Internet scraping to be unethical and downright malicious. I'm not saying
by blantonl 2y ago
This tends to be a very unpopular opinion around here, but in almost all cases I find Internet scraping to be unethical and downright malicious. I'm not saying all cases, but I'm saying almost.
A lot of the actors involved tend to be hustle culture types who think they are OWED your data, regardless of the ethics, laws, being a good citizen, whatever. They will blatantly disregard terms of service and hide behind massive setups such as these to circumvent protection etc.
And the problem is, if you run any sort of business or service that is data oriented, there will be thousands of people that will do this, which will cause you to devote enormous amounts of time, effort, money, and infrastructure just to mitigate the issues involved with data scraping. That's before you are even addressing whether or not these people are "stealing" your data. People who feel they are entitled to the crux of your business aren't bothered by being nice in the way they take it - they'll launch services that will cripple infrastructure.
Whenever I deal with a scraping process that decides it wants my entire business, and it wants all of it RIGHT NOW, or in 5 minutes, I want to find the person and sit them down in a room and tell them "hey, develop your own ideas and business. Ok? Thanks"
And if you think this was a problem before, it's exponentially worse over the past few months with every Tom, Susan, and Harry deciding they must have all your data to train their new LLM AI model. By the thousands.
- vouaobrasil 2y agoI absolutely agree. In fact, I think the problem is that like everything, there is an optimal point for efficiency, and crossing that line by making things "too easy" when it comes to data means too much power for one person to handle ethically. Absolute power may corrupt absolutely, but near absolutely power also corrupts quite nicely, too. In short, we should have limits to amount of scraping possible, simply because humans can never be trusted past a certain point to remain ethical. After all, ethics at its first approximation is only a mechanism to improve societal cohesiveness, and it only works as long as the person doesn't have enough power to "do away" with society.
- jumby 2y agoWould you make the same argument of the inverse: data gathering?
- vouaobrasil 2y agoYes, I would. There is a law of diminishing returns for all technical and scientific inquiry.
- deleted 2y ago[deleted]
- flir 2y agoThere's a lot of local history locked up in facebook's nostalgia groups. I want to archive it in an open format. I want to grab new rental listings and put them in an RSS feed, so I only look at each one once. That's my uses for data scraping right now. If that destroys someone's business, I don't actually care. Maybe it's selfish, but my right to re-format data for my own convenience outweighs their right to make a profit.
- throwaway11460 2y agoNot that I think you shouldn't do it or you're doing something wrong, but describing it as a right irks me the wrong way. You don't have any right to expect someone else's computers to work for you.
- flir 2y agoI'm not sure how to phrase it except in terms of competing rights, but I take your point. At the point where I'm scraping, the data's on my computer though.
- solarkraft 2y agoYou could call them interests . It's often in a business's interest to format data in a specific way to make money, for example interlacing it with ads.
- flir 2y agoNice.
- blantonl 2y agoIf that destroys someone's business, I don't actually care. Maybe it's selfish, but my right to re-format data for my own convenience outweighs their right to make a profit. Exhibit A
- 2y ago
- brigadier132 2y ago> hustle culture types It seems like you have this imaginary strawman that you hate and it seems like that's the foundation of why you dislike this.
- blantonl 2y agoNo. The foundation of why I dislike it is simple. If I own some data, then I get to dictate the terms of how that data is used. Period. “Hustle culture types” is simply a little anecdote about the types that would look you in the eye and tell you they are entitled to disregard what I said above. They’ll usually wrap it in some altruistic bs to justify as well.
- some1else 2y agoServing HTML will get you scraped. Your terms don't overrule fair use.
- throwaway11460 2y agoWhy do you put it on the open internet if you don't want machines to find and read it? ToS is nice but you can't expect that it applies - the user (of the machine doing the scraping) might be a child which makes the potential contract automatically void, for example. Also, there are people under jurisdictions where such things have no power, or that don't recognize your rights to the data. And the whole thing of putting data out publicly and then just expecting machines to see the pile of data and go "oh so where do I sign the ToS?" is weird... Just put it behind a rate limited API key...
- blantonl 2y agoWhat makes you think putting data on the Internet all the sudden means I unilaterally surrender the rights to my intellectual property? If I choose to make my data available to some businesses to make discovery of it easier, and I choose to decline to allow others to unilaterally copy my data to develop a different business, that's my right. And it is unethical and unreasonable for any other person to assume otherwise that they are entitled to the same rights I granted someone else. If I own some data, I get to the be arbitrator of the who/what/when/where on the use of the data. Period.
- deleted 2y ago[deleted]
- malwrar 2y agoIf your business is just that you have a bundle of information and expose it over an open website, I’m not really sure how you’re able to maintain a mentality that you are somehow entitled to ownership of that information. You already put it out there, it’s now public, any illusion to exclusivity is now gone because anyone could come along at any time and make a copy without your knowledge. A moral position on this issue is even more confusing to me. Do you think that you e.g. own the knowledge on which radio frequencies are used where? Do you think you have a moral claim on ownership of (presumably unpaid) user-submitted information? I think the only legitimate moral grievance you have is high traffic volumes from inconsiderate scrapers.
- blantonl 2y agoDo you think you have a moral claim on ownership of (presumably unpaid) user-submitted information? You damn right I do. I own, develop, and maintain the entire system that enabled the body of works to exist in the first place. Do you think that you have a claim on ownership of the data because you drove by, saw what you liked, and decided that now you'll just rip the baton out of my hand?
- malwrar 2y ago> You damn right I do. I own, develop, and maintain the entire system that enabled the body of works to exist in the first place. I don’t think that meets the bar. Running a website is absolutely not equivalent to the collective effort people put in to populate that website with the information that actually gives the overall artifact its value. There is a large history of outrage when similar information repository websites with user-generated content violate expectations of openness. Nevermind the fact that the actual information itself isn’t even private or proprietary, just obscure and distributed. > Do you think that you have a claim on ownership of the data because you drove by, saw what you liked, and decided that now you'll just rip the baton out of my hand? I wouldn’t claim ownership nor want to, when I scrape stuff I usually just want information in a different format. I’m confused as to how you think you can even “own” data to begin with. Suppose that your users uploaded songs instead of RF info, do you believe you own their music solely because they chose to share it on your site? Do you think your users would believe that?
- greenbandit 2y agoI use web scraping to identify and monitor fraud. Exhibit A: https://archive.ph/0ZUA8 https://archive.ph/0ZUA8 This website is used to recruit people to set up "lead generation" Google Business Profiles and leave paid reviews. Exhibit B: https://archive.ph/WWZuw https://archive.ph/WWZuw This is an example of the Craigslist ad used to initially attract people to the website above. Exhibit C: https://archive.ph/wip/7Xig4 https://archive.ph/wip/7Xig4 This is one of the Google Maps contributors which left paid reviews. If you start with the reviews on that profile, you'll find a network of Google Business Profiles for fake service-area businesses connected through paid reviews. Web scraping allows me to collect this type of data at scale. I also use scraping to monitor the status of fake listings. If they are removed, the actor behind them will often get them reinstated. This allows me to report them again.
- blantonl 2y agoI don't care if you use Web scraping to solve the Israeli / Palestinian conflict. You're not entitled to anyone's data, computers, services, etc because you've decided for altruistic reasons that it is appropriate. Cool use case. Love it. Fascinating stuff. But if Google told you to stop, would you? Or would you instead decide to build a 5 server cluster of 200 4G modems spread across continents to continue your work? Because if you did I would assume that you've decided to move on from a cute little altruistic process into a commercial use of someone else's data to make a profit.
- greenbandit 2y ago> cute little altruistic process Maybe it is not the opinion which is unpopular, but the way it is being presented.
- ansc 2y ago>I don't care if you use Web scraping to solve the Israeli / Palestinian conflict. Maybe you should though. It's always worth it to think about which giant's shoulder you're standing on. It's giants all the way down.
- dmkii 2y ago
- tengbretson 2y agoIs it unethical for a mouse to eat the cheese without triggering the trap?
- hipadev23 2y agoI find it aptly hilarious that your own business model at broadcastify.com is recording publicly accessible radio broadcasts and then selling access to those recordings for commercial gain.
- blantonl 2y agoWhy is that hilarious? We developed an entire community, infrastructure, system, architecture, everything, from scratch, and provide access to something that never existed in the first place on the Internet. That's a significant key difference here. This would be analogous to you thinking ancestory.com is "aptly hilarious" for arguing against someone just scraping their site for content. What makes you think you should be entitled to drive by the very unique house that we built, and pointing right at that house and saying "I think I'll take that all of that for myself!"
- hipadev23 2y agoBecause you fail to see the very obvious parallels to scraping. I’m not criticizing your business (I think you provide a valuable service) but your hypocritical stance on what forms of publicly available information are allowed to be gathered and repackaged. Google’s original (and OpenAI’s) business model was also building a scraping infrastructure, system, and architecture, from scratch — and providing access to something that never existed in the first place.
- blantonl 2y agoIt's completely perpendicular, not parallel. Public safety communications are radio waves that are broadcasted and the ability to passively monitor them is enshrined in United States law. That is a massively key difference. If I was sending data into your home from my infrastructure without any action from you whatsoever, and you were reaching up into the air and gathering it and repackaging it, AND the law said that I have no intellectual property rights to said data, then that's a whole different story.
- 2y ago
- dale_glass 2y ago> Whenever I deal with a scraping process that decides it wants my entire business, and it wants all of it RIGHT NOW, or in 5 minutes, I want to find the person and sit them down in a room and tell them "hey, develop your own ideas and business. Ok? Thanks" That's a lot of righteous anger for somebody building a business on top of other people's data. "Broadcastify is the worlds largest source of public safety, aircraft, rail, and marine radio live audio streams." I have no sympathy whatsoever. You're just complaining about the very thing you're doing. If it's fair for you to do that, it's fair for others to do it to you.
- blantonl 2y agoThey volunteer to provide the data to us. Every single last one of them. Nowhere in our business model did we make the conscious decision to say "hey, look at that business, they have something, and I'm going to take it."
- bsuvc 2y agoReading public website data is not "taking it". It is still there. Observing publicly available information is not theft, nor is it illegal. Of course copyright rules apply, but that is for if you reproduce something.
- blantonl 2y agoreproduce something No one is developing a 5 server cluster with 200+ 4g modems to observe publicly available information. They are using said cluster to deliberately work around blocks, rate limits, and restrictions on scrapers who are scraping content solely to reproduce the data and use it for commercial purposes (make money)
- schlipity 2y agoAren't you also volunteering your data? Don't browsers just talk to your webserver and say "Hey, what do you have?" and your site responds in kind.
- juunpp 2y ago> but in almost all cases I find Internet scraping to be unethical and downright malicious. The Web (you said the "Internet", but you meant the Web) was not envisioned to be a commercial space. Your statement is antithetical to the original idea of the open Web. It's when the MBAs joined the party circa 2k and decided to profit out of it that all of these confused and wrong opinions about what the Web should be arose and that lead to the situation today. Your statement is a vast display of zero historical context. MBAs are obviously not very concerned with history. They just want to protect their own little turd for their own little profit and vanity, which is why they now put it behind a paywall, JS, and anti-bot proxies.