8 ms·
I think of web scraping as nothing less than automation of human work. There really is no hacking or unauthorized access involved. It's either me, using a brows
by mnmkng 6y ago
I think of web scraping as nothing less than automation of human work. There really is no hacking or unauthorized access involved. It's either me, using a browser, to read something publicly accessible on the internet, or it's my computer, which I've programmed to read those things for me, so I can focus on creative work or spend time with friends and family.
This applies not only to personal life, but businesses as well. Instead of hiring 100 interns to collect some data manually, you hire one programmer to automate that data collection process.
I think that the ability to automate tasks on the internet is absolutely crucial to further development of our society and limiting it in any way will be detrimental to the world as a whole. The amount of information these days is so vast, that no human labor force could possibly analyze that information and use it to drive our progress.
Disclaimer: We run a web scraping platform (https://apify.com https://apify.com)
- sokoloff 6y agoWhile I have substantial agreement with your point of view, when web content is substantially ad-funded, automated scraping effectively bypasses the “payment”. I say this as someone who runs an adblocker, installs them for family, and doesn’t derive income from hosting ads, so I’m not pro-ad; I just realize that while it’s not a crime for me to dump the whole bowl of waiting room candies into my backpack, that’s going to be frowned upon.
- simplecto 6y agoThis is orthogonal to the purpose of scraping. If I run a search engine which potentially links real human eyes back to you, then should I pay the "ad" toll as well? I don't believe so. I do believe there is an agreeable middle-ground, but Google walked away from that conversation years ago.
- krageon 6y agoAdvertisements are harmful to every layer of society, mostly because they prey on you to instill desires that you probably wouldn't have had on your own (it being the case that this is their entire added value proposition). They should not be tolerated, and the fact that they ever were is a travesty. That being the case, any argument founded on "but advertisements" does not hold water.
- LunaSea 6y agoIt's not founded on "but advertisements" it's founded on "but paying for content". The fact that websites are developed by people (wages), hosted on infrastructure (hardware renting) and require countless other jobs is somehow magically forgotten in these discussions.
- perched_robin 6y agoWhatever can be killed by adblockers should be. 95% of the web consists of bloated javascript made by overpaid frontend devs that's less functional than a plain text file. So the ad money runs dry and nobody can afford to push 20 megabytes of javascript and fonts with every page view anymore. What a tragedy.
- krageon 6y agoIt's not up to me or anyone consuming content to help with figuring out the right way to pay for it. What is up to me is the choice to not be potentially exposed to malware any time I go to read the news. While you may have a point, that point just doesn't matter within the context that we're talking about. Making companies money is not my responsibility.
- LunaSea 6y agoIt absolutely is your responsibility if you are consuming the content. Consuming the content means agreeing to the premise that the content is paid for using ads. If you disagree with said premise you may happily browse another website. Malware is already illegal to install and you are free to sue the website for damages. I don't like ads either but trying to justify that their should be a choice of browsing a website without the ads it hosts is ludicrous.
- krageon 6y ago> Consuming the content means agreeing to the premise that the content is paid for using ads. It is in fact not this way, because the content arrives with or without the ads. In the EU EULAs (the dystopian construct that you'd expect to enforce that bit of lunacy) that purport to apply to content you have already accessed are not legally valid. Leaving the legal interpretation aside, me doing one thing doesn't mean consent for something else. Believing otherwise is both unethical and amoral, stances I don't hugely feel like interacting with.
- shadowprofile77 6y agoWhen you dump the bowl of candies into your backpack, there are no more candies left for anyone else. On the other hand, when you scrape, all the original data and its underlying use case remain just as they are. You simply having scraped took nothing away from anyone.
- mikem170 6y agoI wrote a scraper for myself only to keep an eye on several used truck websites. I was able to do this for craigslist, so I could run it every couple days and check the nearest 75 cities, the list of truck ads, filtering on/out certain keywords, and the pages for the interesting trucks, making a report for myself to help find the vehicle I want. I could not do this for any of other several large sites advertising used trucks, like commercialtrucker.com, truckpaper.com, and machinio.com. I kept bumping into artificial javascript and captcha limitations that were not worth my time to try to work around. The thing is that I will probably find what I want on a craigslist site that I can scrape, I'm getting a lot of great info from them, and I'm not going to bother with any of the ad-based sites, it takes to olong to run all the manual searches. They have effectively done a disservice to their (presumably paying) customers who want to sell their vehicles.
- sokoloff 6y agoIf AutoTempest covers your search criteria, I’ve found them to be excellent to do repetitive searches. Somewhat fortunately for you, they suck for craigslist searches, but have great coverage for all the other major sites.
- artembugara 6y agoI totally agree with the first part where you say it's fine to scrape. It is just an automation. But you don't mention what you do with this data after. 1. You store it somewhere (not in your brain) 2. You extract value out of it (directly or indirectly) That's why I understand why it is a problem for those who publish this data. One thing I always try to say to everyone who argues about web scraping: web scraping is not a problem, problem is what you do with the information you scraped. Disclaimer: We crawl the web for news (https://newscatcherapi.com/ https://newscatcherapi.com/)
- jagged-chisel 6y ago> That's why I understand why it is a problem for those who publish this data. I still don't understand why it's a problem for those who publish the data. If the "data" is facts (e.g. lists of values associated with objects, like the colors available on automobiles), this data is not protected by copyright, there is no remedy if it's republished, and if a business relies on limiting access to this data, that business needs a new business plan. Charge for access to cover the costs of obtaining and organizing the data, but understand that clients are legally allowed to make copies and use them however they see fit, even if that impacts the supplier's ability to charge for access. If the "data" is prose, it's under copyright and republishing without a license has remedies under law. Maintaining copies of articles for the purpose of processing them to obtain other data (e.g. how many nouns? what adjectives are near those nouns? etc...) isn't protected.
- petr25102018 6y agoDatabases are actually protected under some copyright laws in some countries.
- jagged-chisel 6y agoFor copying the database wholesale, yes. But the individual records, formatted on a web page, and scraped out by a bot is not copying the database files or its schema wholesale. They're still individual facts.
- Someone 6y agoI think of placing a zillion cameras in public places as nothing less than automation of human work. Instead of hiring a million agents to record what the population is doing, you hire one programmer to automate that data collection process. Totally different? Yes, but it shows that scale can affect whether something is OK to do or not, at least for some (I accept that people watching the streets can be useful, but also have my doubts about doing that at scale, whether by cameras or by hiring a million agents) With both cameras and copy-pasting web content, there’s the issue what you do with the data. If, for example, I start scraping all the articles on a newspaper’s web site, publish them on a web site, adding my own ads, most people would think that shouldn’t be legal. If you agree, we’re now haggling over the price (https://quoteinvestigator.com/2012/03/07/haggling/ https://quoteinvestigator.com/2012/03/07/haggling/). That’s where things get difficult, but I think the entire spectrum from white (scraping for this goal is fine) via grey to black exists.
- narag 6y agoIf, for example, I start scraping all the articles on a newspaper’s web site, publish them on a web site, adding my own ads, most people would think that shouldn’t be legal. Publishing? Ads?
- fakedang 6y agoRepublishing, duh.
- narag 6y agoOK, let's reformulate that: would people be opposed to just scrapping or to the illegal republishing and adding ads? Another similar trick: "If someone will block ads and then murder the publicist and burn his house, most people would find that it should be illegal".
- fakedang 6y agoAdvertising isn't some zero-sum game that suddenly makes you lose because you happened to see it. Republishing is - the original publisher loses out on the value of his publication because that means either reduced ad revenue or reduced subscription. I'm guilty myself of reading republished articles, but it is a loss for the publisher.
- Wowfunhappy 6y ago> I think of web scraping as nothing less than automation of human work. I agree that scraping should be allowed. We wouldn't have Google otherwise. However, there's something to be said for the fact that some activities are okay at a small scale but become problematic at a large scale—particularly the scale which becomes possible when a task is automated. For instance, I recently bought an expensive camera and have been having fun walking around the city and taking interesting photos. Many of the photos have people in them. I don't think there's any harm in this. I store my photos in Aperture, which automatically performs facial recognition on everything in my library. Perhaps some day, if I take enough pictures, Aperture will notice that the same stranger is present in two completely different images. That might be kind of cool—I can see myself fancifully trying to imagine this person's life story. I don't think there would be any harm in that, either. However, if I aggregated millions of photos from different sources, and used them to track people's movements across the city, that would clearly be a huge problem! Sure, it would be merely automating the work a human could do, but the scale just changes everything!
- bobobob420 6y agoCan the case be made that recording in public is a right (as if always should be) but trying to track where everyone is at every point of time is stalking at a mass scale which should be illegal as stalking on a one to one scale is illegal? For reference this is what one site (findlaw) has given as what constitutes the crime of stalking: The crime of stalking can be simply described as the unwanted pursuit of another person. Examples of this type of behavior includes following a person, appearing at a person's home or place of business, making harassing phone calls, leaving written messages or objects, or vandalizing a person's property.
- travisporter 6y agoThis sentiment exactly. Sure the information is free. But to find my property records, phone number, and "aggregate them" like Spokeo, fastpeoplefinder and similar sites, is akin to digital stalking IMO
- 6y ago
- diegoperini 6y ago> There really is no hacking or unauthorized access involved. I'm not sure about that. If I had a phone book full of phone numbers (those heavy ones from 80s), would calling every number in that book to find the one I'm looking for be legal/ethical? P.S: I agree that "Web Scraping Is Vital to Democracy".
- arthurcolle 6y ago> If I had a phone book full of phone numbers (those heavy ones from 80s), would calling every number in that book to find the one I'm looking for be legal/ethical? Sure, why not?
- diegoperini 6y agoI have no idea, I've never done it.
- fakedang 6y agoI can tell you haven't cold called ever in your life.
- drdeadringer 6y ago> I think of web scraping as nothing less than automation of human work. I did this with Python regarding market prices for ETFs, Stocks, and Mutual Funds. I wrote Python scripts, one of which was to web-scrape current market values, investment distribution, &c for ETFs, stocks, and Mutual Funds for whic I was invested in. I then would have to manually port that into a spreadsheet for "my own special graphs" and such. I'm sure there are places online which would do this for me if I logged in, entered in all of my information and more, and provide that all for me ... but this story should sound familiar. And due to personal reasons, I've been lagging behind for far too long. My Point: This information is publicly available and and it serves my purposes. There's no reason why I should not be able to do this. Yes, it's my fault for not using the information [and yes, the information dies with me], but the point is that I automated the gathering of that information for my own self -- and "everyone" else on the planet has that same information. I never felt like I was illegal doing this. I'm happy to know if there is a "save" way to do this type of aggregate situation otherwise.
- deleted 6y ago[deleted]