15 ms·
Taking action against scraping for hire
- paultopia 4y ago"Scraping attacks" LOL
- sophacles 4y agoWhy not? weev was put in jail over incrementing a number in a url. Surely writing software to put values into urls is even worse.
- sneak 4y agoLet's be clear and accurate: technically weev was put in jail for conspiring on IRC with JacksonBrown. JacksonBrown was the one who wrote a PHP script that incremented a value in a URL (and appended a valid Luhn check digit following incrementation). Conspiracy to access a protected computer system - that is, typing on IRC. weev didn't write any of the code or access the API.
- NelsonMinar 4y agoOctopus sounds really useful; is there an open source equivalent? I'd love to be able to scrape my own data on Facebook. Their data export feature is fairly good but far from complete.
- rustdeveloper 4y ago“This industry makes scraping available to individuals and companies that otherwise would not have the capabilities.” - seems like web scraping companies are doing a good job :)
- jhoelzel 4y agoThe phone charger makes engery available to individuals and companies that otherwise would not have the capabilities. ;)
- theincredulousk 4y agoMaybe some irony here as IIRC Facebook started as essentially a scraping company, pulling student profiles from college websites and re-publishing it for their own profit. The scrapers have become the scrapees. The horror.
- HeckFeck 4y agoData harvesting is moral for me, but not for thee.
- mateuszbuda 4y agoIn general I agree that harvesting public data is moral. I think that in these particular cases it's: 1) extracting data from profiles that opted for not being public (only available to logged in users) and 2) reposting scraped data (publicly?) as belonging to the guy who scraped it without users consent.
- trasz 4y agoIf they are being harvested it makes them public by definition. Unless there was a break-in.
- adolph 4y agoThe state of "opted for not being public" and 'available to any system authenticated person' seem contradictory. I appreciate that 'system authenticated person' is a smaller set than those who can access anything publicly accessible, and that the former is a subset of the latter.
- Alex3917 4y ago> extracting data from profiles that opted for not being public The tool lets you download the contact info of your friends, which you should be able to do anyway. In fact Facebook tries to trick its users into thinking they can do this with their data takeout option, but the downloaded files don't actually include any of the contact info for your contacts. Which makes zero sense, considering the entire point of Facebook is that it's a digital rolodex for storing your friends' contact info.
- slightwinder 4y agoFrom the article, it seems to be service for scrapping data you have access anyway. As long as they only handle those data to the requesting customer, whose login they used, I don't see a difference between general public, and this users personalized "public". If access is still limited to the people who have the access-rights, then I don't see a difference between accessing through the official interface, or via scrapped data.
- jacooper 4y agoThey are will using fb.com domain? I though meta is not FaceBook?....
- Silica6149 4y agoI think it's like Google vs Alphabet. Alphabet is the parent company like Meta. As for why their domain is facebook for their news site, not sure why. It would make for sense for it to be under meta instead.
- pclmulqdq 4y agoThey have to keep the walls up on their garden so they can get maximum value from harvesting.
- deleted 4y ago[deleted]
- pid-1 4y ago> scrapping attack
- mohamez 4y agoThat cracked me up when I read it lol
- Hedepig 4y agoIs this much different from LinkedIn vs hiQ?
- nojito 4y agoLogged in vs not logged in data.
- logifail 4y ago> Logged in Is this actually private data, or is it public stuff that's become annoyingly hard to view anonymously because Meta chose to stick it behind a login box?
- cupofpython 4y ago>public stuff that's become annoyingly hard to view anonymously because Meta chose to stick it behind a login box this one
- nojito 4y agoAnything behind a login gate is private data for that registered user only.
- logifail 4y ago> Anything behind a login gate is private data for that registered user only That's quite the claim, if only the login gate were either always there or indeed always not. Presuambly such "private" data ought not to be being indexed by search engines and returned to users who search? "site:instagram.com" is of the order of 228 million pages on google.com, and "site:facebook.com" is another 422 million.
- nojito 4y agopretty sure you get hit with a login gate if you navigate to the results via site:instagram.com no?
- dangerlibrary 4y agoFingers crossed they eventually get around to suing Clearview AI out of existence. https://www.nytimes.com/2020/01/18/technology/clearview-privacy-facial-recognition.html https://www.nytimes.com/2020/01/18/technology/clearview-priv...
- htrp 4y agoThis is different from LinkedIn v HiQ because HiQ was only scraping publicly available data that was generally accessible to the broader internet. In these two cases, the data is being scraped from FB/Insta using credentials that the client handed over or the mass creation of accounts solely for scraping purposes.
- squaresmile 4y agoYeah, I think this is more like the Cambridge Analytica situation.
- benwad 4y agoDid FB ever take any legal action against Cambridge Analytica? I can't remember anything about it and this sounds very similar to that (although back in those days FB's tools made this incredibly easy).
- lesuorac 4y agoNo. FBs ToS at the time [1] allowed CA to do what they did. Namely, CA didn't resell the data or give it to an ad agency. [1]: https://web.archive.org/web/20180329131546/https://developers.facebook.com/policy https://web.archive.org/web/20180329131546/https://developer...
- Nextgrid 4y agoI wish the Cambridge Analytica FUD would stop. CA's "attack" was to setup a malicious website that convinced idiots to give it access to their Facebook account using the standard oAuth2 flow. Did they misuse the collected data? Sure. But people granted access to that data knowingly. This wasn't really an attack in my view. Facebook wasn’t really complicit and definitely didn’t sell/give away any data.
- Nextgrid 4y ago> the mass creation of accounts solely for scraping purposes. Those accounts wouldn't be allowed to view private data though unless they friend/follow the person first, so they'll only still be limited to data the account holders intend to be public and available to anyone. There's also no evidence that the scraped data was aggregated at scale or commingled in any way, so even if customers provided their actual credentials which grant them access to private data of their friends, the scraper didn't share it with anyone else but them.
- PhilipA 4y ago>Octopus, a US subsidiary of a Chinese national high-tech enterprise, built a cloud-based platform designed to provide paying customers access to on-demand scraping software and services. It is interesting as how they try to position this as a Chinese attack on them.
- MangoCoffee 4y agoit look like Zack is giving up on the Chinese market.
- romanovcode 4y agoI guess after Winnie the Pooh rejected to name his children for him he got sour grapes for China.
- upupandup 4y agoIt must coincide with Christopher Wray's sudden claim that there is an active dragnet of sorts that is trying to subvert America from within much like the recent election interference of a former Tianmen square activist who tried to run for congress I think. It makes me think that there are many people on CCP's dole, rich powerful famous people are somehow beholden to the CCP in some unknown way but we can all guess correctly that they are all old white men who have previously been seen with young females.
- fxtentacle 4y agoOf course, Facebook wants to make it sound like scraping is illegal, when it generally isn't. But account hijacking and mass-creation of accounts just to access private pages are clear violations of the Facebook and Instagram ToS, so they surely can sue for that.
- dementiapatien 4y agoSince when do you get sued for breaching TOS?
- thallium205 4y agoSince when do you get sued for breaching a contract? When the offense is worth it.
- curiousllama 4y agoSince you start a business on the violation. "Since when do I get sued for taking too many free samples from Costco?" -> "Since you started taking millions of them to resell"
- jhoelzel 4y agoim not sure on american law, but if you give me those samples willingly i can do whatever i want with them. Actually this is the reason why many products come with the lable "not for resale" but i have yet to find somebody who cares about it :D
- treis 4y ago>give me those samples willingly Doesn't seem like Facebook is giving them willingly.
- golemotron 4y agoYou can get sued for anything that causes harm. Relevant life lesson: don't do things to people with money that they might perceive as harm. Corollary: Being sued is as much punishment as losing a suit for most people.
- cosmiccatnap 4y agoI would consider this appropriate if one of the largest offenders of scrapping weren't the one pretending to be the offended.
- trasz 4y agoWe need to update the law to make sure Meta loses in cases like this.
- xvector 4y agoHN is hypocritical - most commenters here are against this because "Meta bad," but at the same time, most commenters wouldn't want their posts shared privately amongst friends to be scraped and made available publicly.
- mpeg 4y agoFor that to happen, one of your friends would have had to willingly allow this tool to scrape their social network, which would include your private posts. Is the scraper to blame here, or the friend?
- Komodai 4y agolol maybe if you don't want that happening you shouldn't be using Facebook
- pawelkobojek 4y agoThere are two cases they brought up, one being web scraping and the other is making a clone website publicly displaying content from Instagram. I think Meta might be mixing up these two cases here on purpose to make it look like web scraping is as bad as stealing photos to publish it on a clone website.
- oefrha 4y ago> most commenters wouldn't want their posts shared privately amongst friends to be scraped and made available publicly. Where's the "posts shared privately amongst friends made public" part? There are two cases here: 1. A service that logs in as the customer (who voluntarily provide their credentials) and scrapes information visible to said customer on their behalf. Nothing about "made available publicly" is alleged. 2. An individual using a pool of bot accounts to scrape posts visible to any logged in user. Nothing about "shared privately" is alleged. To be clear I don't like the method, but I'll also have to admit I've used one of the Instagram "clone sites" in the past thanks to their login wall. Unless I missed something, it sounds like you just made it up.
- postalrat 4y agoWho is scraping their private messages? Themselves or their friends?
- Komodai 4y agoIs it Octopus Data Inc. aka Octoparse they are suing?
- allenleee 4y agoIronically, Octopus reminds me of "Octopus VR" in the Silicon Valley show. https://www.youtube.com/watch?v=ltFB4WBdDg4 https://www.youtube.com/watch?v=ltFB4WBdDg4
- mothsonasloth 4y ago"It's a water animal"
- carride 4y agoIn the early days of FB, they convinced people that pages (or some content, sorry I do not know the FB terms) could be public for anyone to view without needing to login to FB. This was very helpful for small businesses and communities. In many countries this is still the quickest place to make a public page. Though now, every small business or community page I want to visit is locked out unless I login FB. Even if I do login it is impossible to copy paste the important details of a page or post, plus the UI is as ugly as it has always been.
- carride 4y agoI am currently in the USA and when I visit a public FB page e.g. [1], there is a small login header, and a very big annoying footer login. I estimate 15% of the content is blocked. I had spent the past year outside USA until one month ago. When I visited the same sites while traveling outside the USA, the annoying login footer moves to the middle of the page blocking almost all content. I do not have proof at the moment, but that was my experience trying to read 95% of government, business, and community pages who are almost all on FB. [1] https://www.facebook.com/ParquesNacionalesdeArgentina
- i_have_an_idea 4y ago> After paying for access to the scraping software, customers self-compromised their Facebook and Instagram accounts by providing their authentication information to Octopus "self-compromised" lol clearly these people just wanted an automated way to access their own data
- antonf 4y ago> clearly these people just wanted an automated way to access their own data GDPR and CCPA (and probably many other national/state privacy laws) forces facebook/instagram/etc to let you download and/or delete your data without using third party websites. Usually people self-compromise their accounts in exchange for money: https://www.buzzfeednews.com/article/craigsilverman/facebook-account-rental-ad-laundering-scam https://www.buzzfeednews.com/article/craigsilverman/facebook...
- throwaway5959 4y agoWasn’t Meta stealing news articles and not paying news organizations for them?
- iandanforth 4y agoCollecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be able to provide their customers (advertisers) with reliable data about individuals. "protect people from scraping" Monopolists will protect themselves and call it protecting you. They will attempt to make you afraid of some other actor using your data in harmful ways so as to detract from how they monetize you and use your data in harmful ways. "deter the abuse" Monopolists don't want to argue about what constitutes abuse. Anything they write in their TOS is entirely for their benefit and only constrained by local law (if that). They will abuse you to the fullest extent they can get away with while arguing that any action to use your rights is "abuse." "safeguard people against clone sites" Monopolists want to maintain their monopoly, there is no greater threat than a direct challenge to that monopoly by allowing data to move freely. -- More subtle but even more ironic rhetorical points "for hire" / "paying for access" Emphasizing that people making money (gasp) for providing this service, is bad. "industry leader in taking legal action" + "across many platforms and national boundaries, also requires a collective effort from platforms, policymakers and civil society" Monopolists can pay high priced marketers to rebrand them as patriotic hero figures fighting valiantly for the little guy.
- blantonl 4y ago
- jjoonathan 4y ago[flagged]
- lcnPylGDnU4H9OF 4y agoIf simp is supposed to be short for simpleton, you might want to consider how simple your thoughts are.
- jmyeet 4y agoI'm torn on Web scraping because the extreme of each end of the spectrum on this issue both seem unreasonable. On one side, you have people who say any form of scraping is be disallowed, even prosecutable. This went so far that the Department of Justice on behalf of AT&T prosecuted a case of URL modification [1]. One of the few bright spots for this psychotic Supreme Court was to curtail the government's power under the CFAA by limiting what constituted "unauthorized" access [2]. On the other hand, there are those who think that any level of scraping should be fine and I think that's untenable too. Consider Yahoo indexing of Stack Overflow [3]: > In the meantime, since Yahoo (via Slurp!) is about 0.3% of our traffic, but insists on rudely consuming a huge chunk of our prime-time bandwidth, they’re getting IP banned and blocked. Do these "scraping extremists" think such actions should be illegal? It's actually not that far-fetched given the Ninth Circuit decided LinkedIn wrongly blocked HiQ scraping [4]. Like if you change your website with the intent that it'll make scraping more difficult, is that a problem? What if it's an unintended side effect? Additionally, companies like Meta, Google and Apple are going to be way more acountable to abiding by data retention laws and regulations than any scraper. If it's OK to scrape FB.com completely, that information is out there forever. I certainly think the government shouldn't prosecute on behalf of companies. At least that should expose to people how the government's #1 priority is in fact to protect the true constituents: corporations and the capital-owning class. [1]: https://www.techdirt.com/2013/09/30/dojs-insane-argument-against-weev-hes-felon-because-he-broke-rules-we-made-up/ https://www.techdirt.com/2013/09/30/dojs-insane-argument-aga... [2]: https://en.wikipedia.org/wiki/Van_Buren_v._United_States https://en.wikipedia.org/wiki/Van_Buren_v._United_States [3]: https://stackoverflow.blog/2009/06/16/the-perfect-web-spider-storm/ https://stackoverflow.blog/2009/06/16/the-perfect-web-spider... [4]: https://blog.ericgoldman.org/archives/2019/09/ninth-circuit-says-linkedin-wrongly-blocked-hiqs-scraping-efforts.htm https://blog.ericgoldman.org/archives/2019/09/ninth-circuit-...
- ConstantVigil 4y ago> So much about this case is ridiculous, and it’s complicated by the fact that nearly everyone agrees that weev is a world-class jerk. But, you need to separate that out from the details of what he did here, to note that it was nothing particularly special, and it involved the sort of thing that security researchers do all the time, and which all sorts of non-security researchers do quite often. Yeah... uhm... I used to do exactly this sort of thing... When I was a teenager, I would look at the URL of whatever site I was on, and would change a number here, or a letter there; and see what I got. Sometimes you get nothing, sometimes you get something. Sometimes that something is quite interesting.
- ConstantVigil 4y ago
- viburnum 4y agoOne of Facebook’s earliest acquisitions was a scraping company called Octazen.
- throw20220707 4y agoFrom GDPR point-of-view this kind of 3rd party data collection is not acceptable (assuming it covers personal information, for example names of people and what they have posted). The difference with Meta's own data collection is that the users have relationship with Meta and users have given their permission for Meta to handle the data. Users also know they can contact Meta and ask them to remove the data. 3rd parties don't have the consent from users. Users don't even have an idea these companies might be holding their data.
- Nextgrid 4y agoFrom a GDPR point of view the scraper would be acting as a data processor on behalf of their customer, no different from using a cloud storage service for your contacts. It's fine as long as the third-party doesn't misuse the scraped data or share it with third-parties and there's no evidence they did so in this case.
- danuker 4y ago> and there's no evidence they did so in this case. Indeed; the users probably wanted to make the data public, if scraper accounts could see it. There is a GDPR allowance for data "manifestly made public by the data subject". https://gdpr-info.eu/art-9-gdpr/ https://gdpr-info.eu/art-9-gdpr/ Here, it's just Facebook wanting to keep the data inside a walled garden. For the same reason, I quit LinkedIn and made my own site. I don't want people to have to sign in to see my profile.
- oxff 4y agoPretty rich idea coming from FB, lol. They do human scraping.
- samsoftstuff 4y agoIt's like they don't know that courts just made it legal: https://techcrunch.com/2022/04/18/web-scraping-legal-court/ https://techcrunch.com/2022/04/18/web-scraping-legal-court/
- blantonl 4y ago"Legal" doesn't make it ethical, nor does it shield you from liability if you willfully violate contract law (terms of service)
- brushfoot 4y agoFrom the article: "[T]he Ninth Circuit reaffirmed its original decision and found that scraping data that is publicly accessible on the internet is not a violation of the Computer Fraud and Abuse Act." The key phrase is "publicly accessible." This wasn't that. The scraping was done by automating Facebook accounts, which have terms of service, which forbid scraping. ToS/EULAs make a big difference. They're the reason Blizzard could shut down bnetd's StarCraft server. They're why no one can legally reverse engineer Oracle to create a drop-in replacement, despite interoperability provisions. More and more platforms are putting the majority of your user-generated content behind auth walls with ToS because that's how they prevent competitors from swiping it.
- Nextgrid 4y agoDoes it go into detail about the actual meaning of "publicly accessible"? Because most content on Facebook/Instagram requires any valid login (as opposed to a specific account) and that data people intend to be public (especially on Insta). In this case, the account requirement would be a technicality and the data, for all intents and purposes, would still be considered "publicly accessible" if anyone with an account can access it.
- upupandup 4y agoPutting a login screen that any public member can bypass isn't private information. Private info would be Onlyfans videos. So far there is no such feature on Instagram
- 4y ago
- samsoftstuff 4y agoIt's like they don't know that courts made it legal: https://techcrunch.com/2022/04/18/web-scraping-legal-court/ https://techcrunch.com/2022/04/18/web-scraping-legal-court/
- almog 4y agoIronically, around a year ago I disclosed (using their White Hat bug bounty program) that I'm able to access recruitment data (candidates details mostly) using very cheap form of scraping against a 3rd party service provider, they dismissed it and instructed me to report it to the 3rd party that operates that service (which I did beforehand but the issue has had not been fixed). Sorry for being vague here, I haven't publicly disclosed it yet, but will probably have to if it don't get fixed.
- jascii 4y agoSo, Facebook doesn't want to share the data it wants us to share with them? Figures...
- deleted 4y ago[deleted]
- neya 4y agoEvil Big Co. that literally STEALS people's personal information everywhere they go even after they've indicated they want to be left alone is now offended when someone does the same to them? Well, color me surprised /s Fuck Facebook. Meta. Or whatever you want to call it.
- romanovcode 4y ago> Meta is an industry leader in taking legal action to protect people from scraping and exposing these types of services, which provide scraping as a service across multiple websites. Sure, as long as Meta is not the one selling the data to Cambridge Analytica it's wrong.
- typon 4y agoGoogle has turned Google Search into a walled garden by scraping people's content and serving it up on their own platter. Is anyone going to stand up to them?
- Litost 4y agoAnyone else heard of Tim Berners-Lee's idea of hosting your data in pods outside the relevant corps wanting access to it and you controlling what's shared and how? This is such a completely different way of doing it, I'm not sure of all the implications, be that from admin (how much effort) to security (would this be a massive hacking opportunity) etc. https://www.theregister.com/2022/01/20/tim_bernerslee/ https://www.theregister.com/2022/01/20/tim_bernerslee/
- nicholasjarnold 4y agoFunny story from the early days of TheFaceBook, probably around 2005ish: I was a webmaster of a set of servers on a major university's network. I also had access (enough to run arbitrary programs that had pretty much full ingress/egress to the public internet) to a number of machines across the campus's network. Through some of my coursework and ACM chapter activities I met some other similarly minded technical people with similar levels of access. We decide that it would be fun to use our superpowers (access + programming abilities + curiosity) to sign up for various accounts on FB and essentially scrape and friend as much as possible. At the time they had some rate limiting, some IP banning (which wasn't terrible because the Uni gave public IPv4 addrs to all machines on campus by default) and then added some early CAPTCHA which we ended up breaking pretty trivially with some python and image recognition code. Never got sued... :) Never really did much with the scripts or data except test that they worked. Fun times.
- ok123456 4y agoRemember back when facebook grew their little network by scraping your gmail contacts. Google blocked them. There was animus between the two companies that resulted in Facebook not making an official android app until 2010.
- upupandup 4y agowhoa wasn't there somebody on HN that ran a web scraping shop that were boasting they can scrape instagram a while back? are these the same guys??? I don't know how far Facebook can get with this, thought Linkedin's court ruling made scraping legal de-facto
- throwaway_meta 4y agoPeople that are criticizing this probably were also critical of the Cambridge Analytica scandal, but it would be useful to compare what happened there and here. With Cambridge Analytica: - Facebook allowed users (with informed consent) to allow external developers to access their data and limited data about their friends, in order to build social-enabled apps. - CA exploited this to scrape basic profile data from a large number of users. It broke the ToS by doing so (in particular by using the data for purposes different than stated) Here the same is happening: - people are giving a third company access to their profile, which includes access to friends' data (in fact a lot more than what the app platform allowed to do) - the company is scraping all the data. At the time of CA, the criticism was that Facebook didn't do enough to enforce its ToS (or maybe that the data sharing should have not been allowed in the first place? But the terms were common knowledge and the attack potential became clear only in hindsight), here people are criticizing that Facebook is in fact enforcing its ToS. Also note that strong enforcement against scraping is one of the mandates that came from the FTC settlement. It seems inevitable that any news about Facebook/Meta is read in the worst possible light these days, even when the criticism is self-contradictory. I would expect less superficial commentary from HN.
- unosama 4y agoThe real reason most people were upset about Cambridge Analytica was it revealed to the public how advertising and PR companies manipulate us. The fact they violated facebook ToS is moreso the excuse for the press covering it when they wanted to write another anti-Trump piece. If you were accusing a specific newspaper of hypocrisy based on two article I might agree. But you're referring to general public sentiment, and I really don't think most people cared or were surprised about the data collection. The shock and scandal was the realization that targeted advertising campaigns and information bubbles have the potential to sway elections.
- throwaway_meta 4y agoI'm referring to the HN crowd, I'm not sure that can be equated to "general public sentiment". I agree with your first paragraph, and my point is that it is not possible to argue at the same time that Facebook should share data more broadly and allow scraping, and at the same time be critical that Facebook allowed CA to happen in the first place. If the CA scandal was a wake-up call, it appears it was not internalized enough for people to understand the implications of what they're suggesting in this thread?
- Nextgrid 4y agoSo much bad faith in this press release but not surprising from such a disgusting company, with of course some China-related fear-mongering despite no evidence of wrongdoing. > After paying for access to the scraping software, customers self-compromised their Facebook and Instagram accounts by providing their authentication information to Octopus. They didn't "self-compromise" their account. They trust Octopus to act on their behalf, and unlike Facebook, Octopus' interests are most likely more aligned with their users' since their service is paid. This is no different from handing your Facebook credentials to your social media manager or secretary. There's no evidence that Octopus misused this access in any way. > Octopus designed the software to scrape data accessible to the user when logged into their accounts, including data about their Facebook Friends such as email address, phone number, gender and date of birth, as well as Instagram followers and engagement information such as name, user profile URL, location and number of likes and comments per post. This is either information people intend to be public or information they trust their friends to keep private. Now if Octopus was leaking the private information to third-parties it would be one thing, but so far I see no evidence Octopus was disclosing the scraped information to anyone but their customer (who is already authorized to access it). > Meta is an industry leader in taking legal action to protect people from scraping and exposing these types of services Translation: Meta is an industry leader in protecting its disgusting business model that hinges on making public data behind a walled garden with an unacceptable "privacy" policy. There wouldn't be a market for Octopus (or other scrapers) if Facebook already allowed customers to efficiently access information they're already entitled to, but that would be against their interests as their entire business hinges on information being held hostage. They've created a problem, are selling the cure (well in this case monetizing it via ads) and are now pissed off that someone else is selling the cure for cheaper.
- postalrat 4y agoHey instagram/facebook/linkedin/etc: It's not your data.
- dmje 4y agoOr Facebook could just open up their data. Oh wait, not their data, silly me. Everyone else's data. Keep on scraping, I say.
- uhtred 4y agoFuck off Facebook you scumbags
- rmbyrro 4y agoThe fact they're wasting time on that is a sign that Facebook decay phase has already started.