12 ms·
[Throwaway for privacy.] I know this was hashed out on the other threads a bit, but can someone please explain to me why folks are so up in arms about this, co
by throwawyaaccoun 5y ago
[Throwaway for privacy.]
I know this was hashed out on the other threads a bit, but can someone please explain to me why folks are so up in arms about this, compared to, say, studies that scrape user data without consent (something the IRB allows all the time by saying that no human subjects are involved)? Is it simply because there is no visibility into this practice (i.e., no email sent?) Scraping user data from public profiles, aggregating it into a model, and publishing a paper or whatever -- that seems demonstrably more invasive to individuals, storing and keeping their user data, than an email quoting a statute.
I agree that the deception was unnecessary, but that's it. It doesn't feel any wronger than that.
Especially because these researchers really were acting in "meta" good faith trying to probe the privacy ecosystem, I fear there may be a chilling effect. Consumers deserve privacy rights and privacy knowledge in the asymmetric surveillance economy we find ourselves in, IMO.
I'm open to being wrong.
- halpert 5y agoYour question is essentially whataboutism. Both things can be wrong. We can care about this instance without diluting the conversation talking about something else that is also bad.
- throwawyaaccoun 5y agoIt's not intended to be whataboutism (sorry about that, I edited this in to clarify) -- I agree that the deception was wrong. But there seems to be something about this particular event that is riling people up, and that's what I am getting at. I am not trying to whatabout, to be super clear.
- runnerup 5y agoTo clarify. I don't think people would be riled up about individuals sending out these emails. Individuals are required to be legal, not 'ethical'. The people who are riled up believe that University studies should be performed ethically. They know that IRB's exist to prevent researchers from doing unethical, but legal, things. In this case, they feel the harm caused should have been prevented. Scraping data silently doesn't cause stress/harm to the participants directly, as they are unaware of any potential threat. It's not "human experimentation should be banned" its "human experimentation should be heavily scrutinized to prevent harm to participants as much as possible. And definitely never cause harm to unwilling / unwitting participants".
- dcow 5y agoWhat bothers/riles me is that there doesn't seem to be a consistent ethical framework applying to these complex situations. Of course things should be ethical but ethics aren't defined as “whatever people on HN and Twitter feel like isn't slimy”.
- tylermenezes 5y agoBecause the end of the email (wrongly in most cases) demanded a response by law and implied they were open to legal action, which caused a bunch of people to hire lawyers to check into their liability.
- dahfizz 5y agoMaybe the problem is the laws which create unknown liability for anyone hosting websites.
- eli 5y agoIn this case the law wasn't the issue. The email message asserted a legal obligation that does not exist.
- dahfizz 5y ago>The controller shall provide information on action taken on a request under Articles 15 to 22 to the data subject without undue delay and in any event within one month of receipt of the request[1] The legal obligation may not have applied in this case, but it absolutely exists. If someone submits a request to you for their data, you are legally obligated to respond. [1] https://gdpr-info.eu/art-12-gdpr/ https://gdpr-info.eu/art-12-gdpr/
- kstrauser 5y agoThe request I got was about the CCPA. It said: > I look forward to your reply without undue delay and at most within 45 days of this email, as required by Section 1798.130 of the California Civil Code. First, the CCPA doesn't apply to my site. It's non-commercial, has many fewer users than required to invoke the CCPA, and zero revenue. No provisions of the CCPA require me to do anything. Second, the questions were about how I'd handle a CCPA request, and weren't actually a request at all: > 1. Would you process a CCPA data access request from me even though I am not a resident of California? > 2. Do you process CCPA data access requests via email, a website, or telephone? If via a website, what is the URL I should go to? > 3. What personal information do I have to submit for you to verify and process a CCPA data access request? > 4. What information do you provide in response to a CCPA data access request? The CCPA doesn't obligate anyone to explain their internal processes. It obligates covered entities to respond to the requests themselves, but not to random drive-by questions. So basically, that sentence was completely wrong. The CCPA doesn't apply to me, and even if it did, the law doesn't say what the researchers claim it did.
- _jal 5y agoThe issue here was not primarily about deception. It seems mainly to be that (a) at least one recipient interpreted their mail as a legal threat, and (b) it was a mass-mailing. Spend a minute thinking through the implications if that were true, and you get a firestorm. I suspect visibility plays a role in the comparison you're making; out of sight, out of mind and all that. But much more importantly, someone sending you what you think is a legal threat is a lot more salient.
- throwawyaaccoun 5y agoInteresting. Ok, so let's say the deception wasn't the problem, suppose for the moment. Would the study have been more palatable if the researchers had more properly vetted the email list to ensure, say, >95% or perhaps even 100% were corporations that did fall under the law?
- _jal 5y agoUnless you're studying how people react to online legal threats, why would you not try to avoid this problem with your study entirely?
- s1artibartfast 5y agoDeception is a necessary part but not the key. The key is potential for distressing a real human being. The problem is that we live in a legal Society where everyone is at risk of life-altering legal consequences.
- throwawyaaccoun 5y agoOh, our society, especially America's, is overly litigious. I agree. But, pushing back a bit (in good faith), do you think asking an entity for your data, or asking them to delete it, should really be considered unusual and panic provoking? I said in another comment the same thing, but do you think this could be a moment of cultural learning?
- 5y ago
- kstrauser 5y agoPart of it was that the did no (or poor) screening. They got their list of target sites from a research list of the popular websites. I got a letter, and my little not-for-profit, not advertised, purely for fun website was around number 350,000 on that list. First, I sincerely doubt my site is even that popular. Second, if I got the mail, so did lots of people in a similar situation. They weren’t spamming Fortune 500 companies. They were spamming a huge number of single-person sites that aren’t subject to the CCPA at all and who certainly don’t have legal departments to ask about it.
- throwawyaaccoun 5y agoI mean this all in good faith: What is the difference between 100,000 individuals emailing 3-5 websites on that list, with their real identities, asking for things to be deleted (such that all 350k are covered)? Where is the meaningful difference between this situation and the one here, ignoring the deception for a moment (unless that is the only issue)? Could this be a moment of cultural learning for everyone? That's kind of how I am looking at it, frankly, but I am open to being wrong. That is, perhaps small entities will learn, in one or two instances, to just ignore this kind of thing?
- rectang 5y agoYou seem extremely unconvinced that any harm was done to the people who were sent scrambling by this alarm. It's as though no matter how convincing the email was, no matter how much of the recipient's time was wasted, no matter how many thousands of dollars they spent on lawyers, you ascribe all blame to the recipient for not having realized they were being deceived — and ascribe no blame whatsoever to the email's author for being deceitful. This whole discussion was had in the old thread, and there was one person who used the same rhetorical device of belaboring the same question over and over again. It was tiresome.
- throwawyaaccoun 5y agoI should have been more clear, so let me correct that. I am convinced. I agree that harm was done, and suffer from generalized anxiety disorder myself, so I empathize with the panic attacks that people received. It is because I believe that harm was done, but also because I am a privacy nut myself, that I am trying to, for my own sake, characterize how I should approach sending emails like this in the future. The study may not go on, but individuals still will send these emails as long as CCPA/GDPR exist. (Just to add some color: It's my anxiety which is causing my to want to delete everything from the internet. If there's minimal info about me online, I can rest easy. It's why this is a throwaway that I will abandon shortly.) Reading everyone's thoughts is what changed my mind. I now understand to have underestimated the emotional and legal effects CCPA/GDPR requests could have on small website operators, and will be more judicious in the future (like this study should have been) in pre-filtering and my wording. Reactions like kstrauser's (elsewhere in thread) were initially surprising to me (perhaps because of the faceless nature of the internet), so I hope you take my about face as genuine. Where do you think this balance lies? I still believe consumers, in general, should have right to ask those with their data about their processes; to give it to them; and, to upon request, delete it. And further, in general, I think these interactions are the kinds of things that researchers might legitimately want to study. I found your other comments to be thoughtful, so I am curious what you think explicitly.
- eli 5y agoYou shouldn't lie to people to trick them into collecting data for you without at least considering the impact on those people. That's nothing like web scraping. (Though IMHO web scrapers should also use an honest User Agent so if website owners have a problem or question or want to block it, they can)
- codazoda 5y agoThe post here on hacker news mentioned the down sides for one receiver. That person was stressed out thinking that they were about to be sued. They considered retaining council, which could have cost them a few thousand dollars, in order to get ahead of the threat. It didn’t come to that, so it’s a “what if”, but I could see myself trying to retain council too. Hopefully, a lawyer would have talked me down and advised me to wait it out. On the flip side, they may have offered to respond on my behalf (which would cost money). I would not respond to such an email myself, ignoring it until I was able to defer to an attorney. I publish a simple personal blog and I worry about the _worldwide_ legal implications of doing so. As one example, I have some old information about making model rocket fuel at home. At the time I had carefully reviewed U.S. law and knew how much I could legally make and have in my possession. Then I got questions from people in other countries and I got spooked. What if I break a law somewhere else?
- kstrauser 5y agoI assume that I’m breaking other countries’ laws all the time, say be criticizing the actions of their governments. I don’t worry about that. I’m much more worried about, say, CCPA compliance while living and working in California. (Not that I’m especially worried it. My personal projects don’t meet any of the criteria which would make it apply to me.)
- codazoda 5y agoYeah, me too. I don’t collect stats on visitors anymore (using Google Analytics for example) because I now understand the privacy implications of doing so. I do use a simple impression counter but I capture no information (not IP, not browser, nothing). I definitely think about the CCPA and ADA laws, but I’m relatively sure they don’t apply to me. Still, I certainly think about them.
- dhimes 5y agoI go as far as saying in the TOS that my sites are for users in the US.
- mint2 5y agoScraping dating is not imposing work, worry and cost on additional people. The victims of scraping are not going to do any additional work unless the scraped data is used irresponsibly, but that is separate from the act of scraping. This email required people to do work and caused worry due to the legal threat that the email tried to lead people to believe was applicable to them. They may have had cost if they called a lawyer and it definitely took their time. Scraping -> no work forced upon victims. That email -> work forced on unwilling victims. Is there something I’m missing? People including that poster aren’t reaching this same conclusion but it seems very apparent so am I missing something?
- nightpool 5y agoWell, the argument of the GP is that "extra work" is not the only form of harm that is possible. When comparing the harm of extra work and stress due to this email to the harm of have your privacy violated by large, publicly-scraped datasets that include your personal information. For example, once your twitter post is collected in a "posts of Twitter users about X political event" dataset, it's now impossible for you to ever delete that post, which could be harmful for you in the future. it's unclear whether one type of harm is categorically worse then the other.
- mint2 5y agopublic posts on the internet being aggregated is not out of the ordinary, if one group doesn’t do it, another may. Scraping private posts would be wrong or gaining access to posts under false pretenses. This would be wrong, although different than the email. The email forced work on people and made legal threats causing work and other effects that would not otherwise happen.
- otrahuevada 5y agoThe wording on the main driver of the experiment, their especially bad emails, leads website operators to think there is a problem where there is none. This, on top of the research being entirely devoid of consent between the human parties involved, makes it a _very_ bad study, one that could well cause both the university and the research team to lose money if some of the 'subject' parties actually had to go get a lawyer to have a look at their shoddy emails. In better studies what is supposed to happen is, you propose taking part in the experiment, you get a signed agreement of some sort, and only then actually start experimenting. What happened here is more like some kind of youtube prank than a useful information gathering procedure.
- sennight 5y agoScraping public data doesn't result in compelling another person to work under a false premise. Sure, you could argue that scraping introduces load that may draw an operator's attention... but the comparison is a pretty big stretch. How these things pass board review I don't know... it seems pretty obvious to me that creating work for somebody who didn't volunteer to it is, at best, antisocial behavior.
- adolph 5y ago"meta" good faith != good faith > why folks are so up in arms about this The implicit legal threat is similar to the harm described in the Prenda saga: https://arstechnica.com/tag/prenda-law/ https://arstechnica.com/tag/prenda-law/ It is wronger than the deception because the PI "Jonathan Mayer" is not just a run of the mill academic focused on "publishing a paper or whatever." This is an activist with an ax that won't grind itself. Reviewing his work mentioned in Wikipedia I'm impressed and appreciate the contributions Mayer has made. Mayer can't be not aware of the problems with the approach.
- UncleMeat 5y agoI personally know Jonathan and hugely respect his work. I could believe that because he is an actual lawyer it was harder to imagine the panic that recipients who have no understanding of the law would experience. But I think that more likely is that the response was a bit of a fluke. Way stranger stuff has been done by security and privacy researchers with the go-ahead from their IRB. This feels to me like this is a methodology that isn't universally agreed on but is not especially uncommon that tripped a response from the internet. The conclusion is more that people should not necessarily take the existence similar research as indication that the broader community is okay with these methodologies.
- adolph 5y agoI suspect Meyer's work is in part preparatory to lawfare in order to force websites to pay for lawyerly services. The letter is akin to a fire insurance company knocking on doors while carrying a torch. "Of all tyrannies a tyranny sincerely exercised for the good of its victims may be the most oppressive." https://quoteinvestigator.com/2019/12/19/intentions/ https://quoteinvestigator.com/2019/12/19/intentions/
- UncleMeat 5y agoFrankly, that's stupid. He's got a PhD and a JD from Stanford and has chosen a faculty position and has done a nontrivial amount of unpaid work for various privacy rights organizations. He obviously isn't motivated by money.
- rubylark 5y agoDemanding a subject to actively participate in your study upon pain of vague and mostly incorrect legal threat is ethically wrong. Passive participation (like scraping) without consent is morally wrong, but since it doesn't cause undue distress to the subjects, it is not as big of a story. The IRB in this case didn't consider this ethically suspect because "websites aren't people". And yet the study disproportionately targeted small websites where there is, in many cases, only one person involved.
- btown 5y agoIn US legal code there is actually a definition of a human subject in https://www.hhs.gov/ohrp/regulations-and-policy/regulations/45-cfr-46/revised-common-rule-regulatory-text/index.html#46.102 https://www.hhs.gov/ohrp/regulations-and-policy/regulations/... (EDIT: to clarify this is a guideline for federal researchers and to my knowledge is not legally binding on private institutions, but seems to be used as a basis for private IRB policies): """ (e)(1) Human subject means a living individual about whom an investigator (whether professional or student) conducting research: (i) Obtains information or biospecimens through intervention or interaction with the individual, and uses, studies, or analyzes the information or biospecimens; or (ii) Obtains, uses, studies, analyzes, or generates identifiable private information or identifiable biospecimens. (2) Intervention includes both physical procedures by which information or biospecimens are gathered (e.g., venipuncture) and manipulations of the subject or the subject’s environment that are performed for research purposes. (3) Interaction includes communication or interpersonal contact between investigator and subject. (4) Private information includes information about behavior that occurs in a context in which an individual can reasonably expect that no observation or recording is taking place, and information that has been provided for specific purposes by an individual and that the individual can reasonably expect will not be made public (e.g., a medical record). """ The argument is that scraping of public data, already recorded by data systems for general (e.g. not specifically medical) purposes, is neither intervention, interaction, nor private information. On the other hand, IMO the researchers here clearly interacted with their subjects. While the email was sent to a privacy@ address, not only are emails different from HTTP GET in how likely they are to be read by humans, but this went a step further and implied legal action would be forthcoming unless a human replied to the message. That's interaction. That makes the recipient a human subject. (IANAL and the above is not legal advice.) EDIT 2: I've had the pleasure to meet one of the researchers here. They are a staunch defender of online privacy, and I believe the team sincerely wanted to measure how effectively businesses are adapting to the changing winds beyond their legal obligations. But I also think the team, and the Princeton and Radcliffe IRBs, should have done more to consider the impact on the people who operate these businesses themselves. I'm sad and disappointed that the systems in place didn't catch this.
- mindslight 5y agoI believe the real issue isn't the research ethics per se, but rather pent up frustration on the larger topic. I posted this in one of the original threads: https://news.ycombinator.com/item?id=29607123 https://news.ycombinator.com/item?id=29607123
- deleted 5y ago[deleted]
- ineedasername 5y agoEthical guidelines on research exist to prevent an adverse impact on participants. This study had adverse impacts: fear, stress, time & money in consulting lawyers. It was therefore defacto an unethical research study. Speculation as to why the protocol slipped through the IRB cracks are that the language used in the study proposal (at least the part made public) dehumanized the protocols by referring to "websites" rather than humans that would be responding to the inquiries. The IRB ruled this was not a human subject piece of research, but that is contradicted by the deception protocol. Deception was justified as necessary because people's behavior might change if they knew it was a research request. That acknowledgement made it implicit that human behavior and potential changes to it due to the experiment was a core factor in the study-- ergo, it had human subjects. Behavioral research on human subjects is required to go through a much more rigorous IRB oversight process precisely to anticipate and mitigate potential adverse reactions. Some people are focussing on the deception, but that is, under some circumstances, allowed by research ethics. The more serious problem was adverse impact which, again, is the primary motivator for why we now have laws and regulation-mandated IRB processes to make sure it doesn't become an issue.
- belorn 5y agoI wonder if IRB ruled in this way because of the assumption of algorithmic response for requests like DMCA take down notices. I can imagine that even for GDPR/CCPA requests, there is still no human involved for website like Google, facebook, youtube and other major sites that is primarily operated through automation. If there is no humans involved then there is no humans to have an adverse impact on. But as you said, researchers however must have suspected that responses would be made by humans or else the email would have included the fact that it was a study.
- throw10920 5y ago> Ethical guidelines on research exist to prevent an adverse impact on participants. Not relevant in this case, because (1) it's clearly not human research[1] and (2) all kinds of other clearly-non-human-experiments still have adverse effects on humans tangentially involved, so that's not unique to human research, either. When scientists were testing the first particle accelerators, they caused some people a lot of stress who were worried that they would destroy the world - does that mean that those tests were human experiments? (clearly not) [1] https://news.ycombinator.com/item?id=29656792 https://news.ycombinator.com/item?id=29656792
- kortilla 5y agoBecause it comes across as a vague legal threat to a website operator! That’s in no way like scraping databases. This cost real legal resources (there are Twitter threads of internal legal counsel hiring outside firms to evaluate this).
- dataflow 5y agoI don't see the similarity. Scraping doesn't involve a thinly veiled threat of a potentially costly lawsuit.