27 ms·
Microsoft, OpenAI sued for ChatGPT 'privacy violations'
- hospitalJail 3y agoWhen this happened to Stable Diffusion, it was easy for me to consider it a necessary evil to progress humanity. When this happens to closedAI, it just seems like a profit grab. Not that it changes the legality of it. Just optics. Wonder if that matters in court.
- cheschire 3y agoFirst they came for the graphics artists, but I did not speak out because I was not a graphics artist. Then they came for the writers, but I did not speak out because I was not a writer. Then they came for me, and there was no one left to speak for me... well, except ChatGPT.
- 111111IIIIIII 3y agoI did not speak out because copyright is farcical nonsense that fetishizes the profit motive at the expense of humanity.
- pessimizer 3y agoThat's why copyright violation should be brutally cracked down on when the copyrights of Microsoft are violated, and lawsuits against Microsoft for intentional and widespread copyright violation should be laughed off. Because capitalism is bad. edit: corporate LLMs have pulled the "one death is a tragedy, ten thousand deaths are a statistic" ploy off fully. If you want people to question whether you're even violating copyright, make sure you violate all of them at the same time. They'll just decide that you're an act of god and not covered under earthly laws.
- 111111IIIIIII 3y agoPrimacy of capital is literally the culprit of the inequality you're complaining about, and the reason you cannot win short of reorganizing society.
- hospitalJail 3y agoIt seems people prefer power distributed by capital, rather than military might or factionalism/leaders/politics. Not that all capital is distributed by merit, plenty of people used military might or factionalism/leaders/politics to obtain disproportionate amount of capital. But if you are against the last 2 happening, I don't see what you expect a reorganization of society to accomplish since you are going to get a power structure of factionalism/leaders/politics taking priority. (Sorry bud, no an-com utopia ever existed, they all had factionalism/leaders/politics, thus defeating the entire purpose of removing class.) I think most of us think we can capture/retain power easier with money, than having to climb up inter-party politics.
- 111111IIIIIII 3y agoCapital primacy is maintained by the capitalist state, i.e. the monopoly on violence. This is literally military might. I don't necessarily disagree with your later points. I do, however, disagree with giving up.
- hospitalJail 3y ago>Capital primacy is maintained by the capitalist state, i.e. the monopoly on violence. This is literally military might. At least its equitable (based on value of output), ofc there are legacy issues as well. Some demagogue can swoon the masses and take it all if not for capital. That demagogue could be Trump or Stalin. Know the consequences of what you are advocating for.
- 111111IIIIIII 3y ago> At least its equitable (based on value of output), ofc there are legacy issues as well. It's not. By definition, it's based on control of capital. That's why it's called capitalism. In other words, those aren't "legacy" issues; they are literally the system as designed.
- gl-prod 3y agoAs a language model I cannot speak for you. But I can help you express your thoughts and views. I can generate words and sentences in many ways.
- jwx48 3y agoI am not a lawyer, just a Sysadmin; but with that said, the linked pdf of the complaint is absolutely fascinating to me. It's worth it (to me) for the list of resources it cites.
- numbsafari 3y agoIt’s okay, I’m sure everything is going to be fine when Microsoft and ChatGPT hot mic your next doctor appointment. https://news.ycombinator.com/item?id=36498294 https://news.ycombinator.com/item?id=36498294
- thumbuddy 3y agoAssuming Google, Amazon, etc haven't already been doing exactly that.
- bannedbybros 3y ago[dead]
- dmix 3y agoThat says it's using GPT4 but it's not clear if it has anything to do with feeding back into ChatGPT. > Nuance has strict data agreements with its customers, so patient data is fully encrypted and runs in HIPAA-compliant environments Additionally Epic seems to already be storing these clinical notes in databases and Nuance which Microsoft owns has already technically been a 'hot mic' in these same doctors office for some time. The new offering is an AI-draft note generator. I'm personally skeptical that model output would suddenly be under different rules than the other voice-to-text AI model output?
- ldehaan 3y ago[dead]
- replwoacause 3y agoA tangential question...but does anyone know what software is used to generate legal documents that look like the PDF linked in the article? I’ve played with LaTeX templates a bit, but I seriously doubt law firms are futzing around with LaTeX for documents as complex as this. They must have some software that produces this formatting.
- deleted 3y ago[deleted]
- frakt0x90 3y agoIn my sample size of one, an attorney I talked to said that Microsoft Word was the most important software he and his colleagues used. So my guess is they're just really good with Word.
- bushbaba 3y agoYep that pdf can be made using ms word
- replwoacause 3y agoThanks! That surprises me but maybe it shouldn’t. I figured it was some purpose-built software for attorneys.
- ghaff 3y agoWord has pretty good revision tracking and support for footnotes which are probably the main things lawyers use more than most average people do. And remember that lawyers communicate a lot with clients, etc. too so there would be a lot of friction associated with a non-standard tool. When I worked on an expert witness report for a big law firm we just used Word.
- svat 3y ago`pdfinfo` on the file says: Creator: Acrobat PDFMaker 23 for Word Producer: Adobe PDF Library 23.3.247; modified using iText® 7.1.6 ©2000-2019 iText Group NV (Administrative Office of the United States Courts; licensed version) So it was likely made in Word and exported to PDF. (One can anyway guess from the "look" of the paragraphs that they're not using anything like Knuth–Plass line-breaking, which rules out things like *TeX and InDesign.)
- zug_zug 3y agoHard to understand how this is a crime, or how they came up with 3 billion dollars of damage. Seems like if it's legal for a person to do it should be legal for software to do for the most part.
- lionkor 3y agoLike operating motor vehicles, carrying guns in some US states, sueing people and companies, submitting content to wikipedia, writing children's books, and writing and voting on laws? Surely, there is some pretty large subset of things where "if it's legal for a person to do it should be legal for software" does not hold up? So how about the default is "not allowed"
- data-ottawa 3y agoI can personally memorize and recite copyrighted works all I want, but when ChatGPT does it then it’s in a commercial context and they’re liable to be sued for infringement. If you ask ChatGPT the rules for D&D, the private sourcebooks are all in there.
- hughesjj 3y ago> and recite copyrighted works all I want ...wait, isn't that false? legitimately asking. or is it because it was done by a corporation that makes it illegal? im thinking of how restaurants dont sing happy birthday and fair use restrictions etc
- bena 3y agoLike most things, it depends. If I recite them to myself, in my home, it's fine. If I do it at a gathering at my house where we're playing D&D, fine. If I do it as a performance, in front of a crowd, or as a recording, now I'm no longer fine. Context matters in a copyright cases. Not to mention, to claim fair use, you do have to claim you violated copyright. Fair use is just an allowed violation. As to Happy Birthday, that's actually ok for them to do now. The person/group that held the copyright to Happy Birthday was found to have not actually have held them in the first place. Happy Birthday is actually an older song called "Good Morning to All". Swap "Good Morning" with "Happy Birthday" and "children" with "dear [PERSON]" and you have the lyrics. This was not deemed a substantive change. And since the copyright on "Good Morning to All" has lapsed, Happy Birthday is in the public domain.
- WA 3y agoInteresting, for once this doesn't have anything to do with the GDPR. It's by 16 (US) individuals, filing the complaint in SF.
- ttul 3y agoIt had to happen eventually. There is so much money to go after. This is a case of lawyers creating their own income stream.
- knaik94 3y agoCalifornia is the only state that has active data privacy laws. Although, I don't think there's any financial transactions, it's just public data scraping. I wonder if the company can even be held liable for the output of these LLMs. There's no direct hosting of any static data. https://iapp.org/resources/article/us-state-privacy-legislation-tracker/ https://iapp.org/resources/article/us-state-privacy-legislat... https://leginfo.legislature.ca.gov/faces/codes_displayText.xhtml?division=3.&part=4.&lawCode=CIV&title=1.81.5 https://leginfo.legislature.ca.gov/faces/codes_displayText.x...
- submeta 3y agoWell, it was too good to be true. Reminds me of the early days of music sharing and Napster.
- bannedbybros 3y agoNapster was people sharing files. That's a crime! But now it's corporations so it should be legal.
- elforce002 3y agoWell, I think the main catalyst here is that corporations need to pay creators, etc... Napster cut the middleman, hehe.
- Cthulhu_ 3y agoWhich was never legal in the first place, but it was great because it liberated music and content to the masses. It was the necessary precursor to what is now Spotify and the like, instant access to billions of songs. The music industry didn't like that (Napster & co) because they wanted purchases and to get paid every time the music they owned was played.
- lionkor 3y agoDiscord rolled out a ChatGPT based bot that can be used in (and thus can read) all private conversations. Not surprised there are issues with it.
- barathr 3y agoRather than there being lawsuit after lawsuit of this sort, we wrote an op-ed this morning that says there should be a simple, compulsory licensing fee that AI companies pay to the public -- something we called the AI Dividend: https://www.politico.com/news/magazine/2023/06/29/ai-pay-americans-data-00103648 https://www.politico.com/news/magazine/2023/06/29/ai-pay-ame...
- jbarrow 3y agoThe order of magnitude of suggested pricing is really interesting: $0.001/word is significantly more expensive than, say, OpenAI's pricing of GPT-3.5-turbo ($0.002/1k tokens, ~750 words, so ~$0.000003/word, assuming I got my zeros correct). So this would increase the cost of running GPT-3 by about 300x. In terms of implementation, I wonder about a few things: Do models trained on more data have to pay more? LLaMA was trained on 1.5T tokens, the original GPT-3 was trained on ~300B tokens. And this is only partially related to model quality, LLaMA 13B and LLaMA 65B were trained on the same data, but the 65B model is better. What's the incentive to ever use the 13B model, if the licensing cost is 100x-1000x the model inference cost? Who defines a word? Each model uses a different tokenizer. I'm personally amused by the idea of a government-mandated tokenizer. What about generations that never see human eyes? As an NLP researcher, I've generated millions of tokens for training and automatic evaluation purposes -- are those subject to licensing as well?
- barathr 3y agoYeah, the idea is that it's much more expensive than current OpenAI pricing but much less expensive than what even a low-end marketing copy writer would charge per word. Its side effect would be to push such tools towards more valuable uses. The idea is to keep it simple, so it wouldn't be based upon the specifics of training, just whether or not it used public data. Anything else would require companies to divulge trade secrets and that won't fly. And words are defined here as, well, words -- English words. There'd be a separate fee per pixel/voxel, and then a catchall for non-language/non-image models.
- Workaccount2 3y agoCan't wait for the deluge of AI generated content dumped en masse on the internet purely to harvest "AI Dividends".
- knaik94 3y ago>For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model. I don't expect this lawsuit to lead anywhere. But if it does, I hope it leads to some clear laws regarding data privacy and how TOS is binding. The recent ruling regarding web scraping makes the case against OpenAI a lot weaker. [1] Data scraping publicly available data is legal. People didn't need consent to having their data be used, there was an implicit assumption the moment the data was published to the public, like on reddit or youtube. I keep seeing this idea reoccur in the suit: >Plaintiff ... is concerned that Defendants have taken her skills and expertise, as reflected in [their] online contributions, and incorporated it into Products that could someday result in [their] professional obsolescence ... Anyone is able to file a suit, I wish people stopped assuming that a news report automatically means it has merit. 1. https://www.natlawreview.com/article/hiq-and-linkedin-reach-proposed-settlement-landmark-scraping-case https://www.natlawreview.com/article/hiq-and-linkedin-reach-...
- tedivm 3y agoThe lawsuit is far more nuanced than you're letting on. There are several aspects that come into play- * Was it published publicly? This is basically defined in the courts as "if you make an unauthenticated web request does the data return?". This is where scraping comes in- if you make the data available without authentication you can't enforce your TOS, because you can't validate that people actually even accepted the TOS to begin with. * Is the data able to be copyrighted? This is where things are interesting- facts can not be copyrighted, which is why a lot of scrapers are able to reuse data (things like weather, sports scores, even "for hire" notices can be considered factual). * If it would typically be considered covered by copyright, does fair use come into play? * Are there any other laws that come into play? For example, GDPR, CCPA, or other privacy laws can still add restrictions to how data is collected and used (this is complicated by the various jurisdictions as well) * Was the work done with the data transformative enough to allow it to bypass copyright protections? This goes back to when Google was scanning books. Because they were making a search engine, not a library, their search tool was considered transformative enough to allow them to continue. It's not enough to say "because it's on the internet, it's fair game for everyone to use". This is a really complicated area where things are evolving rapidly, and there's a lot of intersecting law (and case law) that comes into play.
- mbgerring 3y agoWow, I really don’t get it, if I were to memorize billions of pages worth of people’s private messages and medical records, then recited them live in the Internet, would that be a crime??
- Cthulhu_ 3y agoYeah, unless you had permission from the authors to do so.
- deleted 3y ago[deleted]
- pessimizer 3y agoWhat exactly is the difference between that and downloading billions of pages worth of people's private messages and medical records, and putting them in a torrent? If there is a difference, I should be able to make a disability discrimination case under the ADA and erase that difference, because I don't have the memory to do that without the aid of a prosthetic (i.e. my laptop.)
- Workaccount2 3y agoYes, it would be. Thankfully AI doesn't work by memorization.
- chasing 3y agoI mean, it ingested all of the content from my blog. Without my permission. It's not a major part of their corpus of data, but still -- I wasn't asked and I don't really care to donate work to large corporations like that. So the technology is cool, but I'm firmly of the stance that they cut corners and trampled peoples' rights to get a product out the door. I wouldn't be entirely unhappy if this iteration of these products were sued into the ground and were forced to start over on this stuff The Right Way.
- scarface_74 3y agoYou don’t get to make information publicly available. But not publicly available. If you want your blog to be restricted, put it behind a login
- chasing 3y agoYes I do. I own the work I create, even if it's publicly available. I do get to decide what happens with it.
- 6bb32646d83d 3y agoAnyone can read your blog and then post their own blog post using knowledge they learned while reading yours. ChatGPT "learned" from your blog that same way
- tumult 3y agoComputers aren't people. Software isn't humans.
- mrtranscendence 3y agoSince the way GPT "learns" is not materially similar to how a human learns, I don't see why this talking point is particularly relevant. Nothing stops the courts from distinguishing between an AI and a human with regard to what may be permissible.
- shubhamgrg04 3y agoIf we dissect this case, it seems to revolve around two central questions: what constitutes 'public' data and to what extent can AI models leverage such data without infringing upon individual privacy. This lawsuit may well set a significant precedent in defining the boundaries of AI ethics and data privacy.
- ravenstine 3y agoDoes every business in the 21st century need to be some form of low-level scam in order to make headway and grow enough to satisfy VCs or investors?
- tsunamifury 3y agoYes, that’s where the disruption comes from. In all Seriousness that’s the advantage left in an efficient market. Look at all the share economy players it boils down to offload the risk, labor and debt but keep the margin.
- gumballindie 3y agoIt does seem like. Playing by the rules limits growth. Stealing, cheating, lying, manipulating, are the endless money cheat, particularly in societies where most people abide by the rules. Once they hit big they find willing politicians to adjust laws to their favor. Rinse and repeat.
- elforce002 3y agoOnly 3? They should go for the whole 10, and settle for 1. Now that the gates are open, we'll probably be entering the "free money" cycle soon.
- seaerkin 3y agoDo we think this is related to media platforms seemingly walling themselves off? Requiring accounts to view content, removing API access. It seems if they can silo data off and make it difficult to access at a large scale, then they are the gatekeeper of the data and can control usage and pricing.
- BlueTemplar 3y agoNo, this always happens with platforms once they feel they have attracted enough users : for instance it happened with Twitter in 2013, or see also what happened to XMPP after Google and Facebook have adopted it, or Reddit going closed source in 2017...
- I_am_tiberius 3y agoWhat I noticed is that the privacy setting which should prevent OpenAI to use my data for training purposes, was already deleted twice and I had to set it again. No idea what that means and if the data that I entered before I noticed that setting was gone is now being owned by OpenAI. Anyway, it is obvious that privacy is no priority to them. Also, it's known that YC companies are informally being told they should not worry about privacy while scaling up. Open AI is not a YC company, but its culture is definitely derived from it.
- CuriousSkeptic 3y agoAs I understood it that setting is an opt-out cookie. So must be set on all new browser sessions. Seems to be a blatant violation of GDPR. So I assume they’ll be fined for it sooner or later and forced to cleanup the training data anyway.
- scrollaway 3y agoHow is that a GDPR violation? GDPR doesn’t prevent opt outs of this kind of thing.
- CuriousSkeptic 3y agoIn the sense that consent requires active opt-in. The passive “opt-in” by failing to set the cookie doesn’t count as consent. So if they’re claiming they have the right to process data on the legal basis of consent, and they claim the absence of that cookie constitutes that consent, then they have no legal basis, and are thus in violation of the law.
- m3kw9 3y agoWhy not sue for 30 billion instead if you are to go full stupid on the price
- janvanlooy 3y agoWe talk to a lot of companies and many want to start using generative AI but are afraid of litigation. As long as it is not clear on which data a given model has been trained and that it is explicitly licensed permissively by the owner you are not sure what can happen. We are actually working on a tool to create billion-size free-to-use Creative Commons image datasets and prepare them for training models like Stable Diffusion. There is a blogpost about it here: https://blog.ml6.eu/ai-image-generation-without-copyright-infringement-a9901b64541c https://blog.ml6.eu/ai-image-generation-without-copyright-in...
- anigbrowl 3y agoFishing expedition. Will probably get thrown out because no particular injury can be enunciated. OpenAI scraped HN as well, and I don't consider my HN posts private because anyone can come here and read them, including artificial intelligences.
- zer0c00ler 3y agoMaybe as a result OpenAI will have to publish how they trained and what data was exactly used.
- ouraf 3y agoIf anything major comes out of this, is probably EVEN MORE prompts and popups asking for permission to use your data. even with GDPR, data collection and sales never stopped, it just made things more annoying by transforming every webpage into a granular term of service to continue doing the same. It isn't even turned off by default. Many sites just give you an "i accept" button or even if you want to manage the preferences, the "accept all choices" button is where the "confirm my choices" should be. Bigger companies will just append this to their TOS and push it down the customer's throat. That if MS doesn't settle out of court and the case gets thrown together with any major oppositon to the data mining
- locallost 3y agoI wonder if we'll see a license for content that forbids its use for training of language models.