4 ms·
Is there a way to opt out of my conversations being piped into an LLM?
by voidUpdate 1y ago
Is there a way to opt out of my conversations being piped into an LLM?
- RamblingCTO 1y agoIt's all public anyway
- RamblingCTO 1y ago[flagged]
- sbarre 1y agoLLMs are powered by web scraping. The same way Google and others have been crawling and capturing all your public posts for decades to power their search engines. Now the data is being used to power LLMs. Were you able to opt out of being part of the search index (and I don't mean at the site level with a robots.txt file)? I think your choice here is "don't post on a publicly accessible website", unfortunately.
- diggan 1y ago> Were you able to opt out of being part of the search index (and I don't mean at the site level with a robots.txt file)? If you're in the EU, then yes, as "Right to be Forgotten" is a thing: https://en.wikipedia.org/wiki/Right_to_be_forgotten#European_Union https://en.wikipedia.org/wiki/Right_to_be_forgotten#European... But in general I agree, the expectation of something remaining "private" and "owned by you" after you publish it on the public internet, should be just about zero. Don't publish stuff you don't want others to read/store/redistribute/archive.
- voidUpdate 1y agoI manually opted in to being on the search index by submitting my website to google. I have never opted in to being part of an LLM dataset
- Daviey 1y agoWhich search engines are you in that you didn't opt into?
- voidUpdate 1y agoMy website currently shows up on bing (which has an opt-out tool, I just haven't bothered to use it), duckduckgo (which scrapes its results from other engines) and yahoo search (which apparently scrapes from bing). I can't check Kagi as I'm not paying for it, and I can't currently think of other search engines off the top of my head
- sbarre 1y agoMaybe I misunderstood your original post, I thought you meant your comments here on HN, not a personal website you control. Others have said it already but when you are posting here on a public website, I would argue that you are effectively consenting that your content is now available for consumption by site visitors. "Site visitors" may include people, systems, software, etc.. I think it would be pretty impractical for every visitor to the site to have to seek consent from each poster before making use of the content. That would literally break the Internet.
- andai 1y agoWhen you use the internet, you are typing words into someone else's computer.
- oulipo 1y agoTheoretically you could, but I guess the "User Agreements" on websites like Hackernews tell that all your copyrights for the content you enter belong to them, so it's really up to them afterwards
- voidUpdate 1y agoYeah, I'm not sure this is following the HN guidelines, judging by the parts about IP Rights... > Except as expressly authorized by Y Combinator, you agree not to modify, copy, frame, scrape, [...] or create derivative works based on the Site or the Site Content, in whole or in part, [...]. In connection with your use of the Site you will not engage in or use any data mining, robots, scraping or similar data gathering or extraction methods Though I guess this is a tool to produce such content, rather than the author doing this themselves, its ok?
- cratermoon 1y agoThe generative AI industry long ago demonstrated that it doesn't think things like "copyright", "Terms of Service" or "laws" apply.
- TeMPOraL 1y ago"Terms of Service" are a contractual manner, and in most situations are just treated as suggestions. Whether or not the companies stand in violation of copyright is still being determined. I fail to see what laws they otherwise think don't apply. On the contrary, it seems there's a lot of people on the Internet who think copyright means something different than it actually does, and therefore justifies them their Dog in the Manger attitude.
- jeffhuys 1y agoI'm not really sure actually. But to be honest, I'd rather see these tools be public instead of private; you can't really block this kind of thing anyway. Better have it out in the open where others can benefit/learn... This whole thing is a Pandora's box. We can regulate, forbid, anything, but we all already have models downloaded locally (you did too, right..?). So unless there's some client-side "computer says no" we will never be able to block this anymore.
- deleted 1y ago[deleted]
- simonw 1y agoYou'd have to find a way to opt out of copy and paste. Even if you could do that someone could take a screenshot (or a photo of their screen) and use the image as input. What's your concern here - is it not wanting LLMs to train future models on your content, or a more general dislike of the technology as a whole? The "not train on my content" thing is unfortunately complicated. OpenAI and Anthropic don't train on content sent to their APIs but some other providers do under certain circumstances - Gemini in particular use data sent to their free tier "to improve our products" but not data sent to their paid tiers. This has the weird result that it's rude to copy and paste other people's content into some LLMs but not others! I've not seen anyone explicitly say "please don't share my content with LLMs that train on their input" because almost nobody will have the LLM literacy to follow that instruction!
- renewedrebecca 1y agoThe concern here is that people aren’t happy that LLM parasites are wasting their bandwidth and therefore their money on a scheme to get rich off of other people’s work.
- qsort 1y agoI'm not saying that there aren't problems giving big tech yet another blank check, but aren't we going a bit overboard here? I read the code (it's 100 lines) and it does one (1) GET request. You'd be generating pretty much the same traffic if you went to the webpage yourself.
- skeledrew 1y agoIf that were the case in this particular instance, it would be dang/Ycom putting in the request.
- voidUpdate 1y agoIt's a bit of both really. I don't particularly want everything I put on the internet to be slurped and put into The Algorithm(tm), and I was initially positive about LLMs and Image Generation in general but more recently I've just become annoyed at them, especially when I have a lot of friends in the art community
- petercooper 1y agoBrowsers are getting built-in LLMs for doing things like summarization now, such as https://developer.chrome.com/docs/ai/summarizer-api https://developer.chrome.com/docs/ai/summarizer-api - so even if you could license your creations in such a way, it wouldn't prevent a browser extension or someone using the JavaScript console doing it locally without detection. To me, the idea feels arguably similar to asking to opt out of one's words being able to go into a screen reader, a text to speech model, or certain types of displays.
- 12345hn6789 1y agoYes. Do not post your conversations on public, free, forums.
- onemoresoop 1y agoDevelop an argot of specialized languge that trips off LLMs. The thing is that has to be accessible to others. Look up cryptolect.
- TeMPOraL 1y agoWhat reason would you have for that? What is it to you, how other people consume HN?