5 ms·
Ask HN: Can I stop OpenAI from using my website data as training data?
So OpenAI says that we cant use the data generated by OpenAI API to train competing models. Can I have a similar ToS for my website so that I don't allow OpenAI to train on my website data for the next versions of GPT?
- deleted 4y ago[deleted]
- MikeTheRocker 4y agoCurious why?
- JohnFen 4y agoI can't speak for OP, but I personally don't want any of my web contents to be used to train these things. Why? Because I'm not convinced these systems are good things in that sort of use case, and until/unless I am, I don't want to contribute to them.
- nutanc 4y agoFor the same reason OpenAI is not allowing text generated by its system to train competing systems. I want to maybe build a competing system with my data and don't want OpenAI to be a competitor in that space.
- dmux 4y agoThere were some recent (but limited) discussions about this: - "Ask HN: Prevent ML like GPT from using public posts like this one?" https://news.ycombinator.com/item?id=33980566 https://news.ycombinator.com/item?id=33980566 - "Ask HN: Does ChatGPT respect Robots.txt?" https://news.ycombinator.com/item?id=35027823 https://news.ycombinator.com/item?id=35027823
- deostroll 4y agoWell, I certainly hope that the Hacker News FM podcast on spotify doesn't come to an end. But as long as stuff is on the internet and accessible, I don't see the point of denying people visiting the webpage. Restricting access helps a little. But in the end the kingdom of humanity is all good and bad. Your life will only be worse if all you think or worry about is the bad. PS: for reasons unrelated to the topic I do believe that podcast might curtain soon.
- deleted 4y ago[deleted]
- ktbwrestler 4y agoCorrect me if I’m wrong,but there’s no law saying that you can’t crawl a site if it has a robots.txt configured, it’s just a convention, and you wouldn’t be able to have any recourse telling a company they can’t pull your info if it’s on the public web.
- ktbwrestler 4y agoYour best path forward might be trying to use some sort of reverse proxy that denies their egress IPs, but you could just be playing whack a mole
- JohnFen 4y agoThis is correct. However, Apache (and, I assume, all other webservers) does let you examine the IP address and user agent strings and make decisions about how to handle them. This is how I stopped crawlers from my sites. I used robots.txt, but backstopped that with detecting undesirable IP addresses and user agent strings. Those would get an error page instead of the actual page. It does mean you need to keep an eye on your logs to spot new crawlers coming around, though.
- darreninthenet 4y agoI'm curious as to why it's not considered to be a DMCA circumvention of a technology.
- JohnFen 4y agoIANAL, but here's the relevant part of the law: “No person shall circumvent a technological measure that effectively controls access to a work protected under this title.” I'm guessing that because robots.txt does not actually attempt to prevent access, it also doesn't "effectively control access" and so isn't covered by this. The clear intent of the law is to prevent actual cracking of access controls.
- darreninthenet 4y agoSo playing devils advocate, isn't it supposed to prevent crawlers (etc) from reading (ie accessing) those web pages?
- sourcecodeplz 4y agoYes, make it private.
- MontgomeryPi2 4y agoAs the answer would seem to be 'no' (crawlers can ignore robots.txt rules), I'm wondering if GPT is going to usher in an era where new content is simply not made available to view/crawl on the web. E.g. Want to know all the cool events happening this weekend in NYC? Ask our vertical gpt-chat site to find out.
- speedgoose 4y agoOpenAI uses the common crawl dataset and I would expect them to respect the robots.txt rules.
- JohnFen 4y ago> I'm wondering if GPT is going to usher in an era where new content is simply not made available to view/crawl on the web. I've stopped adding new stuff to the public areas of my websites. Not as a result of GPT specifically, but as a result of a dramatic increase in the amount of scraping being done to support things that I don't approve of. AI training in general being one of them.
- dev_0 4y agoUse simple authentication headers
- perrygeo 4y agoUltimately, whether web scraping for AI training data purposes is considered legal comes down to whether web scraping for ANY purpose is legal. This isn't clear one way or the other, though recent rulings indicate that US courts are firmly poised to declare web scraping of public content fully legal. IE there is nothing you can do to stop it other than making it not public.
- zzo38computer 4y agoI want to allow it, but only if they also allow data generated by them to be used to train competing models (and data from those competing models to also be used to train competing models of them too, etc). And if someone wants to copy the data for another purpose (e.g. making a database, or making a movie), they cannot restrict that purpose either. (I expect that the AI will still probably have similar kinds of problems than it does now regardless of whether or not it is allowed, though.)
- motbus3 4y agoWe were discussing this in my job some days ago. Legal team is still checking, but technically you can copyright and define usage license for your content. But if you write about common subjects or non copyright material you may be better off the internet