3 ms·
Are you blocking OpenAI access to your website?
I read over Twitter that a lot of founders are blocking GPTBot (or OpenAI web crawler) access to their websites or startups, are you? Why?
- sysadm1n 3y agoYou can block OpenAI, but you should really be blocking all bots and bad actors who may be scraping your site, scanning for vulns, or the new kid in town: using your site as training data. Robots.txt is not enough, and whilst the major players (Google, Bing etc) honor robots.txt, it can be completely ignored by other actors.
- tikkun 3y agoHow to? Is there an easy thing with cloudflare or similar and what’s the pricing
- 7moritz7 3y agoJust turn on Bot Fight Mode, that should be sufficient. If you have a very high risk scenario permanently turn on the under attack mode and it will throw a Turnstil captcha at all suspicious requests. In that scenario you should probably set up a WAF rule to block all Tor requests too (it's under the "country" selector).
- KomoD 3y agoNo, not intentionally at least, very possible they get stopped by Cloudflare though.
- version_five 3y agoI encourage it, I'd be happy for anything I do to make it into the repertoire of an language model. What is the downside, I'm sharing it on the internet anyway. It's people trying to protect dying business models (ads, thankfully) or who are upset someone else found a use for their data and want to retrospectively rent seek that get worried about crawling.
- eureka-belief 3y agoDying business models? I’d disagree. If you spend time making thoughtful content, it’s not unreasonable to look for ways to monetize it. The issue with scrapers and chatgpt is that it’s essentially ripping off your innovation with no compensation. If people don’t protect their content, then what incentive is there to produce something new?
- version_five 3y agoBefore chatgpt people were putting stuff online and were fine with their lot. OpenAI found something else to do with that data, and it feels like people are upset they aren't getting a cut, even if their economics haven't changed. I can see that potentially having info available on chat gpt would mean less page views for content that exists to show ads, but I'm not sympathetic to that business model. People who have models other than trying to get views for ads are just as incentivized to create new content as they were before.
- meiraleal 3y agoHumans are creative long before the sad advent of copyright and patents and we are finally starting to see an end for this craziness of patenting ideas. If your idea can't make you money without using the state legal system to prevent others from copying it, your idea isn't good. And the government shouldn't subside it.
- ChildOfChaos 3y agoAs long as it's not just copying and pasting the content I disagree. If I read the article and absorbed it, I have 'learnt' from the content. Then maybe I go on to write an article or create content based on what I learnt from your article. Most writing is just this, mixed with people connecting some dots together. How is this different from an AI reading and learning from your content? AI is happening, it's silly to block AI reading your content and fairly egotistical to think that your content really matters. If you are putting it out there for others to learn from, why can't an AI be trained on that data, learn from it and then help others learn from it by using the AI tool, you are still adding to the human pool of knowedege.
- Lariscus 3y agoTraining LLMs on data that you don't own the rights to is copyright infringement. Why should I continue to feed a machine that already violated my rights?
- muzani 3y agoTwo different things there though. It's a good example of something that is unethical but not (yet) illegal. Because it's legal, it makes more sense to block it.
- version_five 3y agoTraining LLMs on data that you don't own the rights to is copyright infringement. No it isn't. There is nothing indicating this is true and lots of evidence that it's either fair use or more likely not a use at all. Just pretending or asserting that it's a violation doesn't change it's legal status.
- johneth 3y agoI'd block it on all of my sites, except in the limited cases where it's advantageous for me to let them scrape it. So, blog posts, things like that: no. Things like technical documentation, that users of ChatGPT might find useful, and that would benefit me if those users can access if it's included there: sure.