13 ms·
Ask HN: How do I stop companies from scraping my site?
This article says that one of the datasets for chatGPT was obtained by scraping all links with reddit with more than 2 upvotes:
https://www.searchenginejournal.com/how-to-block-chatgpt-from-using-your-website-content/478384/ https://www.searchenginejournal.com/how-to-block-chatgpt-fro...
I don't want big companies to scrape my content and then sell it on their platform.
Novelty of LLM output may be an open question, but input is just someone else's stuff. I assumed that default copyright protects from this kind of bullshitttery. That it says that work can not be used, adapted, copied without creators permission. (I can only guess that it was allowed to happen, because that's the first time someone stole IP in this particular manner on this scale?) But now that we know that it's a thing, how can we maintain ownership of the inputs legally and engineering wise?
- nonrandomstring 3y agoMake some of your content, which is invisible to normal users, so exceptionally toxic to the bots that they will begin to avoid your site by choice.
- thrlsxbfn 3y ago[dead]
- lnalx 3y agoHow can doing this cannot affect your SEO?
- altdataseller 3y agoServe different content to non search engine bots
- speedgoose 3y agoI often read that doing that was a big no-no as Google would flag your website as cheating. But that may be a SEO legend.
- rchaud 3y agoIf SEO's a concern, then there's nothing you can do about the LLM bit. Either put the website behind a login, or deal with robots sucking everything up. Pre-LLMs, there were already many websites showing scraped versions of other people's sites, and ranking higher in search results.
- thrlsxbfn 3y ago[dead]
- quickthrower2 3y agoYou can gate it behind a login. Even a simple self made captcha (what is 2 + 7?) to reveal the content would probably stop LLMs. But hurt seo so you have to not rely on that. Do a medium and show a paragraph first then the login/captcha to continue.
- Slavaqua 3y agoWouldn't it severely affect SEO? I don't know, but I assume, that scrambling article in to a word soup and exposing it to crawlers would not solve the SEO problem. Would it?
- joshuanapoli 3y agoOpenAI’s web scraping bot will respect your robots.txt. https://platform.openai.com/docs/gptbot/disallowing-gptbot https://platform.openai.com/docs/gptbot/disallowing-gptbot
- nicbou 3y agoIt's a bit late since it already crawled everything once.
- Kim_Bruning 3y agoSpidering and Web scraping in and of itself is permitted by law (eu) or fair use (us). It is not considered illegal and many people do it for many different purposes. There are many common tools and libraries to help with this on linux, mac os and windows. It's even legal to keep the copies in a searchable database. [1] What is not then permitted is to give other people copies, or publish them on your website, or pretend it's your own work etc... When it comes to LLMs or image generation models, they don't keep any copies and they don't generate any copies either, so they consider themselves to be well in the clear. [2] If you want to stop people scraping your stuff anyway, you can always use robots.txt , or put things up behind a login-wall. Do consider the morality of what you are doing though. Personally I feel that published data should be scrape-able where practical. [1] https://en.wikipedia.org/wiki/Authors_Guild%2C_Inc._v._Google%2C_Inc https://en.wikipedia.org/wiki/Authors_Guild%2C_Inc._v._Googl.... (you're even allowed to do this with physical books) [2] https://www.uspto.gov/sites/default/files/documents/OpenAI_RFC-84-FR-58141.pdf https://www.uspto.gov/sites/default/files/documents/OpenAI_R... (With apologies for my crude summary of their actual arguments)
- satyrnein 3y agoPublishing a derivative work of something you scraped is not permitted. The question is whether LLMs are derivative works, which has not been settled in courts yet.
- Kim_Bruning 3y agoTechnically correct on the point that it has not been settled in courts yet. For eg stable diffusion I actually dove a bit deeper on this: The LAION dataset is -in itself- almost indisputably legal: It contains URLs of images, and tags describing those images[1]. To be sure: _It does not contain actual images_. People have made this mistake before, and LAION quite correctly reply that they don't have them. [2] Stable diffusion's models are much smaller [3] than the LAION dataset which they use as their source. Again, the models do not contain images. This is pretty suggestive that SD models should not be considered derivative works. Of course sometimes funny things happen in court, so we can't be sure until a court actually decides. I don't currently run my own GPT or LLAMA, but those should have similar properties. (in fact, OpenAI appears to claim so.) [1] https://laion.ai/blog/laion-5b/ https://laion.ai/blog/laion-5b/ (if you look for the files in any of the variations, you'll see it's in the order of terrabytes) [2] https://www.vice.com/en/article/pkapb7/a-photographer-tried-to-get-his-photos-removed-from-an-ai-dataset-he-got-an-invoice-instead https://www.vice.com/en/article/pkapb7/a-photographer-tried-... [3] https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0 https://huggingface.co/stabilityai/stable-diffusion-xl-base-... : see under "files and versions", it's sd_xl_base_1.0.safetensors. File size is 6.9 GB
- microflash 3y agoI wrote about this a while ago [1]. Unfortunately, with robots.txt, you're at the mercy of crawlers. They may respect it or ignore it altogether. You can block IP addresses but many crawlers may not even use static IP addresses. You can go to extremes and put your content behind a login, as others have suggested. But that would also create friction for your intended audience. [1]: https://www.naiyerasif.com/post/2023/09/30/blocking-ai-web-crawlers/ https://www.naiyerasif.com/post/2023/09/30/blocking-ai-web-c...
- Slavaqua 3y agoIt sounds like a loosing battle to manually keep track of all bots and their deployment IPs that generate datasets that LLM's might use or start using in the future. There must be a legal solution, a licence that forbids the use of content for training without explicit permission from the author.
- hermannj314 3y agoI sell apples at the market. How can I prevent people from buying my apples just to make pies or tarts and pretending those products are theirs? I grew the apple, I should be able to decide what people do with it. I have a sign that says the apples are only for eating, but people are ignoring it.
- firejake308 3y agoI think this is a poor analogy. The better comparison would be if people were buying my apples and extracting the seeds to plant their own trees, then selling their own apples at a lower price than me and putting me out of business. For example, I don't read Medium articles as often as I used to, in part because I can just ask my programming questions directly to ChatGPT. Granted, this is to be expected because ChatGPT can customize it's response based on my unique situation whereas the Medium article is fixed, but it's still unfair to the author of the Medium article that ChatGPT was trained on, because without their effort, none of this would have been possible. And yet, because I am no longer generating as revenue for Medium, the original authors are no longer being compensated for the value they provided me. I personally believe that LLM hosts should be required to pay for training data. With RAG, I think it makes sense to charge per-query, but for base models, content authors would probably need to organize into larger groups to facilitate payment.
- satyrnein 3y agoYour apple analogy is completely legal and ethical. The alternative is Monsanto's attempt at DRM with their "terminator" seeds, which many found to be unethical.
- BlueTemplar 3y agoSadly, this is not even clear cut, though of course this is a particularly egregious example. For instance depending on jurisdiction it might be illegal to sell seeds that can make plants that can produce viable seeds, unless you have a specific authorisation and depending on the species.
- rchaud 3y agoIf it's worth that much, just put it behind a login.
- deleted 3y ago[deleted]
- datavirtue 3y agoYou can take it off the web. Frankly these concerns confuse me. You published so people can find value, no?
- precompute 3y agoYou literally can't. Data collection at scale is one of the three pillars of the current "AI" hype. It's an issue and it was by design (thanks, three-letters!) and now it will never ever be revoked. Assume all data on the internet is logged in some form and is available to people who shouldn't be able to access it and that those people use it to model things at scale, or sometimes, just store it so they can reinterpret it later. Storing data is cheap. Transmitting data is cheap. MITMing the world's data flow? Priceless.
- tikkun 3y ago1. Robots.txt will help somewhat [1] 2. Put it behind a login wall [1]: https://platform.openai.com/docs/plugins/bot https://platform.openai.com/docs/plugins/bot
- ReflectedImage 3y agoCommerical web scraping services exist that can scrap any web content. The best you can do is ask nicely in a robots.txt file.
- mensetmanusman 3y agoAdd a login with login information visible to humans. When logging in, have a prompt that says ‘ai systems are not allowed to login lest you pay $€£’
- aaron695 3y ago[dead]
- golly_ned 3y agoSteve Huffman? Is that you?
- throwawaydghvhv 3y agoOn my site I don't need SEO, so I sprinkled it with invisible links leading to an endless Markov chain generated walls of garbage text. The robots seem to love it! (If you go similar route, don't forget rate limiting)
- Slavaqua 3y agoI don't understand your solution. The actual content is still being scraped, is it not?
- xyzal 3y agoI think they are interested in rendering the corpus of text from their site unsuitable for training LLMs by stuffing it with nonsense more than preventing scraping.
- askiiart 3y agoI have a similar approach: I just send AIs the Bee Movie script!
- JoeyBananas 3y agoOnly distribute your blog posts to people who sign an NDA.
- YaBa 3y agoYou don't. If the information is valueable enough it will be scrapped, you can use captchas, logins and all the tips and tricks available, but at the end of the day, if a dev wants to scrape the info, it will.
- theyknowitsxmas 3y agoMake it a scrollable RTMP stream of a VPS with no controls. or Take a screenshot and convert it to CSS. If you need hyperlinks, use image maps on blank transparent PNG's with top z-index.
- wharfjumper 3y agoCloudFlare Turnstile may help. Presumably it would require some coding by you to only display content once the test is passed by the client. I am not affiliated with CloudFlare in any way. [1] https://www.cloudflare.com/products/turnstile/ https://www.cloudflare.com/products/turnstile/
- xyzal 3y agohttps://kudurru.ai/ https://kudurru.ai/