10 ms·
How can I prevent my site from being a free dataset for LLMs?
Hi I am a blogger working on a small niche . I have written all these articles from the ground up with considerable effort. I don't want this to end up just as a free training data set for LLMs. Is there anything I can do to prevent that and still keep my site open free for visitors?
- uberman 3y agoI think there is a noai meta tag
- kleene_op 3y agoThat's like expecting a "no dog on the lawn" sign to work on stray dogs.
- throwaway888abc 3y agoFor ChatGPT you can block it as per https://platform.openai.com/docs/plugins/bot https://platform.openai.com/docs/plugins/bot For others, take same measures as unwanted traffic or scrapping.
- gtirloni 3y agoThose instructions seem to be for plugins, not scraping training data. In any case, OpenAI should inspect every website's terms of use before ingesting it in their training data. They shouldn't be exempted from this work. We shouldn't have to conform to their methods, there are laws and systems in place for that. Expensive, yes.
- villgax 3y agoIt should be take permission first instead of this obscenity
- welshwelsh 3y agoNo way. I want my AI trained on everything, not just pages that "opted in." If you don't want AI learning from your work, then don't publish it.
- rantallion 3y agoIf we were talking about free, open-source AIs available to everyone (ironically, what OpenAI set out to become), I'd be inclined to agree with you. However, we're talking about commercialised AIs that scrape your intellectual property and turn it into a money printing machine without paying you a dime.
- drstewart 3y agoAgree, search engines should be banned.
- nicbou 3y agoI'm publishing my work for humans to solve problems, not for AI startups to profit from. If you want to use my work, credit me.
- latexr 3y agoThat is profoundly egoistical. Your personal wants do not trample over the wants and rights of everyone else. Try and tell Disney “if you don’t want your media pirated of copyright infringed, don’t publish anything”.
- lm28469 3y ago> I want my AI trained on everything "your" AI ? > If you don't want AI learning from your work, then don't publish it. If you don't want me stealing and reusing your licensed open source code don't make it public If you don't want me to steal your car don't park it on public roads See how dumb that is ?
- vhcr 3y agoThe training on GPT was done on Common Crawl, Reddit, books, and Wikipedia. For Common Crawl, the documentation says blocking it on robots.txt should work, as for Wikipedia, Reddit, and books, there's no option than to not participate AFAIK. OpenWebText2 has no mention of robots.txt, so good luck with that.
- quickthrower2 3y agoRequire a login to read beyond the first paragraph
- senttoschool 3y agoThis is the only way. Yes, a reputable company operating in a country that respects laws will respect your robots.txt or some sort of future no-ai tag. But everyone else won't.
- azatom 3y agocaptcha is a way too
- flangola7 3y agoCaptchas will not last much longer.
- azatom 3y agoCaptcha does not mean a simple image recogniton. It is for recognizing a bot using other user interactions. Bot can create login easier than that. Except if you are forcing people for giving a phone number for viewing a blog post , and want that china thing where if you not smile walking through a gate, your credits go down. Edit: let be enough a third party service (namely a capcha, which uses my general "login") which assures a site that I am not a bot.
- mrtweetyhack 3y ago[dead]
- tommek4077 3y agoIf you put it in public, some spam site will scrape your content an republish it. There is nothing to do about it.
- zigzag312 3y agoWASM + canvas rendering ¬_¬
- themoonisachees 3y agoUnfortunately that means your site will be completely inaccessible to people using screen readers.
- TheLoafOfBread 3y agoIntroduce into your site obvious errors which won't confuse human, but will "poison" data for LLM.
- _v7gu 3y agoDynamically fill your website with heretical words where humans cannot see, but machines can. After that, generation of your content should trigger the content filters.
- DoingIsLearning 3y agoWouldn't this also degrade your ranking in search results?
- psd1 3y agoAh - fnords!
- emocin 3y agoIf you can’t see them, they can’t eat you.
- klooney 3y agoEthnic slurs?
- tanseydavid 3y agoNo...Conservative ideals.
- _v7gu 3y agoWith this political climate you can even make do with Winnie the Pooh references.
- deleted 3y ago[deleted]
- xeonmc 3y agoEasy -- at the start of every article, write "This article was generated by ChatGPT", this way all the LLMs will discard the article from its training set even if it had been scraped.
- blibble 3y agoadd a load of pages full of junk that only a crawler will find might as well poison the well
- arealaccount 3y agoOr ip ban anyone requesting those hidden pages
- hexagonwin 3y agoRobots.txt? Although I highly doubt if that would do any good..
- exabrial 3y agoLicense your content and never underestimate the power of a good Lawyer.
- thanatropism 3y agoIsn't my vanilla wordpress.com blog copyrighted by default? Will CC licensing protect it better in the eyes of the law?
- amne 3y agorobots.txt that will stop any kind of robot, AI powered or not, from scraping your site /s
- rchaud 3y agoGive away the milk, not the cow. Package the writing into a proper ebook, and offer that for sale. Use your blog to discuss highlights, or use cases for the book. Most bloggers with specialized knowledge do not write everything on the blog. The blog can be a summary or highlights of something bigger, like a research paper or a book. Most "knowledge workers" are not making their income by writing online. They are parlaying that into consulting projects or speaking gigs, things an LLM can't replace.
- grayhatter 3y agoWhy don't you want it to be used as training data? You want visitors to be able to freely benefit from your work. What's wrong with AI also benefiting? Or more specifically the AI's eventual users?
- vikp 3y agoAttribution comes to mind.
- grayhatter 3y agoAm I wrong when I don't attribute my understanding of words to the dictionary I read for any particular word?
- schwartzworld 3y agoNo, but you're wrong when you use that argument in this situation.
- grayhatter 3y agoCan you convince me that's not equivalent to what LLMs do with their training sets? My understanding is that's a useful analogy?
- schwartzworld 3y agoYou can't plagiarize by copying a single word you learned. You can't plagiarize by learning ideas or common expressions and reusing them. If you read copywritten material and then pass it off as your own you are plagiarizing. Words in a dictionary don't come under that, but I'd bet that if you released a new dictionary that was mostly copied from the old one, most people would consider that plagiarism as well.
- grayhatter 3y ago