8 ms·
It's interesting to me that OpenAI considers scraping to be a form of abuse.
by Imnimo 6mo ago
It's interesting to me that OpenAI considers scraping to be a form of abuse.
- sabedevops 6mo agoSeriously. The hypocrisy is staggering!
- zer00eyz 6mo ago" Integrity at OpenAI .. protect ... abuse like bots, scraping, fraud " Did you mean to use the word hypocrisy. If not, I'm happy to have said it. I just want to note, that it is well covered how good the support is for actual malware...
- ProofHouse 6mo agoThe irony is thick
- nikitaga 6mo agoScraping static content from a website at near-zero marginal cost to its server, vs scraping an expensive LLM service provided for free, are different things. The former relies on fairly controversial ideas about copyright and fair use to qualify as abuse, whereas the latter is direct financial damage – by your own direct competitors no less. It's fun to poke at a seeming hypocrisy of the big bad, but the similarity in this case is quite superficial.
- AtlasBarfed 6mo agoBecause you say it is? I obviously disagree. I mean, on top of this we are talking about not-open OpenAI.
- bakugo 6mo agoThe cost is so marginal that many, many websites have been forced to add cloudflare captchas or PoW checks before letting anyone access them, because the server would slow to a crawl from 1000 scrapers hitting it at once otherwise.
- nslsm 6mo agoThe issue is that there are so many awful webmasters that have websites that take hundreds of milliseconds to generate and are brought down by a couple requests a second.
- bakugo 6mo agoOpenAI must be the most awful webmasters of all, then, to need such sophisticated protections.
- swagmoney1606 6mo agoAnd yet I have to pay in my time and cash to handle the constant ddos'es from the constant LLM scraping
- karlshea 6mo agoI don’t know what world you live in but it’s not this one.
- not2b 6mo agoI understand why OpenAI is trying to reduce its costs, but it simply isn't true that AI crawlers aren't creating very significant load, especially those crawlers that ignore robots.txt and hide their identities. This is direct financial damage and it's particularly hard on nonprofit sites that have been around a long time.
- stingraycharles 6mo agoThese are ChatGPT and Claude Desktop crawlers we’re talking about? Or what is it exactly? Are these really creating significant load while not honoring robots.txt? Genuinely interested.
- cruffle_duffle 6mo agoI bet dollars to doughnuts that 95% of the traffic is from Claude and ChatGPT desktop / mobile and not literal content scraping for training.
- crote 6mo agoThat wouldn't explain the 1000x increase in traffic for extremely obscure content, or seeing it download every single page on a classic web forum.
- duttish 6mo agoAnd doing it over, and over, and over and over again. Because sure it didn't change in the last 8 years but maybe it's changed since yesterdays scrape?
- miki123211 6mo agoThey seem to mostly be third-party upstarts with too much money to burn, willing to do what it takes to get data, probably in hopes of later selling it to big labs. Maaaybe Chinese AI labs too, I wouldn't put it past them. OpenAI et al seem to mostly be well-behaved.
- 6mo ago
- razingeden 6mo agoIt is direct financial damage if my servers not on an unmetered connection — after years of bills coming in around $3/mo I got a surprise >$800 bill on a site nobody on earth appears to care about besides AI scrapers. It hasn’t even been updated in years so hell if I know why it needs to be fetched constantly and aggressively, - but fuck every single one of these companies now whining about bots scraping and victimizing them, here’s my violin.
- gzread 6mo agoIf you can identify the scraper you should have a valid legal case to recover damages.
- thisislife2 6mo agoOnly if they had a robots.txt for their site.
- razingeden 6mo agoI hadn’t even considered that. Don’t know why that comment is greyed out or downvoted. It’s a static site that hasn’t been updated since 2016—- so it’s .. since been moved to cloudflare r2 where it’s getting a $0.00 bill, and it now has a disallow / directive. I’m not sure if it’s being obeyed because the cf dash still says it’s getting 700-1300 hits a day even with all the anti bot, “cf managed robots” stuff for ai crawlers in there. The content is so dry and irrelevant I just can’t even fathom 1/100th of that being legitimate human interest but I thought these things just vacuumed up and stole everyone’s content instead of nailing their pages constantly?
- gzread 6mo agoNo, it's still illegal to DDoS sites that don't have robots.txt.
- thisislife2 6mo agoYou are right, I hadn't considered that aspect.
- PunchyHamster 6mo ago> Scraping static content from a website at near-zero marginal cost to its server, vs scraping an expensive LLM service provided for free, are different things. I bet people being fucking DDOSed by AI bots disagree Also the fucking ignorance assuming it's "static content" and not something needing code running
- Den_VR 6mo agoI miss the www where the .html was written in vim or notepad.
- holler 6mo agoahh yes, fresh off reading "Html For Dummies" I made my first tripod.com site
- consp 6mo agoJust did that for a test frontend for a module I needed to build (not my primary job so don't know anything about UI but running in browsers was a requirement), so basic HTML with the bare minimum of JS and all DOM. Colleagues were very surprized. And yes, vim is still the goto editor and will be for a long time now all "IDE" are pushing "AI" slop everywhere.
- mghackerlady 6mo agoIt still can be. Do it. Go make your website in M$ Frontpage, for all I care
- alsetmusic 6mo agoHave you not seen the multiple posts that have reached the front page of HN with people taking self-hosted Git repos offline or having their personal blogs hammered to hell? Cause if you haven't, they definitely exist and get voted up by the community.
- nozzlegear 6mo agoAre they, actually?
- sandeepkd 6mo agoLets not try to qualify the wrongs by picking a metric and evaluating just one side of it. A static website owner could be running with a very small budget and the scraping from bots can bring down their business too. The chances of a static website owner burning through their own life savings are probably higher.
- expedition32 6mo agoPerhaps the long play is to destroy all small hobby websites until only a AI directed web is left.
- miki123211 6mo agoIf you're truly running a static site, you can run it for free, no matter how much traffic you're getting. Github pages is one way, but there are other platforms offering similar services. Static content just isn't that expensive to host. THe troubles start when you're actually running something dynamic that pretends to be static, like Wordpress or Mediawiki. You can still reduce costs significantly with CDNs / caching, but many don't bother and then complain.
- jazzyjackson 6mo agoIt's true it can be done but many business owners are not hip to cloudflare r2 buckets or github pages. Many are still paying for a whole dedicated server to run apache (and wordpress!) to serve static files. These sites will go down when hammered by unscrupulous bots.
- ezrast 6mo agoSetting aside the notion that a site presenting live-editability as its entire core premise is "pretending to be static", do the actual folks at Wikimedia, who have been running a top 10 website successfully for many years, and who have a caching system that worked well in the environment it was designed for, and who found that that system did not, in fact, trivialize the load of AI scraping, have any standing to complain? Or must they all just be bad at their jobs? https://diff.wikimedia.org/2025/04/01/how-crawlers-impact-the-operations-of-the-wikimedia-projects/ https://diff.wikimedia.org/2025/04/01/how-crawlers-impact-th...
- heyethan 6mo ago[flagged]
- deleted 6mo ago[deleted]
- lm411 6mo agoThat is ridiculous. You imply that "an expensive llm service" is harmed by abuse, but, every other service is not? Because their websites are "static" and "near-zero marginal cost"? You have no clue what you are talking about.
- camillomiller 6mo agoWell he’s a simp
- the_sleaze_ 6mo ago60% of our traffic is bot, on average. Sometimes almost 100%.
- deleted 6mo ago[deleted]
- AmbroseBierce 6mo agoIt's not like those models are expensive because the usefulness that they extracted from scraping others without permission right? You are not even scratching the surface of the hypocrisy
- not_your_vase 6mo ago> net-zero marginal cost Lol, you single-handedly created a market for Anubis, and in the past 3 years the cloudflare captchas have multiplied by at least 10-fold, now they are even on websites that were very vocal against it. Many websites are still drowning - gnu family regularly only accessible through wayback machine. Spare me your tears.
- make3 6mo agoAbsolutely not, the former relies on controversial ideas to qualify as legal. Stealing the content from the whole planet & actively reducing the incentive to visit the sites without financial restitution is pretty bad.
- SkiFire13 6mo ago> Scraping static content How do you know the content is static?
- wolvoleo 6mo agoIt's more ironic because without all the scraping openai has done, there would have been no ChatGPT. Also, it's not just the cost of the bandwidth and processing. Information has value too. Otherwise they wouldn't bother scraping it in the first place. They compete directly with the websites featuring their training data and thus they are taking away value from them just as the bots do from ChatGPT. In fact the more I think of it, I think it's exactly the same thing.
- expedition32 6mo agoThis leads me to thinking: I ask chatGPT a question and they get the answer from gamefaqs. But what happens if gamefaqs disappears because of lack of traffic? Can LLM actually create or only regurgitate content.
- stefanka 6mo agoThey cannot create original content.
- wolvoleo 6mo agoWell they can make some up, like hallucination. That's an additional problem: when the original site that provided the training data is gone: how can they use verify the AI output to make sure it's correct?
- wolvoleo 6mo agoIt will remain in their scraped data so they can keep including it in their later training datasets if they wish. However it won't be able to do live internet searches anymore. And it will not generate new content of course. Especially not based on games released after the site codes down so it doesn't know. Though it could of course correlate data from other sources that talk about the game in question.
- Aerroon 6mo ago>Can LLM actually create or only regurgitate content. Contrary to what others say, LLMs can create content. If you have a private repo you can ask the LLM to look at it and answer questions based on that. You can also have it write extra code. Both of these are examples of something that did not exist before. In terms of gamefaqs, I could theoretically see an LLM play a game and based on that write about the game. This is theoretical, because currently LLMs are nowhere near capable enough to play video games.
- VadimPR 6mo agoGetting scraped by abusive bots who bring down the website because they overload the DB with unique queries is not marginal. I spent a good half of last year with extra layers of caching, CloudFlare, you name it because our little hobby website kept getting DDoS'd by the bots scraping the web for training data. Never in 15 years if running the website did we have such issues, and you can be sure that cache layers were in place already for it to last this long.
- cicko 6mo agoInteresting how other people's cost is "near-zero marginal cost" while yours is "an expensive LLM service". Also, others' rights are "fairly controversial ideas about copyright and fair use" while yours is "direct financial damage". I like how you frame this.
- gmerc 6mo agoIt’s not for techbros to decide at what threshold of theft it’s actually theft. “My GPU time is more valuable than your CPU time” isn’t a thing and Wikipedias latest numbers on scraping show that marginal costs at scale are a valid concern
- platybubsy 6mo agoBait or genuine techbro? Hard to say
- 9864247888754 6mo ago[dead]
- grishka 6mo ago> Scraping static content from a website at near-zero marginal cost to its server It's not possible to know in advance what is static and what is not. I have some rather stubborn bots make several requests per second to my server, completely ignoring robots.txt and rel="nofollow", using residential IPs and browser user-agents. It's just a mild annoyance for me, although I did try to block them, but I can imagine it might be a real problem for some people. I'm not against my website getting scraped, I believe being able to do that is an important part what the web is, but please have some decency.
- lelanthran 6mo agoI don't think a rule along the lines of "Doing $FOO to a corporate is forbidden, but doing $FOO to a charitable initiative is fine" is at all fair. What "$FOO" actually is, is irrelevant. I'm curious how you would convince people that this sort of rule is fair. The corp can always ban users who break ToS, after all. They don't need any help. The charitable initiative can't actually do that, can they?
- nickphx 6mo agoSpeak for yourself.
- xmcqdpt2 6mo agoAI providers also claim to have small marginal costs. The costs of token is supposedly based on pricing in model training, so not that different from eg your server costs being low but the content production costs being high. And in many cases AI companies are direct competitors (artists, musicians etc.) (TBH it's not clear to me that their marginal costs are low. They seem to pick based on narrative.)
- ungreased0675 6mo agoYou’re describing the tragedy of the commons. No single raindrop thinks it’s responsible for the flood.
- cindyllm 6mo ago[dead]
- unsungNovelty 6mo ago"near-zero marginal costs". For whom exactly???? https://drewdevault.com/2025/03/17/2025-03-17-Stop-externalizing-your-costs-on-me.html https://drewdevault.com/2025/03/17/2025-03-17-Stop-externali...
- ori_b 6mo agoMy website serving git that only works from Plan 9 is serving about a terabyte of web traffic monthly. Each page load is about 10 to 30 kilobytes. Do you think there's enough organic, non-scraper interest in the site that scrapers are a near-zero part of the cost?
- andrepd 6mo ago> Scraping static content from a website at near-zero marginal cost to its server The gall. https://weirdgloop.org/blog/clankers https://weirdgloop.org/blog/clankers
- foobiekr 6mo agoYou are, of course, ignoring the production costs of the static content that OpenAi is stealing. Stop justifying their anti-social behavior because it lines your pockets.
- mcfedr 6mo agoI'm sure the copyright holders would consider your use of their content as direct financial damage
- Aurornis 6mo agoI interpreted scraping to mean in the context of this: > we want to keep free and logged-out access available for more users I have no doubt that many people see the free ChatGPT access as a convenient target for browser automation to get their own free ChatGPT pseudo-API.
- wolvoleo 6mo agoThis is bad why? Well yeah for openai because all they want it to be is a free teaser to get people hooked and then enshittify. Morally I don't see any issues with it really.
- lelanthran 6mo ago> I have no doubt that many people see the free ChatGPT access as a convenient target for browser automation to get their own free ChatGPT pseudo-API. Not that hard - ChatGPT itself wrote me a FF extension that opened a websocket to a localhost port, then ChatGPT wrote the Python program to listen on that websocket port, as well as another port for commands. Given just a handful of commands implemented in the extension is enough for my bash scripts to open the tab to ChatGPT, target specific elements, like the input, add some text to it, target the relevant chat button, click it, etc. I've used it on other pages (mostly for test scripts that don't require me to install the whole jungle just to get a banana, as all the current playright type products do). Too afraid to use it on ChatGPT, Gemini, Claude, etc because if they detect that the browser is being drive by bash scripts they can terminate my account. That's an especially high risk for Gemini - I have other google accounts that I won't want to be disabled.
- heyethan 6mo ago[flagged]
- crote 6mo agoVery few websites are truly static. Something like a Wordpress website still does a nontrivial amount of compute and DB calls - especially when you don't hit a cache. There's also the cost asymmetry to take into account. Running an obscure hobby forum on a $5 / month VPS (or cloud equivalent) is quite doable, having that suddenly balloon to $500 / month is a Really Big Deal. Meanwhile, the LLM company scraping it has hundred of millions of VC funding, they aren't going to notice they are burning a few million because their crappy scraper keeps hammering websites over and over again.
- raincole 6mo agoQuite sure even literal thieves would consider thievery a form of abuse.
- littlestymaar 6mo agoYeah, they know it's bad, they just don't think the rules apply to them.
- kamban 6mo agoYou nailed it.
- catoc 6mo agoIt’s only bad if you’re a closed, for-profit entity </sarcasm>
- lukan 6mo agoWas that sarcasm? Speaking of it, what parts of OpenAI are still open?
- catoc 6mo agoI know, always hard to tell on HN. Added the relevant declarative tag
- deleted 6mo ago[deleted]
- reactordev 6mo agoThe front door…
- deleted 6mo ago[deleted]
- 6mo ago
- axegon_ 6mo agoThe levels of irony that shouldn't be possible...
- miki123211 6mo agoIt's not scraping they're concerned about, it's abusing free GPU resources to (anonymously) generate (abusive) content.
- wiseowise 6mo agoChurch, politicians, moralists are all the biggest hypocrites that want to teach you something.
- newsoftheday 6mo agoI agree on politicians, no idea what a "moralist" is supposed to be but there are good and bad churches and church goers; lumping all church goers into one category calling them hypocrites is wrong. There are many good churches and church goers who help people and their communities.
- RobotToaster 6mo ago"You're trying to kidnap what I've rightfully stolen!"
- gib444 6mo agoAnd have absolutely no reservations about making such an obvious statement on a public forum
- DrinkyBird 6mo agoIt’s funny because the first AI scraper I remember blocking was from OpenAI’s, as it got stuck in a loop somehow and was impacting the performance of a wiki I run. All to violate every clause of the CC BY-NC-SA license of the content it was scraping :)
- jordanb 6mo agoThey don't want anyone to take that which they have rightfully stolen.
- splatter9859 6mo agoExactly! How dare you have access to their stolen content in the midst of them doing the same.
- altmanaltman 6mo agoWell at least they have 1 person working on "Integrity" so can't be too bad
- rsrsrs86 6mo agoThis
- ValveFan6969 6mo ago[dead]