11 ms·
I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how
by Octokiddie 2y ago
I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how much could that possibly bring in?
- gtirloni 2y agoIt'd say it's more like a honeypot for bots. So pretty similar objectives.
- Octokiddie 2y agoSo it served its purpose by trapping the OpenAI spider? If so, why post that message? As a flex?
- Takennickname 2y agoIt's a honeypot. He's telling people openai doesn't respect robots.txt and just scrapes whatever the hell it wants.
- cwillu 2y agoExcept the first thing openai does is read robots.txt. However, robots.txt doesn't cover multiple domains, and every link that's being crawled is to a new domain, which requires a new read of a robots .txt on the new domain.
- queuebert 2y agoDid we just figure out a DoS attack for AGI training? How large can a robots.txt file be?
- a_c 2y agoWhat about making it slow? One byte at a time for example while keeping the connection open
- happymellon 2y agoA slow stream that never ends?
- SteveNuts 2y agoThis would be considered a Slow Loris attack, and I'm actually curious how scrapers would handle it. I'm sure the big players like Google would deal with it gracefully.
- throw_a_grenade 2y agoYou just set limits on everything (time, buffers, ...), which is easier said than done. You need to really understand your libraries and all the layers down to the OS, because its enough to have one abstraction that doesn't support setting limits and it's an invitation for (counter-)abuse.
- starttoaster 2y agoDoesn't seem like it should be all that complex to me assuming the crawler is written in a common programming language. It's a pretty common coding pattern for functions that make HTTP requests to set a timeout for requests made by your HTTP client. I believe the stdlib HTTP library in the language I usually write in actually sets a default timeout if I forget to set one.
- Calzifer 2y agoThose are usually connection and no-data timeouts. A total time limit is in my experience less common.
- gtirloni 2y ago
- everforward 2y agoNo, because there’s no legal weight behind robots.txt. The second someone weaponizes robots.txt all the scrapers will just start ignoring it.
- Retric 2y agoThat’s how you weaponize it. Set things up to give endless/randomized/poisoned data to anybody that ignores robots.txt.
- everforward 2y agoYou mean human users? That is and always will be the dominant group of clients that ignore robots.txt. What you’re talking about is an arms race wherein bots try to mimic human users and sites try to ban the bots without also banning all their human users. That’s not a fight you want to pick when one of the bot authors also owns the browser that 63% of your users use, and the dominant site analytics platform. They have terabytes of data to use to train a crawler to act like a human, and they can change Chrome to make normal users act like their crawler (or their crawler act more like a Chrome user). Shit, if Google wanted, they could probably get their scrapes directly from Chrome and get rid of the scraper entirely. It wouldn’t be without consequence, but they could.
- Retric 2y agoIt’s fairly trivial to treat Google’s crawler differently if you want. https://developers.google.com/search/docs/crawling-indexing/verifying-googlebot https://developers.google.com/search/docs/crawling-indexing/... The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free. People don’t mindlessly click on things like terms of service crawlers are quite dumb. Little need for an arms race, as the people running these crawlers rarely put much effort into any one source.
- everforward 2y ago
- flutas 2y ago> Except the fiist thing openai does is read robots.txt. Then they should see the "Disallow: /" line, which means they shouldn't crawl any links on the page (because even the homepage is disallowed). Which means they wouldn't follow any of the links to other subdomains.
- niutech 2y agoThis robots.txt has Disallow rule commented out: # buzz off #User-agent: GPTBot #Disallow: /
- darkwater 2y agoAnd they do have (the same) robots.txt on every domain, tailored for GPTbot, i.e. https://petra-cody-carlene.web.sp.am/robots.txt https://petra-cody-carlene.web.sp.am/robots.txt So, GPTBot is not following robots.txt, apparently.
- fsckboy 2y agohumans don't read/respect robots.txt, so in order to pass the Turing test, ai's need to mimic human behavior.
- gunapologist99 2y agoThis must be why self-driving cars always ignore the speed limit. ;)
- microtherion 2y agoMore directly, e.g. Tesla boasts of training their FSD on data captured from their customer's unassisted driving. So it's hardly surprising that it imitates a lot of humans' bad habits, e.g. rolling past stop lines.
- roughly 2y agoJesus, that’s one of those ideas that looks good to an engineer but is why you really need to hire someone with a social sciences background (sociology, anthropology, psychology, literally anyone who’s work includes humans), and probably should hire two, so the second one can tell you why the first died of an aneurism after you explained your idea.
- yreg 2y agoAI DRIVR claims that beta V12 is much better precisely because it takes rules less literally and drives more naturally.
- cwillu 2y agoAccessing a directly referenced page is common in order to receive the noindex header and/or meta tag, whose semantics are not implied by “Disallow: /” And then all the links are to external domains, which aren't subject to the first site's robots.txt
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- Takennickname 2y ago> Except the first thing openai does is read robots.txt. What good is reading it if it doesn't respect it
- GaggiX 2y agoIt seems to respect it as the majority of the requests are for the robots.txt.
- flutas 2y agoHe says 3 million, and 1.8 million are for robots.txt So 1.2 million non robots.txt requests, when his robots.txt file is configured as follows # buzz off User-agent: GPTBot Disallow: / Theoretically if they were actually respecting robots.txt they wouldn't crawl any pages on the site. Which would also mean they wouldn't be following any links... aka not finding the N subdomains.
- swyx 2y agofor the 1.2 million are there other links he's not telling us about?
- flutas 2y agoI'm assuming those are homepage requests for the subdomains.
- otherme123 2y agoA lot of crawlers, if not all, have a policy like "if you disallow our robot, it might take a day or two before it notices". They surely follow the path "check if we have robots.txt that allows us to scan this site, if we don't get and store robots.txt, scan at least the root of the site and its links". There won't be a second scan, and they consider that they are respecting robots.txt. Kind of "better ask for forgiveness than for permission".
- jeremyjh 2y agoThat is indistinguishable from not respecting robots.txt. There is a robots.txt on the root the first time they ask for it, and they read the page and follow its links regardless.
- dspillett 2y agoSo, it has worked…
- madkangas 2y agoI recognize the name John Levine at iecc.com, "Invincible Electric Calculator Company," from web 1.0 era. He was the moderator of the Usenet comp.compilers newsgroup and wrote the first C compiler for the IBM PC RT https://compilers.iecc.com/ https://compilers.iecc.com/
- throw_a_grenade 2y agoThis is honeypot. The author, https://en.wikipedia.org/wiki/John_R._Levine https://en.wikipedia.org/wiki/John_R._Levine, keeps it just to notice any new (significant) scraping operation launched that will invariably hit his little farm and let be seen in the logs. He's well known anti-spam operative with his various efforts now dating back multiple decades. Notice how he casually drops a link to the landing page in the NANOG message. That's how the bots will get a bait.
- agilob 2y agoIt's for shits-and-giggles and it's doing its job really well right now. Not everything needs to have an economic purpose, 100 trackers, ads and backed by a company.
- schleck8 2y agoThe books on there are affiliate links I think.
- pflanze 2y agoLinkers & Loaders is their own book (I haven't checked the others). They have a page at https://www.iecc.com/linker/ https://www.iecc.com/linker/ where they used to publish a draft of the book contents, but changed the page to say "Chapters were available in an excessive variety of formats, but are not any longer due to chronic piracy", when it got posted to HN at https://news.ycombinator.com/item?id=18424233 https://news.ycombinator.com/item?id=18424233 and I bundled the files for offline reading. I notified them via email about that asking if they are OK with it but got an unfriendly response that I pirated the files and that wasn't OK, so I took the link down again and they changed that text. (Shrug. I'm not a/the book author, they are. I'll say that I also suggested to them they ask on the page not to do what I did since then I wouldn't have, but they chose their more radical approach.)