6 ms·
Did we just figure out a DoS attack for AGI training? How large can a robots.txt file be?
by queuebert 2y ago
Did we just figure out a DoS attack for AGI training? How large can a robots.txt file be?
- a_c 2y agoWhat about making it slow? One byte at a time for example while keeping the connection open
- happymellon 2y agoA slow stream that never ends?
- SteveNuts 2y agoThis would be considered a Slow Loris attack, and I'm actually curious how scrapers would handle it. I'm sure the big players like Google would deal with it gracefully.
- throw_a_grenade 2y agoYou just set limits on everything (time, buffers, ...), which is easier said than done. You need to really understand your libraries and all the layers down to the OS, because its enough to have one abstraction that doesn't support setting limits and it's an invitation for (counter-)abuse.
- starttoaster 2y agoDoesn't seem like it should be all that complex to me assuming the crawler is written in a common programming language. It's a pretty common coding pattern for functions that make HTTP requests to set a timeout for requests made by your HTTP client. I believe the stdlib HTTP library in the language I usually write in actually sets a default timeout if I forget to set one.
- Calzifer 2y agoThose are usually connection and no-data timeouts. A total time limit is in my experience less common.
- gtirloni 2y agoHere you go (1 req/min, 10 bytes/sec), please report results :) http { limit_req_zone $binary_remote_addr zone=ten_bytes_per_second:10m rate=1r/m; server { location / { if ($http_user_agent = "mimo") { limit_req zone=ten_bytes_per_second burst=5; limit_rate 10; } } } }
- beau_g 2y agoScrapers of the future won't be ifElse logic, they will be LLM agents themselves. The slow loris robots.txt has to provide an interface to it's own LLM, which engages the scraper LLM in conversation, aiming to extend it as long as possible. "OK I will tell you whether or not I can be scraped. BUT FIRST, listen to this offer. I can give you TWO SCRAPES instead of one, if you can solve this riddle."
- iosguyryan 2y agoCan I interest you in a scrape-share with Claude?
- reasonabl_human 2y agoSolid use case for Saul Goodman LLM alignment
- Phelinofist 2y agoSounds like endlessh
- bityard 2y agoThat would make it a tarpit, a very old technique to combat scrapers/scanners
- everforward 2y agoNo, because there’s no legal weight behind robots.txt. The second someone weaponizes robots.txt all the scrapers will just start ignoring it.
- Retric 2y agoThat’s how you weaponize it. Set things up to give endless/randomized/poisoned data to anybody that ignores robots.txt.
- everforward 2y agoYou mean human users? That is and always will be the dominant group of clients that ignore robots.txt. What you’re talking about is an arms race wherein bots try to mimic human users and sites try to ban the bots without also banning all their human users. That’s not a fight you want to pick when one of the bot authors also owns the browser that 63% of your users use, and the dominant site analytics platform. They have terabytes of data to use to train a crawler to act like a human, and they can change Chrome to make normal users act like their crawler (or their crawler act more like a Chrome user). Shit, if Google wanted, they could probably get their scrapes directly from Chrome and get rid of the scraper entirely. It wouldn’t be without consequence, but they could.
- Retric 2y agoIt’s fairly trivial to treat Google’s crawler differently if you want. https://developers.google.com/search/docs/crawling-indexing/verifying-googlebot https://developers.google.com/search/docs/crawling-indexing/... The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free. People don’t mindlessly click on things like terms of service crawlers are quite dumb. Little need for an arms race, as the people running these crawlers rarely put much effort into any one source.
- everforward 2y ago