22 ms·
One of my websites was absolutely destroyed by Meta's AI bot: Meta-ExternalAgent https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/#identify-
by markerz 2y ago
One of my websites was absolutely destroyed by Meta's AI bot: Meta-ExternalAgent https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/#identify-2 https://developers.facebook.com/docs/sharing/webmasters/web-...
It seems a bit naive for some reason and doesn't do performance back-off the way I would expect from Google Bot. It just kept repeatedly requesting more and more until my server crashed, then it would back off for a minute and then request more again.
My solution was to add a Cloudflare rule to block requests from their User-Agent. I also added more nofollow rules to links and a robots.txt but those are just suggestions and some bots seem to ignore them.
Cloudflare also has a feature to block known AI bots and even suspected AI bots: https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/ https://blog.cloudflare.com/declaring-your-aindependence-blo... As much as I dislike Cloudflare centralization, this was a super convenient feature.
- jandrese 2y agoIf a bot ignores robots.txt that's a paddlin'. Right to the blacklist.
- deleted 2y ago[deleted]
- nabla9 2y agoThe linked article explains what happens when you block their IP.
- gs17 2y agoFor reference: > If you try to rate-limit them, they’ll just switch to other IPs all the time. If you try to block them by User Agent string, they’ll just switch to a non-bot UA string (no, really). It's really absurd that they seem to think this is acceptable.
- viraptor 2y agoBlock the whole ASN in that case.
- therealdrag0 2y agoWhat about adding fake sleeps?
- CoastalCoder 2y agoI wonder if it would work to send Meta's legal department a notice that they are not permitted to access your website. Would that make subsequent accesses be violations of the U.S.'s Computer Fraud and Abuse Act?
- jahewson 2y agoNo, fortunately random hosts on the internet don’t get to write a letter and make something a crime.
- throwaway_fai 2y agoUnless they're a big company in which case they can DMCA anything they want, and they get the benefit of the doubt.
- BehindBlueEyes 2y agoCan you even DMCS takedown crawlers?
- throwaway_fai 2y agoDoubt it, a vanilla cease-and-desist letter would probably be the approach there. I doubt any large AI company would pay attention though, since, even if they're in the wrong, they can outspend almost anyone in court.
- Nevermark 2y agoSmall claims court?
- betaby 2y agoCrashing wasn't the intent. And scraping is legal, as I remember per Linkedin case.
- 2y ago
- coldpie 2y agoImagine being one of the monsters who works at Facebook and thinking you're not one of the evil ones.
- throwaway_fai 2y ago[flagged]
- FrustratedMonky 2y agoThe Banality of Evil. Everyone has to pay bills, and satisfy the boss.
- throwaway290 2y agoOr ClosedAI. Related https://news.ycombinator.com/item?id=42540862 https://news.ycombinator.com/item?id=42540862
- Aeolun 2y agoWell, Facebook actually releases their models instead of seeking rent off them, so I’m sort of inclined to say Facebook is one of the less evil ones.
- echelon 2y ago> releases their models Some of them, and initially only by accident. And without the ingredients to create your own. Meta is trying to kill OpenAI and any new FAANG contenders. They'll commoditize their complement until the earth is thoroughly salted, and emerge as one of the leading players in the space due to their data, talent, and platform incumbency. They're one of the distribution networks for AI, so they're going to win even by just treading water. I'm glad Meta is releasing models, but don't ascribe their position as one entirely motivated by good will. They want to win.
- int_19h 2y agoFWIW, there's considerable doubt that the initial LLaMA "leak" was accidental, based on Meta's subsequent reaction. I mean, the comment with a direct download link in their GitHub repo stayed up even despite all the visibility (it had tons of upvotes).
- devit 2y ago[flagged]
- jsheard 2y agoThat's right, getting DDOSed is a skill issue. Just have infinite capacity.
- devit 2y agoDDOS is different from crashing. And I doubt Facebook implemented something that actually saturates the network, usually a scraper implements a limit on concurrent connections and often also a delay between connections (e.g. max 10 concurrent, 100ms delay). Chances are the website operator implemented a webserver with terrible RAM efficiency that runs out of RAM and crashes after 10 concurrent requests, or that saturates the CPU from simple requests, or something like that.
- adamtulinius 2y agoYou can doubt all you want, but none of us really know, so maybe you could consider interpreting people's posts a bit more generously in 2025.
- atomt 2y agoI've seen concurrency in excess of 500 from Metas crawlers to a single site. That site had just moved all their images so all the requests hit the "pretty url" rewrite into a slow dynamic request handler. It did not go very well.
- adamtulinius 2y agoNo normal person has a chance against the capacity of a company like Facebook
- Aeolun 2y agoAnyone can send 10k concurrent requests with no more than their mobile phone.
- bodantogat 2y agoI see a lot of traffic I can tell are bots based on the URL patterns they access. They do not include the "bot" user agent, and often use residential IP pools. I haven't found an easy way to block them. They nearly took out my site a few days ago too.
- newsclues 2y agoThe amateurs at home are going to give the big companies what they want: an excuse for government regulation.
- throwaway290 2y agoIf it doesn't say it's a bot and it doesn't come from a corporate IP it doesn't mean it's NOT a bot and not run by some "AI" company.
- bodantogat 2y agoI have no way to verify this, I suspect these are either stealth AI companies or data collectors, who hope to sell training data to them
- datadrivenangel 2y agoI've heard that some mobile SDKs / Apps earn extra revenue by providing an IP address for VPN connections / scraping.
- odo1242 2y agoChrome extensions too
- int_19h 2y agoDon't worry, the governments are perfectly capable of coming up with excuses all on their own.
- 2y ago
- MetaWhirledPeas 2y ago> Cloudflare also has a feature to block known AI bots and even suspected AI bots In addition to other crushing internet risks, add wrongly blacklisted as a bot to the list.
- throwaway290 2y agoWhat do you mean crushing risk? Just solve these 12 puzzles by moving tiny icons on tiny canvas while on the phone and you are in the clear for a couple more hours!
- gs17 2y agoIf it clears you at all. I accidentally set a user agent switcher on for every site instead of the one I needed it for, and Cloudflare would give me an infinite loop of challenges. At least turning it off let me use the Internet again.
- homebrewer 2y agoIf you live in a region which it is economically acceptable to ignore the existence of (I do), you sometimes get blocked by website r̶a̶c̶k̶e̶t̶ protection for no reason at all, simply because some "AI" model saw a request coming from an unusual place.
- benhurmarcel 2y agoSometimes it doesn’t even give you a Captcha. I have come across some websites that block me using Cloudflare with no way of solving it. I’m not sure why, I’m in a large first-world country, I tried a stock iPhone and a stock Windows PC, no VPN or anything. That’s just no way to know.
- dannyw 2y agoThat’s probably a page/site rule set by the website owner. Some sites block EU IPs as the costs of complying with GDPR outweigh the gain.
- petee 2y agoSilly question, but did you try to email Meta? Theres an address at the bottom of that page to contact with concerns. > webmasters@meta.com I'm not naive enough to think something would definitely come of it, but it could just be a misconfiguration
- TuringNYC 2y ago>> One of my websites was absolutely destroyed by Meta's AI bot: Meta-ExternalAgent https://developers.facebook.com/docs/sharing/webmasters/web- https://developers.facebook.com/docs/sharing/webmasters/web-... Are they not respecting robots.txt?
- eesmith 2y agoQuoting the top-level link to geraspora.de: > Oh, and of course, they don’t just crawl a page once and then move on. Oh, no, they come back every 6 hours because lol why not. They also don’t give a single flying fuck about robots.txt, because why should they. And the best thing of all: they crawl the stupidest pages possible. Recently, both ChatGPT and Amazon were - at the same time - crawling the entire edit history of the wiki.
- candlemas 2y agoThe biggest offenders for my website have always been from China.
- tonyedgecombe 2y ago[flagged]
- viraptor 2y agoYou can also block by IP. Facebook traffic comes from a single ASN and you can kill it all in one go, even before user agent is known. The only thing this potentially affects that I know of is getting the social card for your site.
- ryandrake 2y ago> My solution was to add a Cloudflare rule to block requests from their User-Agent. Surely if you can block their specific User-Agent, you could also redirect their User-Agent to goatse or something. Give em what they deserve.
- globalnode 2y agocant you just mess with them? like accept the connection but send back rubbish data at like 1 bps?
- EVa5I7bHFq9mnYK 2y agoYeah, super convenient, now every second web site blocks me as "suspected AI bot".
- PeterStuer 2y agoMost administrators have no idea or no desire to correctly configure Cloudflare, so they just slap it on the whole site by default and block all the legitimate access to e.g. rss feeds.