3 ms·
Ask HN: The amount of AI bot traffic is out control?
We host a web app which has around 50,000 statically generated public pages, and the amount of bot traffic is insane. Facebook’s crawler requested the same pages over 3 million times in 2 days, and we’ve used Vercel and Cloudflare to block as much as possible but it’s not working. The AI scrapers are routing requests through residential proxies, and even after putting our entire app behind an Auth gate with cloudflare turnstile they’re creating accounts to get access to data that’s meant for humans to read / use. I’ve never seen anything like it before, our hosting costs are up like crazy, and besides all of this, I can’t trust anything I read or see on this web anymore. As a last resort we’ll try switching to hcaptcha tomorrow instead of turnstile which appears to be more difficult but this is all so depressing to me. Want to hear if anyone else is experiencing this, there’s no way I’m the only here who is struggling with sophisticated bots?
- smallerize 19d agoHow do you know it's facebook's crawler if it going through residential proxies?
- flyingcoder 19d agoIt had meta in the user-agent headers, it’s trivial to spoof but I’ve read of many others who are seeing the same thing. It’s not just them though, it’s also parallel web systems, and then a bunch of others who aren’t identifying themselves.
- d3Xt3r 19d agoThat's probably not Meta themselves, but their new agent called Muse. It's likely coming from actual individuals. Not sure why they're scraping your website though. Maybe time to turn off all the SEO stuff to make your site less discoverable?
- avinash147 19d ago[flagged]
- flyingcoder 18d agoThank you, this is all really helpful
- babyDoomer 19d agoI doubt it's really Facebook creating fake accounts and using proxies, why would they keep their real UA then? Anyhow, I occasionally suffer this same problem when happy crawlers find my site and use truesign.ai to block them, so far successfully. It detects residential proxies, fake emails and overall suspicious activity, and - something I really wanted to avoid - there are no captchas.
- verdverm 19d ago> The AI scrapers are routing requests through residential proxies How do you differentiate between an Ai scraper and an end user having an agent perform a search and then fetch all the results to analyze them for relevance? (because search has become so bad I need an agent to deal with it before I look at things)
- flyingcoder 19d agoBecause the sheer volume is insane, and when looking through the logs the requests for a short term IP do not make sense / are not sequential, hope that makes sense
- verdverm 19d ago> Because the sheer volume is insane It sounds like you are reading tea leaves (based on your other comments) and volume alone is not an indicator to the source. My personal web page access has gone up 10x because I have agents doing things, and they are sloppy af, fetching way more than they need to. Insane volume from individuals, by way of their personal agents, is not surprising to me, given what I have at my hands and what I know others have and do in other harnesses. Them not being sequential, or typical scraping patterns, would to me lend credence towards people using deep research agents because google search has become so bad. In other words, the SEO industry went too hard, helped ruin search result quality, and now we the people are using Ai to sift through the noise for what's actually useful, work around all that money that manipulates what we see. For me, what used to be a single google search with useful page snippets is now an agent running multiple queries across multiple SERPs and fetching dozens or >100 pages.
- protoduction 18d agoAt Friendly Captcha [1] (I'm a co-founder) we see this more and more as well, with residential proxies being commonly used now relying on IP reputation alone is (no longer) enough. The main usecase of our captcha was protecting web actions (i.e. stuff you do on a page, like creating an account, making a query or logging in), but some our customers already built a setup where our captcha is part of a system that sits "in front of" the website to tackle this kind of abuse. We're launching a product called Friendly Guard [2] to make this more straightforward this year. If you're up for it you can join our beta to try it out before it goes live, we'd love to learn from the situation. [1]: https://friendlycaptcha.com https://friendlycaptcha.com [2]: https://developer.friendlycaptcha.com/docs/v2/friendly-guard/ https://developer.friendlycaptcha.com/docs/v2/friendly-guard...