4 ms·
> Layer 7 just means they are making ton of HTTP requests from lot of IPS however this is the nature of a web analytics. Exactly! Which is why it was so hard.
by JackWritesCode 6y ago
> Layer 7 just means they are making ton of HTTP requests from lot of IPS however this is the nature of a web analytics.
Exactly! Which is why it was so hard. There's no path pattern. Everything hits "/". So the only way to fight back is to match IP / header patterns (but even then, we have to redact sensitive headers).
> How can the attackers know which websites have your analytics on? Crawling the web is super hard.
The attacker went after some of our more high profile customers. They're known via testimonials or from Twitter.
> The access log should store this `sid` allowing you to just block the account. (or ask them for ton of money if it's a legitimate use :) )
We could certainly temporarily block traffic to a site. The problem is, without some kind of firewall (e.g. WAF), our application has to absorb so much traffic, and that's the issue. We need to block it at the edge.
- throwaway189262 6y ago> The problem is, without some kind of firewall (e.g. WAF), our application has to absorb so much traffic, and that's the issue. We need to block it at the edge. I think you need some high performance ingestion code. One day somebody is going to load test or misconfigure their site. Or you'll get a really big client. Or somebody will get Reddit frontpaged or slashdotted. Blocking huge volumes of traffic isn't always going to be an option. You should be able to handle it. I would get off all this "sexy" stuff like lamdba and SQS and build some old school battle boxes at a co-lo. Run some extreme performance framework like Actix or Vert.X built to handle giant piles of traffic. A single box with those frameworks and a 40 gig line can handle millions of requests a second. Use counting bloom filters synced between boxes for IP blocking. This avoids saving IP's and is needed to prevent ram exhaustion attacks anyways. Use two layers, one for individual IP's and a second for blocking subnets with lots of bad actors in them. Block for ~24h by using dual sets of bloom filters in a sliding window with 12h overlap where IP's get added to both. When a filter is 24h old, discard it. This needs to be done because you can't remove items from Bloom filters, they saturate over time. On these super boxes you can do filter, blocking, batching, and dedup. Then send the data off to wherever for further processing. You mentioned not wanting to use Cloudflare because they're a competitor. That's fine but I would take a look at their blog. They go over the tech they use for mass data ingestion and filtering, and it's basically what I just described.