6 ms·
Awesome, I'd love some advice if you're willing. So here's our situation. * We're getting hit with a huge DDoS attack, repeatedly over 3 weeks, with no sign of
by JackWritesCode 6y ago
Awesome, I'd love some advice if you're willing. So here's our situation.
* We're getting hit with a huge DDoS attack, repeatedly over 3 weeks, with no sign of stopping
* With zero access logs, there was no way to find patterns in the attack, and we had no way to block it
* Our service was going offline during these attacks
* We introduced access logs that are auto-deleted after 24 hours. We redacted all information about the site/page/activity etc. but keep IP & User Agent for pattern matching
* We were then able to identify a pattern and block the attack on Saturday
* Without access logs (even redacted ones), this wasn't possible
I was hoping a more senior engineer on Hacker News would comment and I can't wait to hear how you'd do it. I have no experience in DDoS protection at all, and this seems like the only possible way. Even rate limiting requires storing IP addresses. But if you know a more privacy-focused way to block these attacks, I'm sure I'll buy you a few beers when we hang out.
- ve55 6y agoI would probably start by looking for third-parties to help manage to problems, at least unless they all end up at the application layer. Throw it beyond Cloudflare and talk to them, and if they're not the right fit, try someone else.
- tpetry 6y agoThere is nothing wrong with this approach but exactly this mentality lead to the actual centralized internet where only a few peers handle most of the traffic.
- JackWritesCode 6y agoThey're still going to need to process IP addresses or some kind of PII though, that's the thing :(
- ve55 6y agoI see, I had misunderstood the extent to which they wanted to remain privacy-focused. Definitely can be a lot more work then when you can't outsource that stuff.
- hartator 6y ago> * We're getting hit with a huge DDoS attack, repeatedly over 3 weeks, with no sign of stopping I've read the full blog post, I am not convinced it's a DDoS attack. Traffic patterns for web analytics will come from over the place and will look like a DDoS when it's not. For example, a customer misplacing their analytics in a JS loop and having a moderate traffic blog will generate billions of requests from all over the globe. Event tracking can be billions of requests as well by themselves. Never attribute to malice that which is adequately explained by simpler means. > * With zero access logs, there was no way to find patterns in the attack, and we had no way to block it It's a loss of time and energy. Your system should be able to handle these billions of requests. > * We were then able to identify a pattern and block the attack on Saturday Was it specific accounts? > * Without access logs (even redacted ones), this wasn't possible You can one-way hash the IP. So you can still look for pattern but you've lost the actual IP. And same one-way hash IP can block whichever IP seems devious in your firewall. (like md5 can be enough.) > But if you know a more privacy-focused way to block these attacks, I'm sure I'll buy you a few beers when we hang out. Haha no need to. But come say hi if you ever in Austin. julien _at_ serpapi.com.
- JackWritesCode 6y agoIt was 100% a layer 7 DDoS, the attack was targeted and malicious. I can't say too much but the AWS team confirmed it. But I appreciate how something like you describe could happen. I like the idea of an MD5 hash. Although I'm not too certain why an IP would be a bad thing to log for 24 hours. From a privacy law perspective, the MD5 hash is considered PII. And if we see an IP address in an access log, we know that an IP visited one of the tens of thousands of websites Fathom runs on, but we don't know which one. Edit: Something else, with MD5 there's no way of finding patterns with the IPs, so you'd have to play whack a mole. Whereas raw IPs allow no real privacy invasion whilst allowing pattern detection
- hartator 6y ago> It was 100% a layer 7 DDoS, the attack was targeted and malicious. I can't say too much but the AWS team confirmed it. But I appreciate how something like you describe could happen. Layer 7 just means they are making ton of HTTP requests from lot of IPS however this is the nature of a web analytics. > we know that an IP visited one of the tens of thousands of websites Fathom runs on, but we don't know which one. How can the attackers know which websites have your analytics on? Crawling the web is super hard. You should be able to get back which accounts are affected from your analytics call: https://starman.fathomdns.com/?p=%2F&h=https%3A%2F%2Fusefathom.com&r=&sid=BIABKBRK&res=1440x900 Isn't sid=BIABKBRK the account? The access log should store this `sid` allowing you to just block the account. (or ask them for ton of money if it's a legitimate use :) )
- throwaway189262 6y agoAPI spam is a big problem in Google Analytics too, you might just be big enough now to have become a target. I think the answer is to build a spam filter. You could take the ones made for email and adapt them for analytics requests possibly. This seems like a real place where ML pattern recognition models would shine. Any realistic spam filter will involve sampling a fraction of traffic for deep analysis and IP banning badly behaving subnets for ~24h. And it's definitely better to shadow ban, but smart assholes will have an account to see if their spam is getting through as well. I wouldn't worry about short-term IP caching. AWS's upstream load balancers and your own servers are probably doing it anyways to maintain TCP state tables. Linux kernel's "conntrack". If you don't want to cache IP's you can you a probabilistic data structure like a bloom filter synched between instances. If has a small false positive rate but is very fast and doesn't store whole IP. Bloom filter based IP filtering is used in every big DDOS prevention system I know of. As much as you like lambda, I would ditch it. And the queues. My general advice is that any time you need to add a work queue to something, it's not fast enough. Your analytics endpoint data ingestion should be something lightning fast like Go, Rust, or an async Java back end. Analytics is a lossy process, you lose traces because of browser behavior and plugins all the time anyways so I wouldn't prioritize 100% accuracy. I would focus on power/dollar over reliability. If I was you, my ingest boxes would be load balanced with DNS round robin and sitting at various Colocation providers. Get a fat 40 gig unlimited data pipe. Build some stupid fast Rust/Go/Java backend that can saturate that pipe. And do all your filtering/spam analysis here. I don't think lambda, SQS queues, PHP are the best technologies for this kind of mass data ingestion. I don't even think your ingestion layer should be on AWS. I would follow the lead of other companies doing mass data ingestion and build your own machines. That's how CloudFlare, Netflix etc are able to handle so much traffic without going bankrupt. I would consider yourself lucky that your first DDOS was so small. 10k requests/sec is tiny. ~400k/sec can be generated on a regular desktop with fiber internet connection. Right now, a single user could knock you offline by messing around with JMeter. I think it's a wake up call that you're in the infrastructure business whether you like it or not, and you need to massively beef up your data ingestion layer. Realistically, you should be able to handle ~50 gigabit attack with 10 million requests/sec. I think that's achievable with a couple boxes colocated on 40 gig lines running fast software
- 6y ago