2 ms·
One aspect I haven't seen mentioned yet is that there is a geographic dimension to what ends up in search engine caches; their crawlers end up talking to the ne
by markonen 10y ago
One aspect I haven't seen mentioned yet is that there is a geographic dimension to what ends up in search engine caches; their crawlers end up talking to the nearest Cloudflare point-of-presence so any leaked data is also going to be from requests served in that same general area.
For example, my Cloudflare-using site is in the 1B to 10B requests a month bracket, meaning 112–1,118 anticipated leaks per Cloudflare's calculations. It's enough that you'd expect to possibly find some in the search engines, but we haven't.
One potential explanation is that our usage is heavily concentrated within a geographic area (the Nordics), and this just isn't where the big search engines are doing their web crawling from.
- eastdakota 10y agoYes, that's correct and not something our analysis has taken into account. Search crawlers are more likely to hit certain PoPs. If your content was less likely to be in one of those PoPs then you're less likely to have leaked data to one of them. We also didn't take into account that the distribution of the ~6,500 sites that triggered the bug weren't even across our infrastructure. We distribute load for any site across some fraction of all the servers in any given PoP. Based on whether you're on a cluster of servers with more or fewer vulnerable sites you'd have been more or less likely to leak data. For the general case the stats are directionally correct and, if anything, conservative (we believe, if anything, overstate the risk). But a particular customer's situation could deviate from the expected probabilities for a number of reasons.