4 ms·
What's the intended harm that restricting people from linking to uncontrolled 3rd party assets would prevent/mitigate? Consider: You're CompanyX, and I'm Dodgy
by shabble 8y ago
What's the intended harm that restricting people from linking to uncontrolled 3rd party assets would prevent/mitigate?
Consider: You're CompanyX, and I'm DodgyFontHost.tld
My business model is exploiting and selling as much data as I can gather/mine from my traffic.
You embed (that is, reference/hotlink) some of my fonts on your pages.
If a user visits your page, and as a result makes a request to me for a font, I can log everything about that request, but I don't (afaik) have much/any additional knowledge that makes it particularly useful.
Assume there's no ?UTM=... tracking content in the url itself, you're just referencing a static font file.
I'm not sure offhand if browsers would be passing a referer header by default, or if that could somehow reliably identify the site I'm actually visiting. If so, that'd be one valuable fact.
I might be able to fingerprint the users browser from other headers or their OS from network-level quirks.
Anything else I'm missing?
I feel like 'IP $x made a request for $file' isn't the important thing to be looking at here, it's what I can learn from other things associated with the request that I can exploit.
But yes, if you had a reliable lookup from (ip,timestamp) to legal person, then it's absolutely Personally Identifiable.
Imagine if every browser set a valid, correct 'X-Requestors-Legal-Name: Bob Smith, Sometown, USA' header on every request. That's obviously identifiable. Adding a layer of indirection doesn't make it less so, although it does maybe place it on a continuum of 'cost/effort to identify based on this info'.
It ranges from 'trivial, because it's right there in the content you're sending', through 'not directly, but easily enough via subscriptions to one or more commercial data providers' to 'if someone steals our data and combines it with stolen data from several other sources, they have a non-zero chance of guessing your identity correctly'.
- clarry 8y ago> I'm not sure offhand if browsers would be passing a referer header by default, or if that could somehow reliably identify the site I'm actually visiting. If so, that'd be one valuable fact. They are passing referer unless the context is an encrypted connection and the resource is on plain HTTP. Firefox also strips out the path from the URL for third party requests, but only in private browsing mode: https://blog.mozilla.org/security/2018/01/31/preventing-data-leaks-by-stripping-path-information-in-http-referrers/ https://blog.mozilla.org/security/2018/01/31/preventing-data... I think this should be the default for all third party domains no matter what the mode. (Really, I'd rather see that header just go away.)
- shabble 8y agoThanks. I did a bit of digging around after posting that and found roughly what you describe, that the Referer: is a valuable datapoint, and should probably be a bit more selective. I suspect it's sufficiently ingrained in existing apps to make it hard to deprecate completely, but something like the path stripping might be a decent compromise. For cross-origin requests I think there's also a mandatory 'Origin:' header that would identify at least the domain (but not path) a user request was referenced from. I used to use a firefox addon called RefControl but IIRC it was a casualty of the quantum/webextensions transition. uMatrix has a basic referer spoofing capability, but it's all or nothing for a particular site/scope.