11 ms·
We've been toying with this idea in earlier revisions of the spec, basically using the hash as a cache key and not loading the same file from websiteB if it has
by bugmen0t 11y ago
We've been toying with this idea in earlier revisions of the spec, basically using the hash as a cache key and not loading the same file from websiteB if it has already been loaded form websiteA.
Unfortunately, this could be used as a cache poisoning attack to bypass Content Security Policy.
See the section about "Content addressable storage" at <https://frederik-braun.com/subresource-integrity.html https://frederik-braun.com/subresource-integrity.html>.
(If you can come up with a magical solution to this problem, join the W3C web application security group mailing list and send us an email.)
- deleted 11y ago[deleted]
- benten10 11y agoHow about this: There's a warehouse 'owner' internal to the browser (and not exposed to pages/extension), who 'remembers' the resources the browser has, and the times it took to access them (for commonly-accessed resources). When a page requests the cached resource, the owner 'returns' the resource with a delay of whatever the original access time was, fudged around by some noise. A weakness of this would be that websites would be able to 'communicate' with each other by engineering response times to your browser, and then checking how long it takes your browser to access that. But this is a different scenario than random websites trying to figure out where else you've been: here the pages need to be in collusion with each other. The servers might try to use fancy algorithms to try to figure out if you're using cached versions by hitting different distributed servers and figuring out if the resource load time is an outlier. But that's prone to a lot of noise and other issues, and lesser of an issue than the original concern. Right? This obviously won't help with page load speeds, but will help network load for bot users and servers. One possible issue might be: if you've been in a slow connection previously, all your connections will seem slow even after you are in a faster connection. For that, you can just purge the cache and force the browser to reload the resources. Edited for formatting.
- sbierwagen 11y agoThe danger is not that malicious websites figure out where you've been, the danger is that a malicious website could poison the cache for a critical piece of JS used in, for example, gmail. Visit a malicious site, boom, Russians can read your email.
- pdkl95 11y agoThis would require generating collisions for hash (e.g. sha385). We can trust the hash because SR?I already assumes the hash function works in the integrity="" attribute.
- nitrogen 11y agoI think both of these things (history snooping, XSS), plus the Dropbox problem of injecting a hash without ever actually having the file, will need to be addressed.
- jewel 11y agoHow about an additional attribute named "global", "shared", "public", "use-global-cache", share-with="*", etc. that the developer can use to opt in to the behavior? A site operator would only opt in to the behavior for assets that are not unique to the site. A second idea would be to wait until several unique domains had requested the asset before turning on the behavior for that asset. (By unique domain I specifically mean the part of the domain that's written in black text in the URL bar, excluding subdomains that are in gray.) These are two easiest-to-implement solutions I can think of.
- sbierwagen 11y agoHow about an additional attribute named "global", "shared", "public", "use-global-cache", share-with="*", etc. that the developer can use to opt in to the behavior? Allowing people to opt-in to a cache poisoning vector seems like a bad idea. A second idea would be to wait until several unique domains had requested the asset before turning on the behavior for that asset. This just raises the bar to a cache poisoning attack from "owns one domain name" to "owns a couple". Some gTLDs are $0.99 per year, or free. (The user would only have to visit a single page, which has a dozen other sites open in invisible iframes)
- klodolph 11y agoCould someone explain how cache poisoning would work here? The hash is already being verified, I assume that you would not cache a file if the hash doesn't match, the same way that you would reject a file from the CDN if the hash doesn't match.
- MichaelGG 11y agoThere's no hash to verify. Bad site preloads a bad script with hash=123. On good site, XSS injects a script src=bad.js hash=123. The browser, makes a request: GET https://good.com/bad.js https://good.com/bad.js. BUT! The hash-cache jumps in and says "Wait, this was requested from the script tag with hash=123. I already have that file. No need to send the request over the network." bad.js now executes in the context of good.com. If the hash-cache wasn't there, then good.com would have returned a 404. There's no hash collision because the request is completely elided (which is a large part of the perf attractiveness).
- AndrewDucker 11y agoHi! I'm a bit confused by this. What would the attacker here be doing, and how? I read the piece on your site, and it's not clear to me what the attacker would be updating, and what effect it would have. Can you explain?
- bugmen0t 11y ago0. evil.com hosts evil.js, <script src=evil.js integrity=foo>. 1. you visit evil.com and the browser stores evil.js with the cache key "foo". 2. you visit victim.com which has an XSS vulnerability, but victim.com thinks it is safe because it uses Content Security Policy and does not allow inline scripts or scripts form evil domains. 3. the XSS attack is loading <script src=www.victim.com/evil.js hash=foo> 4. the browser detects that "foo" is a known hash key and loads the evil.js from cache. Thinking that the file is hosted on victim.com - when the file is in fact not even present. 5. the evil.js script executes in the context of victim.com, even though they use a Content Security Policy to prevent XSS from being exploitable.
- sidarape 11y agoI see. Genuine question: I haven't thought this through but why not just execute the file in the context of the original file?
- icebraining 11y agoBecause if the file was legitimate, the site might need it to run under its own context.
- ukd1 11y agoIsn't the solution to only allow caching using the key if it's over https (to stop modification) AND in the original HTML (i.e. not added afterwards by JS). Limiting, but would cover above.
- icebraining 11y ago
- pdkl95 11y agoCache poisoning doesn't make sense when you are using hashes. If someone can generate sha384 collisions in a way that allows them to substitute malicious files in the place of jQuery, we have bigger problems. > Content injection (XSS) If we assume XSS, an attacker could simply inject whatever they want. The cache isn't needed. This still wouldn't poison any legitimate cache keys. > The client still has to find out if the server really hosts this file. So use (URL, hash) as the key in the permanent cache. This removes most of the bandwidth, and using a CDN allows for one GET per file across many sites. So what exactly is the attack? I'm really not seeing how someone could attack a permanent cache without first breaking the hashing functions that we already have to trust. edit: after reading https://news.ycombinator.com/item?id=10311555 https://news.ycombinator.com/item?id=10311555 This would work in the cases where we allow XSS (which is already a compromised scenario). Simply adding the URL (or maybe even just the hostname) prevents this entirely, and we still get almost all of the benefits for local resources, and we get all of the benefits when using a CDN. edit2: There are two issues being discussed. 1) Is the file we loaded form a (possibly 3rd party) site correct? 2) Did we ask for the correct file(s). Cache poisoning is when you can fool #1, while XSS attacks manipulate #2.
- MichaelGG 11y agoIf you add the domain, then how is it any different from existing caching? If using a CDN you're already all set; the CDN can return the file with cache forever headers.
- nickodell 11y ago>Is the file we loaded form a (possibly 3rd party) site correct? But there are also parts important to interpreting the file that aren't part of the hash, like the mime type. I think this problem is a lot more complicated than you're saying.
- airza 11y agoThe idea behind content-security policy is that it allows scripts to come only from whitelisted domains. You can't inline evil scripts and you can't link them from any domain. So, in the case of XSS, the attacker CAN'T just do whatever they want. They need to make the browser think that the script is being hosted on a whitelisted domain. Hence, the attack here is making the victim load that keyed script on a different page, then redirecting them to an XSS hole that links to that script as 'hosted' by a whitelisted domain. Since it seems to be on a whitelisted domain and match the original script's hash, it will execute on the page, which is not ordinarily possible on a page which is running CSP. I hope this encourages you to not immediately assume that large groups of people working on technically complicated problems are stupid in the future.
- MichaelGG 11y agoWhat about using HTTP headers? So websiteB wants to load something by hash, it would have to whitelist it with "AllowedCAS: hash1 hash2 hash3". In fact, CSP already does this for inline scripts. So add another attribute to the CSP header, like 'cache-src' or something, listing good hashes. XSS can't modify the CSP header, so isn't this safe?
- javajosh 11y agoLet a resource have one hash and potentially many source domains. Define a CSP whitelist consisting of trusted domains. This list is applied to resource loading in the browser, and it is also applied to cached resources in the following way: cache = { b0af301e782bf5e2a8ccce919b6ca3b70aa771db: {domains: ['evil.com','airbnb.com'], content: '...'}, 35d778783c4155c20360d269c9dd000fdcd39548: {domains: ['javajosh.com'], content:'...'} } You go to secure.com, but a malicious user has put the b0af301 script in your path. CSP's white list for secure.com is [secure.com, javajosh.com]. The browser dereferences the hash, checks against the associated domains, and rejects if a whitelisted domain isn't in that list. Your browser running secure.com would reject the b0af301 script. (Something I personally would like would be for for orgs like EFF.org to post known-good hashes, so I can always add the EFF hashes to my site's CSP whitelist, and have a warm-and-fuzzy feeling.)
- oconnore 11y agoIf you submit a request with Etag: <integrity>, the server can validate with 304, or deny with 4xx/5xx I think this would also allow the server to "pre-validate" with HTTP2 push.
- devit 11y agoPossible solutions: 1. Add an If-Hash-Mismatch header so you don't need to transfer the body 2. Add a list of hashes to accept to the content security policy headers 3. Add a list of public keys to accept to the content security policy, and allow the content if it's signed by one of those (this requires some standard way of signing things, maybe PGP/MIME or a dedicated HTTP header) 4. Only allow this from <script> and <style> tags that are in the <head>, or that are at "end" of <body> (meaning there are no tags other than <script> or <style> afterwards), or resources referenced from CSS and JavaScript files loaded that way. EDIT: 5. Add an ECC public key (Curve25519?) to the content security policy, and accept hashes where an extra attribute is specified providing an inline signature of the hash with the key The idea of the last one is that XSS would usually happen in in the middle of the body and not in the head or footer. That said, you can XSS with inline script, so it seems this only mitigates XSS vulnerability with length limitations on the payload (EDIT: nope, CSP blocks inline script).
- MichaelGG 11y agoCSP allows blocking inline scripts right? But 1. should exist regardless to complement the existing caching options. It shouldn't be sent by default to avoid adding another tracking method, but if the source page specifies a hash, and you have that hash, then If-Hash-Mismatch is perfect. 2. Bingo, winner.
- striking 11y agoCouldn't you just store the URL and hash together, or salt the hash with the URL?
- kuschku 11y agoAnd what if you’d just check the filesize? In the same way as you check for modified resources with existing caching methods?