9 ms·
Security advisory: Breach and Django
- cschmidt 13y agoI'm sure it will come, but I'd appreciate a layman's terms explanation of this. What is the threat, and how do you go about fixing things in Django?
- steveklabnik 13y agoHere's a simple explanation: http://arstechnica.com/security/2013/08/gone-in-30-seconds-new-attack-plucks-secrets-from-https-protected-pages/ http://arstechnica.com/security/2013/08/gone-in-30-seconds-n... It's not Django, but us over at Rails have been discussing various parts of BREACH and how we'll handle it: https://github.com/rails/rails/pull/11729 https://github.com/rails/rails/pull/11729 The important two comments are here: https://github.com/rails/rails/pull/11729/#issuecomment-22066881 https://github.com/rails/rails/pull/11729/#issuecomment-2206... and https://github.com/rails/rails/pull/11729/#issuecomment-22083669 https://github.com/rails/rails/pull/11729/#issuecomment-2208... > Let's let this stew for a while with security researchers doing their > analysis on various approaches and wait and see what the security community > as a whole recommends. > > My only concern to rushing out a release is that we do something equally > dumb and end up creating a different problem for our users. > > We can roll out fixes as it becomes clear what the consensus is as to the > best solution for a generalised framework like Rails. As http://breachattack.com/ http://breachattack.com/ says: * Be served from a server that uses HTTP-level compression * Reflect user-input in HTTP response bodies * Reflect a secret (such as a CSRF token) in HTTP response bodies These things are easy to tell about your application, but are much harder for frameworks to detect generally, which is why projects like Django and Rails will take some time to evaluate exactly how to best handle this at the framework level.
- skizm 13y agoThe link is like 5 sentences long and 2 of them are recommendations for stopping the attack.
- publicfig 13y agoJust because it's short doesn't mean it's simple to understand. Also, the recommendations aren't about how to fix things in Django, but instead a way to prevent this attack from happening. Hardly a long-term solution, especially because it basically asks you to avoid using GZIP for compression, which many people and organizations rely on in order to process timely responses.
- Hovertruck 13y agoThis advisory is pretty understandable, I think. The big bold text ("BREACH may be used to compromise Django's CSRF protection") is a strong warning of the threat (becoming vulnerable to XSS). They list two steps that they recommend taking; disable the gzip middleware in your settings.py, and disable gzip for responses from your web server.
- jeffasinger 13y agoI'm not all that familiar with BREACH, so please correct me on the parts I'm wrong about, but it seems that it's an attack that allows one to recover some data sent over TLS if compression on TLS and the protocol level is enabled. In Django, this means that attackers could recover the CSRF token that's used to prevent cross site requests. This means anyone between you and a client could later have that client automatically make authenticated requests to your app, simply by visiting a site they control, without the knowledge of the user. To protect yourself, the Django team recommends turning off compression either at the TLS level and at the HTTP level.
- steveklabnik 13y agoThe only part you're wrong about is the compression on TLS part, it's not an 'and': > While CRIME was mitigated by disabling TLS/SPDY compression (and by modifying > gzip to allow for explicit separation of compression contexts in SPDY), > BREACH attacks HTTP responses. These are compressed using the common HTTP > compression, which is much more common than TLS-level compression. This > allows essentially the same attack demonstrated by Duong and Rizzo, but > without relying on TLS-level compression (as they anticipated). and > It is important to note that the attack is agnostic to the version of > TLS/SSL, and _does not require TLS-layer compression._ breachattack.com
- IvyMike 13y agoImagine you're going to send a compressed and encrypted message to a friend, and I (the attacker) can do two things: 1) Append a bit to the message before it is compressed and encrypted. 2) See the size of the final message. So I start by appending the string "4179174b19e0cdc91bf4" to your plaintext message. I see the final encrypted message size is 500 bytes. Then, I redo the experiment, but this time, I append the string "cschmidt@example.com" to the message. The final encrypted message size is now 480 bytes. The string I injected was the same size, but the compression worked better this time, and I can guess it's because the string I picked is redundant with something in your plaintext. Mix in a bunch of complicated math and a bit of javascript, and you've got an exploit. This threat isn't specific to Django: it's being billed as a TLS attack, but any encryption system that uses compression the same way is vulnerable.
- deleted 13y ago[deleted]
- RyanZAG 13y agoSwitch off all GZIP..? That feels very extreme, I'm sure there are better workarounds than that one. EDIT: The following workarounds should be very simple to implement and seem like more viable alternatives for production? Length hiding (by adding random amount of bytes to the responses) Rate-limiting the requests Mitigations 6 and 7 taken from http://breachattack.com/ http://breachattack.com/
- derefr 13y agoAll these framework-vendor guides will recommend switching off Gzip, because it's a content-neutral workaround; it works everywhere, for every instance of the attack, no matter how you've coded your app. There are more specific workarounds, but they require changing how you encode secrets into your page, so there can't really be a vendor guide on how to do that; the vendor doesn't know how and where your app sticks secrets into its views, after all.
- tptacek 13y agoI'm sure everyone is going to come up with workarounds that re-enable compression, but they'll be context-dependent and will involve code; in the meantime, the attack is straightforward and viable. Think of disabling compression as a stopgap.
- illumen 13y agoDefinitely think about it before just doing it though... Disabling compression can break some apps. Especially when they rely on huge compression ratios for text (5-10 times ratio is common for with much json for example). So that is not an app agnostic work around. For example, a 100k of json request, can turn into a 1MB json request. The more data required to send, the more chance of error - especially on 3g/2g networks. For many high end projects, just disabling compression without regard to testing or having an idea of what the application is doing would get you fired or taken to court. Not only would this break apps, but it would also lose business in that there is evidence from Amazon and others that every 100ms extra latency can cost 1% in sales. From SPDY whitepaper: "45 - 1142 ms in page load time simply due to header compression". Remember that headers use the upload part of the link... which means too many headers and you can saturate the upload, therefore making the whole internet connection stall for everyone using it. Common upload limits are only 5-10K/second, so excessive headers combined with many requests can easily DOS many internet connections. I spend a lot of time optimising websites for these reasons, and disabling compression could add 20 seconds of load time for a good percentage of users. So, for many apps, turning off compression is no solution at all. You might as well just disconnect your app from the internet - that will also give you a secure and broken app. A proper risk, and impact analysis should be done first. Too often quick hot fixes to security issues just break things or even make things less secure.
- tomp 13y agoSo, it seems that even if I encrypt everything, a lot of information is still present in the size of encrypted message; in case of VOIP, it's possible to guess speech that is being transferred over an encrypted transport, in the case of text, it's possible to figure out secrets if the attacker can modify an equally-sized part of the message. Is there any general way of preventing this kind of attacks? Inserting random data could work, but it's distribution would have to be exactly right for the attack to be impossible over longer periods of time. For the BREACH case, we could solve it by not compressing user input, but what about the VOIP case? Also, why does the site http://breachattack.com/ http://breachattack.com/ says that "Randomizing secrets per request" is less effective than disabling compression?
- sdevlin 13y ago> Is there any general way of preventing this kind of attacks? Disabling compression is a 100%-effective countermeasure for compression oracle attacks. > Also, why does the site http://breachattack.com/ http://breachattack.com/ says that "Randomizing secrets per request" is less effective than disabling compression? Putting random data in the server response will only slow down the attack. With enough requests, the noise from that random data will wash out. Disabling compression will stop the attack cold. The whole thing is predicated on analyzing the size of the compressed text. No compression, no compression oracle.
- STRML 13y agoHow is it that random data would only slow it down? If I add a random field to my response of variable length, say, from 0 to 50, with completely random characters, it should completely throw off this attack. The length of the output will change from request to request. I suppose, given infinite time, you could send the same request over & over and map the variance of content lengths, and get an idea of what the actual content length was before random padding? But the compression seems to throw that off even more AFAIK - because the data we pad with is random, it could very well accidentally compress well because of the rest of the data in the response, further throwing off any guesses. Edit: From the pdf on breachattack.com: While this measure does make the attack take longer, it does so only slightly. The countermeasure requires the attacker to issue more requests, and measure the sizes of more responses, but not enough to make the attack infeasible. By repeating requests and averaging the sizes of the corresponding responses, the attacker can quickly learn the true length of the cipher text. This essentially boils down to the fact that the standard error of the mean in this case is inversely proportional to p N, where N is the number of repeat requests the attacker makes for each guess.
- brokentone 13y agoCorrect me if I'm wrong, but it appears as though Django isn't the only framework/technology that is vulnerable to such an attack, they're just one of the first to provide a mitigation strategy (resulting in this post).
- steveklabnik 13y agoAny website that * Be served from a server that uses HTTP-level compression * Reflect user-input in HTTP response bodies * Reflect a secret (such as a CSRF token) in HTTP response bodies is vulnerable, regardless of technology. The mitigation strategies were given in the original paper[1], this announcement is just repeat of what's in there. That said, it's exactly the right thing to do, that's not a knock on Django. 1: http://breachattack.com/#mitigations http://breachattack.com/#mitigations
- homakov 13y ago> * Be served from a server that uses HTTP-level compression this is the only must-have. Last 2 exist almost in every website (everyone needs a CSRF token, everyone reflects something somewhere)
- steveklabnik 13y agoMy blog does not have CSRF tokens, and does not reflect user input. Many web _apps_ have these things, but many web sites do not.
- homakov 13y agosince it's attack on secrets then we don't even consider blogs and read only media as targets :)
- steveklabnik 13y agoAbsolutely, I'm just saying that your statement that it's the only must have is factually incorrect. I could include some sort of secret on every page, but since my blog doesn't reflect user input, it would be fine. I could also have a 'search' function that would reflect input, but without the secret in the body, the secret would be fine. All three parts are a must-have, even if they are incredibly common. Saying otherwise is misleading.
- danso 13y agoA few days ago, Meldium's announcement of a Ruby gem that provides an inexpensive partial protection (i.e. not disabling gzip) made it to the HN front page: http://blog.meldium.com/home/2013/8/2/running-rails-defend-yourself-against-breach http://blog.meldium.com/home/2013/8/2/running-rails-defend-y... The two protective measures are masking the Rails CSRF token and appending a HTML comment to every HTML doc to slow down plaintext recovery. How easy is this to include in a Django plugin?
- mhurron 13y agoIs a partial workaround really better than a guaranteed workaround.
- z-factor 13y agoThe attacker has to be able to issue requests on behalf of the user with injected "canary" strings. I fail to see a practical exploit where one can do this and wouldn't have access to the secret in the response anyway. What am I missing?
- veesahni 13y agoI'm in the same boat - if the attacker could inject strings into requests pre-compression, then wouldn't the client already be compromised?
- rlucas 13y agoNo, you're missing that the original GET requests can be performed in some cases over HTTP, either by forgery or by surreptitiously spoofing the user's own browser into doing it. No need to have compromised the SSL/TLS.
- tptacek 13y agoDoes any GET or POST URI endpoint in your application accept parameters? Do none of those parameters impact the output of the application? That set of circumstances is extraordinarily common.
- z-factor 13y agoThe request has to be issued by the attacker from the victim's browser. If the attacker can do that, why is he unable to read the response to that request? Edit: I think I can see a scenario where a third-party website does these requests via an <iframe> or an <img>. I'm not sure there's a way to do POST quite as easily.
- tptacek 13y agoDo you understand how CSRF works? Just think of it in terms of CSRF. Since the attacker is trying to infer page content, they don't care that the server rejects all the probing requests, so CSRF protection doesn't help you as the attacker carries out the BREACH/CRIME stuff. If the result of the attack is an inferred CSRF token, they then cap the whole exploit off with a (now working) actual CSRF attack.
- homakov 13y agoim a rabbit
- homakov 13y agoRails is not vulnerable to cookie forcing btw. But your authentication can be (Devise bug http://blog.plataformatec.com.br/2013/08/csrf-token-fixation-attacks-in-devise/ http://blog.plataformatec.com.br/2013/08/csrf-token-fixation... )
- JshWright 13y ago>And i was trying to report it, but didn't find a contact/email. How hard did you try? Django uses the industry standard security@ address for reporting security issues. A quick googling results in this page pretty easily: https://docs.djangoproject.com/en/1.5/internals/security/ https://docs.djangoproject.com/en/1.5/internals/security/ EDIT: I described the link as 'first' in the Google results, but that was because Google was being helpful and promoting a page I've visited a lot before... In reality, it's a few links down.
- homakov 13y ago1) open https://www.djangoproject.com/ https://www.djangoproject.com/ 2) find /contact/, /email/, user group
- steveklabnik 13y agoI've never really looked at Django's site before, picked 'community' on the upper right, and it says > Report potential security issues in Django via private email to > security@djangoproject.com, and not via Django's Trac instance or the > django-developers mailing list
- JshWright 13y agoThe community page (where all the mailing lists and contact addresses are listed) say: >Report potential security issues in Django via private email to security@djangoproject.com, and not via Django's Trac instance or the django-developers mailing list
- STRML 13y agoCould somebody help me understand how this attack would be viable? It seems like the attack has the following requirements: 1. You want a secret that appears in the response body, like a CSRF token. 2. The web server always responds with the exact same response for a request. 3. The response body contains data that you send to the server, e.g. url params. 4. The attacker has access to an environment where he can send requests under your browser session (otherwise, the user would be unauthenticated and there would be no secrets to steal). Given (4.), how is this a real concern? If I, an attacker, am able to make 3000+ requests while logged in under your session and modify the request character by character pre-encryption, doesn't it logically follow that I have your cookies anyway?
- Erwin 13y agoThe #4 is not that difficult without compromising the user's browser -- as long as the user can visit the site under your command or see some HTML under your commend you can make the browser do a HTTPS request to anywhere at all. Maybe you buy some targetted ads served in an iframe. Maybe you send the user an email where his email server either always shows images, or you trick the user in clicking 'display images' with promise of kittens. You won't be able to see the results directly, but if you can observe how long the encrypted responses will be, you'll know whether your reflected input could make use of the compression dictionary (meaning your reflected input matches the secret) or not. I wonder if there is any way to even do this without the passive network snooping -- like some kind of internal browser stats API call that tells you # of HTTPS bytes transferred. It could be innocent enough so it's not protected.
- tptacek 13y agoNo, it does not logically follow. The attack means that someone who can (for instance) poison the DNS can (for instance) evade CSRF protection for the domains they've poisoned.
- Erwin 13y agoSo to be clear: 1) The attacker must be on the same network as you, or at least be able to detect how large the compressed and encrypted replies are. If you are on the same network it seems to be there are far more MITM and whatnot attacks that are more likely to succeed, if you do not use HSTS (or secure DNS if that helps). 2) The attacker must be able to get your browser to rapidly generate many (how many?) requests from your browser to the site. It takes "30 seconds" they claim, but is that at a rate 100 requests per second? 3) Each request must carry something that will be reflected by the body of that particular page when it's rendered. I suppose it could be an error message or search string that's echoed. It seems to me that unless you generate a CSRF token unconditionally on every page, the subset of pages that both reflect something with no protection (e.g. search results) and have a protected form (e.g. change my email address to XYZ) might be small. 4) The secret that can be extracted is what's in the reply body and not the headers -- headers are not compressed, since the TLS compression is now universally disabled post-CRIME. Personally I use Referer header checking as well. IME all the browsers of my users do send them. So if you extract the CSRF token, it's useless by itself unless you also can make the browser send the right Referer header (and AFAIK, all the holes such as Flash have been plugged). Other than that -- it seems that if you are normally generating e.g. a 32 byte CSRF key, you could interleave it with 32 bytes of good randomness per request?
- JshWright 13y ago>Other than that -- it seems that if you are normally generating e.g. a 32 byte CSRF key, you could interleave it with 32 bytes of good randomness per request? Which would be pretty easily averaged out with a few more requests. It makes the attack a little harder, but not substantially so.
- d0m 13y agoOne could argue that when talking about security, it's always about making things harder to breech, not a full proof protection.. I'm not sure how adding 32 bytes of "good randomness" would help.. because the size might be very similar since the randomness might not get properly reduced. And thus, the slight variation in size will still be very relevant. However, adding between 1 and 32 bytes of randomness might be a pretty good counter! I.e. If you request the page with the "guess letter A", and then request again the same page with the same "guess letter A" and you get +/- 32bytes of different encrypted stuff, it's very hard to assume something was better compressed. The cool thing about that is that it's fairly trivial to do with most implementation of CSRF. Thoughts?
- level09 13y agoThis would cause a big problem for us. we mobile web service serves around 3-4k concurrent requests on average. without compression our API would take 300% - 900% increase in the delay. is there any alternatives ? would like to know what Cloud Flare would do as their CDN is based on compressed nginx responses.
- pudquick 13y agoAs gzip compression only applies to the content of the page, not the headers, I would assume that prefixing your page with content that is variably compressible and of varying lengths would throw a monkey wrench in the attacks. The compressed content of any part of a page very much depends on what came before it. Altering the content to include a script comment block full of random text and various common HTML and JavaScript elements (Markov chains anyone?) would definitely change how a page is compressed. If the compressed length of the replies varies significantly with every request - even if the request content is identical - attacks like this can no longer reveal hidden information. Edit: You could improve this significantly by including false positive matches as well. If your HTML content has: csrf="45a7..." in it, you could hash that content into enough material to generate 19 or so identical looking code blocks embedded in a script comment. You've now provided a 95% chance they attack the wrong one / increased the number of attacks they'll need to try by 20x. This method (minus the above part) would actually be cacheable by smart CDNs like Cloudflare.
- Daniel_Newby 13y agoRandom padding can be averaged out. It increases the work factor of the attack, but not by much.
- sehrope 13y agoHow about having the CSRF token change with each request? If it's encrypted/signed by the server for each request with a random IV then it would be different in each request. It would be a bit more processing on the server (decrypt vs just HMAC verify) but it would be completely different each time. It seems kind of belt and suspenders as you're encrypting data within an encrypted channel but I think it gets around this issue.
- dvogel 13y agoIf the CSRF token changes with each page view then opening a second page (perhaps an explanation for a form field) in a new tab/window would invalidate the form in the original tab/window.
- krapp 13y agoMaybe only have it generated for views of the page with the form on it?
- sehrope 13y agoNot necessarily. The token can be used to simply verify that the request came from a legit page and not cross site request. The encrypted CSRF need only be verified by the server to see if it's not expired. The server can store the expiration in the CSRF token itself (encrypted and signed). It does not need to maintain a list of the CSRF tokens. I wrote about this a little while back. Comments are here: https://news.ycombinator.com/item?id=5971464 https://news.ycombinator.com/item?id=5971464
- dvogel 13y agoWouldn't the lifetime of the token have to be <30 seconds, according to the claims made in the paper?
- sehrope 13y agoI don't think so. Here's the snippet from the linked PDF[1]: > DEFLATE [2] (the basis for gzip) takes advantage of repeated strings to shrink the compressed payload, an attacker can use the the reflected URL parameter to guess the secret one character at a time. By encrypting the CSRF token (or any other "secret" data you want to roundtrip from server to client and back) with a random IV per request this wouldn't work. The value sent by the client would not be the same as the new token generated by the server (since each has a random IV). Even though the decrypted value of each token is the same, the values presented to the client in the response body are each different and not predictable (to the client). [1]: http://breachattack.com/resources/BREACH%20-%20SSL,%20gone%20in%2030%20seconds.pdf http://breachattack.com/resources/BREACH%20-%20SSL,%20gone%2...
- homakov 13y agoOfftopic: this is very simple mitigation for any website, requires JS: https://gist.github.com/homakov/6147227 https://gist.github.com/homakov/6147227
- pquerna 13y agoHas anyone looked at mitigating the attack by changing the behavior of chunked transfer encoding? Chunked Transfer encoding is basically padding that a server can easily control, without having to change content or behavior of a backend application. A web server could easily insert an order of magnitude more chunks, and randomly place them in the response stream.
- donaldstufft 13y agoI'm not sure I fully understand the proposed fix here, how does it differ from the application simply including random chunks of data inside the response? This area of things isn't my strong suite, but assuming that this is analogous to just adding random data to the response, I believe that simply adding random data to the response can be worked around by doing more requests as using statistics to factor out the noise introduced. If my understanding is wrong then excuse me :)
- pquerna 13y agoYes, it functionally the same. The difference is that it is extremely easy to add at an http proxy or load balancer level, and could potentially turn 30 seconds into hours; I would love a way to figure out the math on how many bytes of random response length changes the number of requests needed?
- lnanek2 13y agoI have seen random workarounds at the app level as well, where the app adds a random length HTML comment on the end of the page. But if random can be statistically removed, then they shouldn't add a random amount. Maybe just track the max size of the returned response and always add enough to reach that max size. Therefore the lengths of all the pages will always be the same. This is still better than turning off compression completely. A typical max for a detail page in an app might just be the size of the page plus 256 bytes per app output field.
- pquerna 13y agoYou can basically add as many or as few bytes as you like, by abusing the chunk-extension in chunked encoding: http://tools.ietf.org/html/rfc2616#section-3.6.1 http://tools.ietf.org/html/rfc2616#section-3.6.1 So you could make all http responses round into 128 byte chunks, by appending 1 to 128 bytes at the end of every response. Effectively it gives you length hiding at an http layer; Still attackable.
- dangayle 13y agoDisable compression altogether? That's craptastic.
- gojomo 13y agoCurrently, the Django templating tag: {% csrf_token %} ...results in an insert like... <input type="hidden" name="csrfmiddlewaretoken" value="566e4606b2094c7c48e5d04b58236f51"> I suspect that the particular mitigation strategy the BREACH authors' describe as "Randomizing secrets per request" could be implemented by having {% csrf_token %} instead emit: <input type="hidden" name="random_data" value="91178a84e0bc6e08a2fda853eef2d2c8"> <input type="hidden" name="csrfmiddlewaretoken_xor" value="e0b594e902c7fe6b1748d13aefaf63aa"> ...where the random_data changes every response, the emitted csrfmiddlewaretoken_xor is the real token XORed with the random_data, and upon submission the server will again XOR the two values together to get the real CSRF token. There may be other secrets that need protection in other ways, and maybe this would make any random-source issues more exploitable... but this would seem to protect the CSRF token, in a cheap and minimal way. UPDATE: Thinking further, though, maybe the attacker can probe for both values at the same time, and thus determine the probability of certain pairs, and thus this only slows the attack? I'd appreciate an expert opinion, as this was the first mitigation that came to mind, and if it's wrong-headed I'd like to bash my intuition into better shape with a clue-hammer.
- homakov 13y agoYour UPDATE is right, attacker can probe a..z 2 times, and just choose the letter that was compressed in both of them, ignoring random compressions
- gojomo 13y agoThanks, but can you clarify... does that mean probing (a..z)×(a..z) (one pair per probe), so there's at least a giant increase in probing required per character? And perhaps even more each character in, since probing for the Nth character now requires (a..z)^(N-1) × (a..z)×(a..z) ? (I'm guessing also, though, it may be possible to probabilistically probe multiple ranges of the secret at once... in a process that seems vaguely similar to forward-error-correction coding.)
- homakov 13y agoalso we need a reflector with 'value="' in the beginning https://twitter.com/homakov/status/364872768921165824/photo/1 https://twitter.com/homakov/status/364872768921165824/photo/...
- sbov 13y agoJust to make sure I understand this correctly: is this only a security issue if you include sensitive information on a page by default? For instance, if you had a search field, the contents of what users puts in that search field will not be compromised. However, if you include a csrf token with the search field form, that can be compromised since it will be there every time the attacker gets the victim to make a request.
- softbuilder 13y agoThis attack works very much like the game Mastermind. http://en.wikipedia.org/wiki/Mastermind_(board_game) http://en.wikipedia.org/wiki/Mastermind_(board_game)
- chopin 13y agoThat's true for almost all side channel attacks using user controlled input (eg. padding oracle attack).
- e12e 13y agoLooking at https://github.com/django/django/blob/ffcf24c9ce781a7c194ed8722b850e7873922f6b/django/middleware/csrf.py https://github.com/django/django/blob/ffcf24c9ce781a7c194ed8... I'm a little confused about how the csrf-token is generally used in Django -- but if I understand the code correctly, it looks for a cookie with the csrf_token, and compares that to a POSTed value (or x-header in case of an Ajax request). If the system has a decent random-implementation there is no secret involved, just a (pseudo)random string -- essentially a csrf cookie is given the client on one request, and compared on the next request(s). Is there any reason one couldn't simply use the rotate_token()-function on every (n) request(s)?
- lpomfrey 13y agoI've knocked up a package that provides CSRF token masking and length modification that may help mitigate this. If anyone wants to vet it and submit pull requests, you're more than welcome. https://github.com/lpomfrey/django-debreach https://github.com/lpomfrey/django-debreach