16 ms·
“Should you encrypt or compress first?”
- tomp 10y agoI don't understand... Why couldn't you do CRIME with no compression as well? Assuming you can control (parts of) the plaintext, surely plaintext+encrypt gives you more information than plaintext+compress+encrypt?
- ontoillogical 10y agoCrime relies on compression --- the "CR" stands for "Compression Ratio" The idea is that the DEFLATE compression algorithm used in the TLS compression mode CRIME attacks will build and index of repeated strings and compress by providing keys to that index. Here's a beautiful demonstration of another similar compression algorithm: http://jvns.ca/blog/2013/10/24/day-16-gzip-plus-poetry-equals-awesome/ http://jvns.ca/blog/2013/10/24/day-16-gzip-plus-poetry-equal... So, if you control some subset of the plaintext you can make guesses as to what the secret you're trying to get at is, and if the size changes after compression you know that you got two hits to that bucket in the index, so your guess is right. You can use this technique to guess some string character by character -- reducing your seach space to n*m instead of n^m for a string of length m with character set of length n.
- vog 10y agoExcellent explaination, thanks!
- nightcracker 10y agoNope, secure ciphers are resistant against attack-controlled plaintext attacks. The reason that plaintext -> compress -> encrypt has issues is because the length of the ciphertext gives away the plaintext if the attacker can control parts of the plaintext.
- laughinghan 10y agoNo, because the length of the plaintext+encrypt depends only on the length of the plaintext and not on the contents of the plaintext; whereas the length of the plaintext+compress+encrypt depends on the length of the compressed plaintext, and the length of the compressed plaintext depends on the contents of the plaintext. So, different plaintext of the same lengths, which would've resulted in plaintext+encrypt of the same length, instead results in compressed plaintext of different lengths which results in plaintext+compress+encrypt of different lengths, leaking additional information!
- Asooka 10y agoNot at all - with compression thrown in, you get a measure of the entropy of the original plaintext from the length of the compressed message.
- geofft 10y agoIf you control parts of the plaintext, a well-designed cipher doesn't tell you anything about the rest of the plaintext. But if you control parts of the plaintext, a compression algorithm will absolutely tell you information about the rest of the plaintext. That's its job—identifying similarities between different parts of the plaintext. And it will leak that information in the form of the size of its output, and the standard definition of "well-designed cipher" doesn't do anything to mask sizes; the size of the encrypted data is the same as the size of the data passed to the cipher. So now both the data and its size are sensitive, and nobody's encrypting the size. (You could sort of work around this by padding, but then you've basically removed the point of compression.)
- mjevans 10y agoActually in the context of security compressing the data you're about to encrypt still matters. The problem exposed by the CRIME exploit (and anything similar) is that the size of the payload also can be an indication of the data within it. To combat this, systems which are exchanging data interactively (in a stream) should further pad messages /up to/ a target size (which might be random per message). As I mentioned elsewhere there's no reason that data can't also be potentially useful (though it does become difficult to ensure that this isn't also a source of attack).
- geofft 10y agoThat reduces, but doesn't eliminate, the amount of information you're leaking. If you pad the data to a multiple of some fixed block size, you'll still learn something across the boundary between two block sizes. An attacker can do a CRIME-style attack that includes a wrong password guess plus some of their own padding, and vary the amount of padding until it is just big enough to take n + 1 blocks instead of n. Then they can vary the password until it goes back to taking n blocks. Randomizing the padding limit also reduces the leaked data, but doesn't eliminate it: your random numbers come from some distribution, so the attacker just has to repeat each inbound message many times, and do some stats. If a correct password guess gives them, say, 1000 bytes with standard deviation 500, and an incorrect one gives them 1001 with standard deviation 500, they simply need to issue a ton of requests. (If you're padding the data to a fixed absolute size, period, and you know no message is smaller then that, then sure, please pad the message, but there's also not a whole lot of point in compressing it at all. Leave it uncompressed, and pad that.)
- JoshTriplett 10y agoIf you have a blob of encrypted data, that should normally tell you nothing about the plaintext, except its approximate length (assuming it wasn't padded). Normally, the length of the plaintext should tell you very little about its content; it doesn't do you much good to know "the message is about 4096 bytes". However, if you compress first, then you know the approximate length of the compressed data, which means you know something about the content of the plaintext. As one of several possible attacks: imagine the attacker could supply data that will get fed back to them in the encrypted blob, along with your own data, and all compressed together. The blob will compress better if your data matches their data. So, repeatedly feed in your data and get back blobs, watching the total size of the compressed-then-encrypted blob, and you can predict enough about their data to guess it.
- usloth_wandows 10y agoI thought this was common sense. Compress then encrypt. Encryption leads to higher entropy, therefore less effective compression.
- danielweber 10y agoThe article is about why that can be wrong.
- mozumder 10y agoBut in almost all cases it's right.
- danielweber 10y ago"Don't compress at all" is the better default answer. Noting that "compression after encryption is stupid."
- mozumder 10y agoThe correct default answer is encrypt after compress. "Don't compress at all" doesn't help you if you need to reduce bandwidth. What good is a secure channel if no one uses it because of its high bandwidth requirements?
- danielweber 10y agoA lot, possibly a majority, of the major breaks in crypto systems (certainly the interesting ones) in the past decade have been because of compressing before encrypting. If someone wants to compress first, demand that they justify the reduced bandwidth usage.
- IncRnd 10y agoThe interview in the article is for a security position. You answer is incorrect in that situation.
- biokoda 10y agoIf you're compressing audio, the simple solution is to compress using constant bitrate.
- VLM 10y agoUnfortunately one way to define variable bit rate is its compressed CBR. Rather than defining at the protocol level "insert comfort noise here" at the compression level you get bitstream level "I donno what this stream is at a higher level, but replace the next 1000 bits with zeros". That's if you do simple sampling. I donno about weird higher level vocoders. I think you could create a constant bit rate vocoder that really is constant. But it'll likely be uncompressible if its a good one because a vocoder basically is a compressor that's specifically very smart about human speech input. If your vocoder output is compressible its not a good one. I think if you replaced your compression step with run it thru a constant rate vocoder you'd get what you're looking for. Probably.
- dcposch 10y agoNo, he's saying you compress CBR, then encrypt. Not compress CBR -> gzip -> encrypt or something silly like that. CBR audio codec, then encrypt gives you a constant-bitrate stream indistinguishable from randomness. That's pretty much the gold standard. (Of course, that still only encrypts content, not metadata. You can encrypt a phone call in such a way that a watcher gets mathematically zero information about what's being said, but the watcher still sees who is calling whom, when, and for how long. Hiding that is much harder.)
- gravypod 10y agoLogically speaking, an encrypted file should have a high entropy set of bits within it. Compressing it would be low return, but higher security since the input file contained more "random" bits. Compressing the source material will yield smaller results but will be more predictable as the file will always contain ZIP headers and other metadata that would possibly make decryption of your file much easier.
- danielweber 10y agoIt's not just the headers. You could strip those off and count on the other side to put them back on. It's that the act of decompressing arbitrary data can leak very important information to the attacker.
- tansey 10y agoMost (good) encryption schemes have the property that knowing part of the message, or some other information like message length, will not help you decrypt the message.
- pizza234 10y ago> Compressing the source material will yield smaller results but will be more predictable This depends on the susceptibility of the ciphers to known-plaintext attack; I'm not sure if today there are scenarios where using such ciphers makes sense.
- tankenmate 10y agoBut the metadata is a form of structured data and as you state is semi known plaintext; i.e. the structure is known, and to a lesser extent also the data (some fields have a known or limited range of possibilities). Aside from the headers the compressed data it self often has structure; take the Lempel Ziv class of dictionary encoders rely on repeating data to compress. It is just this fact that it is repeating means that you can guess with much higher probability than normal what certain bytes will _not_ be (because longest words that match the dictionary are chosen to tokenised to maximise compression); i.e. bytes that _don't_ match a word suffix in the dictionary will restart the search for a new matching word / token pair. Having said that the plaintext itself is almost never random; but the key thing does the attacker have any crib that might be used to have a good first guess that can reduce the amount of work required. So which is more guessable; the semi-structure, at the byte level, of compressed data, or the possible semi-structure of the original plain text? If you are dealing with a protocol that specifies compression (and in particular which compression method) you may have given away part of the game. One way is to "bump up" the entropy and add "chaff" to the compressed stream before encryption; i.e. add some entropy but less than what would be less than the amount saved through compression. My gut feel is that the efficacy of this would vary depending on plaintext, compression method, and encryption method. You also run the risk of side channel analysis via CPU, RAM, power usage etc.
- nightcracker 10y agoThere's no compress or encrypt _first_. It's just compress or not, before encrypting. If security is important, the answer to that is no, unless you're an expert and familiar with CRIME and related attacks. Compression after encryption is useless, as there should be NO recognizable patterns to exploit after the encryption.
- vog 10y agoThis is exactly what the article says.
- loeg 10y agoThe troublesome part is the headline.
- JshWright 10y agoYeah, but the article uses a lot more words to say it...
- hinkley 10y agoDoes this really need to be said though? I may be too close to the problem. I've had to explain this to project managers and customers of course, but this is hacker news. It feels like a three page article on why you should put your socks on before your shoes and not after.
- yellowstuff 10y agoIt's not just you. I've never used encryption or compression in any serious way, but the right answer seems obvious if you know the definitions of encryption and compression.
- draugadrotten 10y agoThis blog is an interesting way to advertise to their target market: us.
- BrainInAJar 10y agoBy eliciting damnfool responses from people that haven't read the article?
- phillmv 10y agoThat part was… unexpected.
- arielweisberg 10y agoSo what does this mean if I am using an encrypted SSL connection that is correctly configured? Is this kind of problem not already dealt with for me by the secure transport layer? It would be a shame if the abstraction were leaky. My understanding of the contract is that whatever bits I supply will be securely transported within the limits of the configuration I have selected. If I pick a bad configuration then yes shame on me, but a good configuration won't care if I compress right?
- kstenerud 10y agoSo if the length of the resulting message is leaking information, salt it by adding some extra random bits to the end to increase the length by a random amount.
- peterwwillis 10y agoWhich may be useful, unless you can use a padding oracle attack or timing attack, or you're using something stupid like ECB mode, or you aren't authenticating your ciphertext. In general, it is safe to assume that whatever countermeasure you are thinking of has already been defeated by an attacker, unless you have researched for a really long time and found no possible alternative.
- vog 10y agoA more interesting question is whether to compress or sign first. There's an interesting article on that topic by Ted Unangst: "preauthenticated decryption considered harmful" http://www.tedunangst.com/flak/post/preauthenticated-decryption-considered-harmful http://www.tedunangst.com/flak/post/preauthenticated-decrypt... EDIT: Although the article talks about encrypt+sign versus sign+encrypt, the same argument goes for compress+sign versus sign+compress. You shouldn't do anything with untrusted data before having checked the signature - neither uncompress nor decrypt nor anything else.
- sirk390 10y agoI think, you mean whether to encrypt or sign first.
- vog 10y agoGood catch! Although the article talks about encrypt+sign versus sign+encrypt, the same argument goes for compress+sign versus sign+compress.
- b101010 10y agoWhy is the debate about "compress/encrypt then sign" vs "sign then compress/encrypt"? Is there a non obvious problem with sign then compress/encrypt then sign again? (overcomplicated or unnecessary?)
- mikeash 10y agoIt's pointless. If you sign the encrypted data, then once the signature is verified in the receiver, you know that the decrypted data is also good. Repeating the signature just wastes time and space.
- x1798DE 10y agoExcept what about the case where someone can spoof that an encrypted message came from them? In that case, you want the signature somewhere inaccessible, so that they can't selectively strip it off.
- kinofcain 10y agoAlso interesting is which compression algorithm you're using. HPACK Header compression in HTTP 2.0 is an attempt to mitigate this problem: https://http2.github.io/http2-spec/compression.html#Security https://http2.github.io/http2-spec/compression.html#Security
- panic 10y agoHas there been any research into compression that's generally safe to use before encryption? E.g., matching only common substrings longer than the key length would (I think?) defeat CRIME at the cost of compression ratio.
- zielmicha 10y agoI'm working (finishing paper) on an algorithm compatible with deflate/gzip that is safe to use before encryption (i.e. it is guaranteed to not leak random secrets cookies). It's a bit more complex than your suggestion - matching common substrings longer than key length would still be vulnerable as substring boundary may still fall inside secret.
- cvwright 10y agoCool! Please post a Show HN with a link to your ePrint when it's done.
- jakozaur 10y agoWould adding some tiny random size help? Based on my poorly understanding, if after compress, but before encrypt we add random 0 to 16 bytes or 1% of size that could defeat quite a lot of attacks (like CRIME).
- phlo 10y agoThat'll make an attack significantly more time-consuming, but won't prevent it. Instead of instand feedback whether they guessed correctly, an attacker would instead need to send a bunch of requests to determine if the average request size has decreased.
- dllthomas 10y agoWhat if the number of padding bytes is a function of the contents?
- lmm 10y agoThat would add noise to the information the attacker gets, but ultimately the "output" length has to be correlated with the input length, so you can't ever get the amount of information the attacker gets down to zero which is what's needed to make a cipher secure.
- dllthomas 10y ago'ultimately the "output" length has to be correlated with the input length' But perhaps not relative to the size of the secret data, right? (Note that I'm not saying, "Oh, obviously we should just do this to avoid that attack" - I'm hoping to learn by carefully understanding where it breaks down).
- phlo 10y agoCRIME works because the compression/encryption algorithms aren't aware of which data is secret and which isn't. They treat the whole HTTP stream as a stream of text and operate on that. To them, text within the page is just the same as cookie data in the corresponding header field or (for example) the URI in the header. If such a distinction were possible, the system could just not compress secret data. But if that were the case, nonsecret data wouldn't even need to be encrypted; it's nonsecret after all.
- justinzollars 10y agotl;dr
- vox_mollis 10y agoA lot of comments here suggesting that encryption increases entropy. While true, it only adds the key's entropy to the plaintext's entropy. In most real-world cases, len(m) >> len(k), so this is usually an insignificant increase of entropy. Compression also adds a trivial amount of entropy (specifically, the information encoding the algorithm used to compress, even if that information is out of band).
- srgb 10y agoI believe you confuse entropy with Kolmogorov complexity
- mjevans 10y agoWhere everyone seems to be getting confused is handling a live flow versus handling a finalized flow (a file). * Always pad to combat plain-text attacks, padding in theory shouldn't compress well so there's no point making the compression less effective by processing it. * Always compress a 'file' first to reduce entropy. * Always pad-up a live stream, maybe this data is useful in some other way, but you want interactive messages to be of similar size. * At some place in the above also include a recipient identifier; this should be counted as part of the overhead not part of the padding. * The signature should be on everything above here (recipients, pad, compressed message, extra pad). . It might be useful to include the recipients in the un-encrypted portion of the message, but there are also contexts where someone might choose otherwise; an interactive flow would assume both parties knew a key to communicate with each other on and is one such case. * The pad, message, extra-pad, and signature /must/ be encrypted. The recipients /may/ be encrypted. I did have to look up the sign / encrypt first question as I didn't have reason to think about it before. In general I've looked to experts in this field for existing solutions, such as OpenPGP (GnuPG being the main implementation). Getting this stuff right is DIFFICULT.
- arknave 10y agoI picked up on the reference to Stockfighter, but does anyone know if the walking machine learning game mentioned at the end of the article exists? Sounds like a fun game.
- ontoillogical 10y agoHeh, I was referring to OpenAI Gym (https://gym.openai.com/ https://gym.openai.com/), specifically https://gym.openai.com/envs#mujoco https://gym.openai.com/envs#mujoco
- hueving 10y agoThat quoted voip paper isn't actually as damaging as it sounds. IIRC that 0.6 rating was for less than half of the words so if you're trying to listen to a conversation to get something meaningful, it's probably not going to happen.
- ontoillogical 10y agoThat's a good point, and the .6 is an optimistic score (I think it was from running a subset of their model) I got the sense that the authors felt that this proves an attack of this sort was possible and viable, but that their model wasn't quite there yet. OTOH this should be enough to acknowledge that the voice encryption/compression scheme they are attacking is not secure.
- cvwright 10y agoFair enough. The larger point, however, is that this stuff was supposed to be encrypted. The adversary shouldn't be able to learn even one bit about the plaintext. After all, nobody would buy a new encryption module that advertised it could protect you 40% of the time.
- js2 10y agoThe paper cited in this article (Phonotactic Reconstruction of Encrypted VoIP Conversations) really deserves to be highlighted, so I submitted it separately: https://news.ycombinator.com/item?id=11995298 https://news.ycombinator.com/item?id=11995298 http://www.cs.unc.edu/~fabian/papers/foniks-oak11.pdf http://www.cs.unc.edu/~fabian/papers/foniks-oak11.pdf
- FuturePromise 10y agoGiven the real risk of CRIME attacks, are there "compression aware" encryption algorithms?
- zielmicha 10y agoI don't think it is possible - it's hard to imagine an encryption algorithm that doesn't preserve content length.
- jtolmar 10y agoIf I compress each component (ie: attacker-influenced vs secret) separately, concatenate the results (with message lengths of course), then encrypt the whole message, is that secure? It seems like it should be, but I'm not an encryption expert. The compression should be pretty good, though.
- jayd16 10y agoWould be great if Apple understood this and compressed IPA contents before encrypting. Instead, when you submit something to the AppStore, you end up with a much bigger app than the one you uploaded. To add insult to injury, if you ask Apple about this fuck up you get an esoteric support email about removing "contiguous zeros." As in, "make your app less compressible so it won't be obvious we're doing this wrong."
- arjie 10y agoNone of this seems to apply to documents you generate to supply to someone else you trust. Compress and encrypt seems perfectly fine.
- rpearl 10y agoIf you're encrypting it, it is to hide information from some sort of attacker, not the trusted recipient of the document. If there is literally no possibility at all of someone else intercepting the document, then why are you bothering to encrypt in the first place?
- arjie 10y agoBecause the wire is unsafe.
- rpearl 10y agoIn which case the considerations in the featured article apply entirely to any attacker with access to the wire/transit. Your original comment makes no sense.
- arjie 10y agoNo. The considerations in the article apply to certain kinds of input under certain conditions. Those caveats mean it doesn't apply to self-generated text documents encrypted all at once, in general.
- poelzi 10y agoif your compression can compress your encrypted data, you should change your encryption mechanism to something that actually works...
- IncRnd 10y agoDespite the question being flawed. The correct answer is a series of questions: Who is the attacker? What are you guarding? What assumptions are there about the operating environment? What invariants (regulations, compliance, etc) exist? There may be compensating controls that invalidate the perceived needs for encryption or compression, for example. i.e. don't design in the dark. Of course, the interviewer may just want a canned scripted answer - but the interview is your chance to shine, showing how you can discuss all the angles.
- em3rgent0rdr 10y agoWhat if you compress and then only send data at regular periods and regular packet sizes? That way no information can be gleaned. E.g. after compressing you pad the data if it is unusually short, or you include other compressed data too, or you only use constant bit-rate compression algorithm.
- spatulon 10y agoThat was a fun read. Do I detect a nod to tptacek's "If You’re Typing the Letters A-E-S Into Your Code You’re Doing It Wrong"? https://www.nccgroup.trust/us/about-us/newsroom-and-events/blog/2009/july/if-youre-typing-the-letters-a-e-s-into-your-code-youre-doing-it-wrong/ https://www.nccgroup.trust/us/about-us/newsroom-and-events/b...
- deleted 10y ago[deleted]
- cm2187 10y agoCan't you just add some random length data at the end. You are defeating compression a little bit, but are also making the length non deterministic. I thought pgp did that.
- Murk 10y agoSo long as you do not compress the random data you add it will work. Compressing it will not help the situation at all since the compression ratio will be constant for the random data. The problem is you have to then seriously negate the effects of compression. Consider compressing "silence" in a telephone call.. You can't compress it well if you also need to have the non silence elements be indistinguishable from it by adding random noise. You must add enough random noise to cover up any compression differential, otherwise statistical artefacts still persist. That amount of random noise will be up to the maximum compression ratio you can achieve.
- peterwwillis 10y agohttps://en.wikipedia.org/wiki/Padding_(cryptography) https://en.wikipedia.org/wiki/Padding_(cryptography)
- gameofdrones 10y agoThe OP should take https://www.coursera.org/learn/crypto https://www.coursera.org/learn/crypto
- boriselec 10y agoThis question was in one of quizzes. Expected answer: compress first.
- Qantourisc 10y agoMaybe we need encryption that also plays with the length of the message / or randomly pad our date before encryption ? I am however no expert, so I have no clue how feasible, or full of holes this method would be .
- cvwright 10y agoYes, this is one option that we proposed back in 2009 after we first pointed out the VoIP traffic analysis attacks in 07 and 08. If you want more information, see our paper on traffic morphing from NDSS 2009 [1]. It works pretty well for VoIP; not as great for trying to counter similar attacks on web browsing traffic. [1] http://www.internetsociety.org/doc/traffic-morphing-efficient-defense-against-statistical-traffic-analysis http://www.internetsociety.org/doc/traffic-morphing-efficien...
- Animats 10y agoThis is why military voice encryption sends at a constant bitrate even when you're not talking. For serious security applications where fixed links are used, data is transmitted at a constant rate 24/7, even if the link is mostly idle.
- deleted 10y ago[deleted]
- khc 10y ago> The paper Phonotactic Reconstruction of Encrypted VoIP Conversations gives a technique for reconstructing speach from an encrypted VoIP call. The technique to reconstructing speech clearly had its limitations.
- dietrichepp 10y agoWow, what a trainwreck. So many comments in here talking about whether it would be possible to compress data which looks like uniformly random data, for all the tests you would throw at it. Spoiler alert, you can't compress encrypted data. This isn't a question of whether we know it's possible, rather, it's a fact that we know it's impossible. In fact, if you successfully compress data after encryption, then the only logical conclusion is that you've found a flaw in the encryption algorithm.
- itsnotvalid 10y agoI am always thinking, if the compression scheme is known, you would need some good noonce to avoid known plaintext (for example, compression format's header is always the same), and also by CRIME, which is to remover the dictionary of the compression. I think it is best to use built-in compression scheme by the compression program to do the encryption first, as those often take these into account (and the header is not leaked, since only the content is encrypted).