5 ms·
Given that algorithms are designed to avalanche (e.g. changing one bit of the plain text causes tons of bits to change in the ciphertext), I don’t see what they
by makecheck 5y ago
Given that algorithms are designed to avalanche (e.g. changing one bit of the plain text causes tons of bits to change in the ciphertext), I don’t see what they expect to find from the encrypted data itself. Two similar encrypted sets of data would be similar essentially by accident.
They would have to get information from other observations, e.g. the fact that you send (encrypted but nonetheless observable) messages to certain people at certain times and frequencies, telling them something about the apparent importance of that relationship.
- g_p 5y agoI could see 2 ways they could do something here. Not suggesting they're feasible, but just in principle: 1. Length correlation of messages - a block cipher will be padded to the nearest block, a stream cipher could (absent padding) give you an exact length while preserving confidentiality. Length lets you infer a very high level category of content (for example, that's likely a short video, that's likely a long video, that's likely a gif, that's probably a still image). 2. Homomorphic encryption as the article suggests, or something simpler - client-side signalling of targeting parameters. Let's say I have the IAB list of 1000 ad targeting categories and I make a list of keywords for these. You could scan messages in plaintext at client-side, via the app that holds the keys. For a given message (the provider knows who you're communicating with, unless the platform is designed more like Signal, to avoid metadata). You can assign a few potential IAB keyword categories to a given conversation. Those can be encrypted into a little block "addressed" to the service provider, and appended to the message sent from the device. Good luck spotting this without static analysis of the app. If you could do this, you could add ads to a given conversation view based on IAB "categories" assigned to that conversation or group chat. You could also infer categories for a given user if there's a clustering of categories present throughout a series of conversations. And that's all just based on client side keyword scans. That's technically extracting meaning/value from encrypted data, without attempting to go homomorphic. And while FHE is cool technology, I'm not sure it makes sense commercially to pursue that if you could just quietly sneak client side targeting in and get away with it due to user apathy and regulatory stagnation. In essence if you use FHE and FB can gain any inference about information contained therein, it's no longer E2EE as the ciphertext conveys information about the plaintext absent the key. If you wish to make a leaky cipher, you could just client-side add the metadata for advertising or whatever other purpose, as either way you need to "backdoor" the client to achieve this information leakage. As you rightly say, the avalanche property of a good cipher, combined with basic principles of cipher mode initialisation and non-reuse of use-once parameters, and use of message authentication schemes that don't leak a MAC of the plaintext at ciphertext level (like MAC-then-encrypt does), means all an observer should see in common between two identical messages is the length. The rest should be entirely avalanched (with exception perhaps of a sequence header in the transport layer depending on the protocol, but even that could be encrypted)