12 ms·
Given that a UUID identifier fits in a single cipher block, and the whole point is that these are unique by construction (no IV needed so long as that holds tru
by oconnore 3y ago
Given that a UUID identifier fits in a single cipher block, and the whole point is that these are unique by construction (no IV needed so long as that holds true), it seems like a single round of ECB-mode AES-128 would enable quickly converting between internal/external identifiers.
128 bits -> 128 bits
- dajonker 3y agoThat's an interesting idea, how would you deal with the bits in the UUID that are used for the version? Setting them to random bits may cause issues for clients that try to use the identifier in their own database or application, as mentioned in the article.
- notpushkin 3y agoIs there a way to encrypt 122 bits -> 122 bits? If so – do that and set version to 4. Alternatively, just say it's a random string ID and not an UUID.
- bajsejohannes 3y agoI'm actually working on encrypting database keys like this, and I opted for a string ID with "url safe" base64. It avoids the ambiguity of looking like a UUID when it's not, and I prefer "o3Ru98R3Qw-_x2MdiEEdSQ" in a URL over "a3746ef7-c477-430f-bfc7-631d88411d49". (Not that either one is very beautiful)
- matja 3y agoCurious on the preference there, especially when characters in base64url encoding can look similar/ambiguous in some fonts.
- bajsejohannes 3y agoThe preference is purely about the amount of characters. I hope no one will ever actually read or write these URLs, but you never know...
- matja 3y agoAh in that case, base58 can be useful because it doesn't use characters that may be ambiguous in some fonts (doesn't use: 0/O/I/l , but does use: 1/i)
- lifthrasiir 3y agoFor your information, yes, you can [1]. For example if you have a good enough 128-bit block cipher (e.g. AES-128-ECB), start with a block of 128 bits where specific 6 bits are filled out and others are filled with the plain text. Repeatedly encrypt the block until those specific 6 bits are reached again (and do the same thing in reverse for decryption). This is possible because a good block cipher is also a good pseudorandom permutation, so it should have a small number of extremely long cycles (ideally just one) with a 2^-6 probability of allowed values in average. [1] https://en.wikipedia.org/wiki/Format-preserving_encryption https://en.wikipedia.org/wiki/Format-preserving_encryption
- KMag 3y agoCycle-walking FPE has a (nearly) unbounded upper-bound on latency. Breaking the 122 bits into two 61-bit halves and using AES as the round function for a Feistel cipher gives you a constant 3-encryption latency instead of the expected 64-encryption average latency of the cycle-walking format preserving encryption. Alternatively, use AES in VIL mode ( https://cseweb.ucsd.edu/~mihir/papers/lpe.pdf https://cseweb.ucsd.edu/~mihir/papers/lpe.pdf ).
- akoboldfrying 3y agoWhat a cute observation! If there's any way for the client to influence the input, it may be prone to DoS attacks: By my calculations, with a million random attempts, you would expect to find a cycle of length at least 435, which is over 13x the average. (Mind you, multiplying the number of attempts by 10 only adds about 72.5 to the expected cycle length, and probably no one has the patience to try more than 100 billion or so attempts.)
- KMag 3y agoThe properties of the permutation are dependent upon the encryption key, so a client being able to select malicious inputs to get long cycles implies either that the client knows the AES key, or that the client has broken AES. In any case, as I mentioned in a sibling comment, with 3 AES encryptions one can construct a 122-bit balanced Feistel cipher with a constant amount of work.
- deleted 3y ago[deleted]
- lysium 3y agoNeat idea. I’m afraid you won’t be able to ever rotate that key, would you? Since it’s result is externally used as an identifier, you would have to rotate the external identifiers, too.
- the_arun 3y agoWhat is the use case for rotating uuids? Aren’t they immutable?
- Raed667 3y agoi think they meant rotating the encryption key not the internal uuids
- oconnore 3y agoI think you could but it would further complicate the id scheme (would need some sort of a version mask to facilitate a rotation window).
- aurbano 3y agoAssuming you have a table where the identifiers are stored you'd have the internal one (UUIDv7) and the encrypted version from it (external id). You could rotate encryption keys whenever you want for new external id calculation, so that older external ids won't change (as they are external, they need to stay immutable).
- lilyball 3y agoIf you're storing the identifiers then you don't need encryption, you just generate a UUIDv4 and use that as the external identifier. And then we're back at where the blog post started.
- fodkodrasz 3y agoNo IV, ECB mode... why bother with encryption at all? Just expose the internal id.
- remram 3y agoBecause the internal ID exposes timing/sequence information, as per jonhohle's comment.
- bajsejohannes 3y agoYou would still use a secret key, so it's impossible for the end user to decrypt it.
- adrian_b 3y agoEncrypting an internal id with ECB into an external id continues to allow the comparison for equality of 2 ids, to determine whether they are the same or not, but except for this it removes all the information contained in the structure of an UUID.
- computerfriend 3y agoECB is perfectly secure when you use it on a single block.
- fodkodrasz 3y agoAnd you have no problem with the same data being encrypted being identifyable (no salt). i see now why it might be useful in this case, though I still don’t like the idea, it feels bad somehow. (Yeah I get that you save storage for some computation)
- isbvhodnvemrwvn 3y agoYou are encrypting a single block of unique information. No other encryption mode gives you any advantages whatsoever.
- MattPalmer1086 3y agoThis seems overly complex, and you need some kind of key too. Why not just hash it with pretty much any hash function?
- bajsejohannes 3y agoBecause a hash is (by definition) a one-way function. You need to be able to go the other way for incoming ids.
- MattPalmer1086 3y agoAh, good point. Hadn't thought about the opposite direction.
- oconnore 3y agoA hash is not reversible, so you’d need a database index to recover the original efficient-to-index identifier, which misses the whole point ;) If you didn’t care about index clustering then just use UUIDv4
- masklinn 3y agoAlthough you could use a hash index to avoid the more deleterious effects of insertion, as long as you don’t need a unique constraint anyway…
- dragontamer 3y agoWhy not just use the AES-128 result as the UUID then? What's the benefit of the internal structure at all? If AES-128 is an acceptable external UUID (and likely an acceptable internal one), then you might as well just stick with a faster RNG.
- oconnore 3y agoThat would be the same as using a random identifier (UUIDv4, for example) with the associated indexing issues when stored in a database. The whole point here would be that you can expose something opaque externally but benefit from well behaved index keys internally.
- amelius 3y agoStorage is cheap, you might as well store the extra integer.
- masklinn 3y agoStorage is cheap, updating indexes is not.
- techdragon 3y agoThis is why I’ll probably just always use a UUIDv7 primary key and a secondary UUIDv4 indexed external identifier… which is extremely close to how I tend to do things today (I’ve been using ULID and UUIDv4)
- turtles3 3y agoBut you still need an external->internal lookup, so doesn't that mean you still need an index on the fully random id?
- RhodesianHunter 3y agoWhy not use snowflake IDs?
- nindalf 3y agoOne key for all tokens or one key per token? If it’s the latter a simple XOR would do because it would be the equivalent of a one time pad.
- robertlagrant 3y agoI don't think it can be a key per token, or it will scale appallingly.
- Nevermark 3y agoOne key per token would require a table matching internal tokens to their key for forward conversion, and another table matching external keys to their key for reversing. Might as well just use randomly generated external keys and have one table if you were doing that. So, one key per all tokens.
- andix 3y agoI like the idea, but I think it's not possible to rotate the key with that approach, without introducing a breaking change. Eternal secrets are usually a very bad idea, because at some point they are going to be leaked.
- remram 3y agoSure you can, just prefix the encrypted identifier with a version number. https://app.example.org/file/1:abcdef12345 -> decrypt abcdef12345 with key "1" to yield UUIDv7 key of file (no matter what the latest key is)
- andix 3y agoAn incrementing version number would once again leak time information. Even a not incrementing version number would leak that kind of information, because if you know the timestamp of another ID with the same version. I think there is no good alternative to random external identifiers.
- remram 3y agoOh that's a good point.
- londons_explore 3y agoBut the private data you are protecting (the user's account creation time) has the same properties as an eternal secret. Therefore there doesn't seem to be much downside in this specific case.
- mort96 3y agoIf you were using something like UUIDv4, you wouldn't be exposing that information at all though, neither in cleartext or ciphertext. It seems weird to say, "the user ID contains secret information so we encrypt it with an eternal fixed pre-shared key then share the ciphertext with the world", when you could've just said "the user ID contains no secret information". It feels like the right solution here is to pick between: use UUIDv7 and treat the account creation time as public information, or use an identifier scheme which doesn't contain the account creation time.
- nhoughto 3y agoYep have used an approach just like that, worked quite well if you have a strong pattern to easily translate from one to the other. Gives you an id with the right properties for internal use, efficient indexing etc, and in its encrypted form gives you the properties you want from an external identifier being unpredictable etc, all from one source id. It is true that now your encryption key is now very long lived and effectively part of your public interface, but depending on your situation that could be an acceptable tradeoff, and there are quite a few pragmatic reasons why that might be true as has been described by other comments. Edit: you can even do 64bit snowflakes internally to 128bit AES encrypted externally, doesn’t have to be 128-128 obvs
- phkahler 3y ago>> It is true that now your encryption key is now very long lived and effectively part of your public interface No need to encrypt, just store the external key in a table. Not that you're likely to change algorithms.
- phkahler 3y agoLate edit: I meant to say No need to encrypt on the fly. Do it once and save it.
- nhoughto 3y agoTrue you could rotate by persisting the old value and complicate your lookup/join process, not my idea of an acceptable solution but yep totally possible and worth it for some set of tradeoffs.
- kevincox 3y agoYou are basically describing BuildKite's previous solution.