11 ms·
If you're using them for unguessable random strings then yeah, they're not ideal. If you're using them for providing a unique id in a distributed system, with
by cetra3 4y ago
If you're using them for unguessable random strings then yeah, they're not ideal.
If you're using them for providing a unique id in a distributed system, with very little chance of collision & fitting them in a db column, then they are great.
- corytheboyd 4y agoYeah I don't really get the point of this article, if you need random values of a specific size don't use uuid, it's literally specified to be one exact length and format.
- ejb999 4y ago>>Yeah I don't really get the point of this article, To get clicks?
- corytheboyd 4y agoYou're not wrong lol
- sieabahlpark 4y ago
- tejtm 4y agoone exact length and five "versions" of the format (so far) https://en.wikipedia.org/wiki/Universally_unique_identifier#Versions https://en.wikipedia.org/wiki/Universally_unique_identifier#...
- adileo 4y agoI made a comparison list with the most known uuids out there, a couple of days ago, it was quite fun discovering all the different kinds of uid and their pros/cons. https://adileo.github.io/awesome-identifiers/ https://adileo.github.io/awesome-identifiers/
- djbusby 4y agoULID example should be in uppercase. Love this chart tho.
- JimDabell 4y agoKSUIDs are fairly popular and missing from your list: https://github.com/segmentio/ksuid https://github.com/segmentio/ksuid
- 8n4vidtmkvmk 4y agowhat's the resolution on those? 32 bits, 100 years.. that seconds right? doesn't sound excellent for time ordering. 100 years also seems a little short but at least I'll be dead
- sekh60 4y agoDon't look at it as being your problem in 100 years, but as helping employment in 100 years and helping the economy ;)
- bmn__ 4y agohttps://datatracker.ietf.org/doc/html/draft-peabody-dispatch-new-uuid-format#page-4 https://datatracker.ietf.org/doc/html/draft-peabody-dispatch...
- BerislavLopac 4y agoIt is also highly recommended that you include a check digit into it, to minimize the chance of a collision. I've used https://arthurdejong.org/python-stdnum https://arthurdejong.org/python-stdnum for that purpose.
- sokoloff 4y agoI don't see how a check digit minimizes the chance of collision. (Here, I'm assuming that a check digit is calculated from the other digits. What am I thinking about incorrectly?)
- georgemcbay 4y agoLooking at the docs for the library linked, it appears to be a Verhoeff algorithm check digit... so yeah, you're correct. This is effectively a simplistic stand-in for a CRC type system -- useful to detect if the data has been corrupted, but not useful to avoid collisions.
- TedDoesntTalk 4y agoAnd if someone is worried about UUID collisions, they need to rethink their priorities in life.
- BerislavLopac 4y agoYou are correct, this should teach me not to write comments when I'm too tired. :/ The check digit wouldn't really help with collisions, since if the strings are the same the digit will be too. They are primarily useful when we need to ensure correctness on human input.
- throwawaymaths 4y agoAlso most well-designed systems only use the UUID as the representation format and use raw bits in performance-critical parts.
- lolinder 4y agoThe raw bits are the UUID, the hex string is just a human-readable representation that also plays nicely with JSON.
- throwawaymaths 4y agoTell that to Django (well 5 years ago anyways iirc, don't know what it does now). Pretty sure it used to store uuids as strings columns in your sql.
- lolinder 4y agoYep, looks like it does the right thing in PostgreSQL but not anywhere else [0]. https://docs.djangoproject.com/en/4.1/ref/models/fields/#uuidfield https://docs.djangoproject.com/en/4.1/ref/models/fields/#uui...
- throwawaymaths 4y agoI feel like it did strings in postgres too, not too long ago and I had a <brain explode> moment when I worked on a codebase and had to figure out why queries were terrible
- lolinder 4y agoSupposedly the behavior hasn't changed since at least version 1.8: https://docs.djangoproject.com/en/1.8/ref/models/fields/#uuidfield https://docs.djangoproject.com/en/1.8/ref/models/fields/#uui... It may not have worked correctly on your project for some reason?
- 4y ago
- _3u10 4y agoUse uint128_t instead.
- scott_w 4y agoThe number of comments saying "using UUIDs for secrets isn't that bad" suggests this article needs to be written...
- hn_user2 4y agoMy only wish is that UUIDs were sortable and still contained their timestamp. When bug hunting, sometimes things become a little more obvious when there is an exact start and end to ids with issues.
- myvoiceismypass 4y agoThere are KSUIDs that aim to satisfy this A go ref impl: https://github.com/segmentio/ksuid https://github.com/segmentio/ksuid
- scrollaway 4y agoAlso, UUIDv6, v7 and v8. Still a draft. https://datatracker.ietf.org/doc/html/draft-peabody-dispatch-new-uuid-format#page-8 https://datatracker.ietf.org/doc/html/draft-peabody-dispatch...
- dexwiz 4y agoDepends on the version used. Some of them do encode time. But since people don’t like to leak information they use the random version (4).
- andreareina 4y agoThey're little endian so not sortable
- sgtnoodle 4y agoWhat does that have to do with anything?
- andreareina 4y ago>>>> My only wish is that UUIDs were sortable and still contained their timestamp. When bug hunting, sometimes things become a little more obvious when there is an exact start and end to ids with issues. >>> Depends on the version used. Some of them do encode time. Encoding time isn't enough, it has to be big endian (unless you write a special sorting function for uuids). Timestamped uuids store the timestamp as [timestamp_low, timestamp_mid, version(!), timestamp_high][1] which doesn't sort right. [1] https://en.m.wikipedia.org/wiki/Universally_unique_identifier https://en.m.wikipedia.org/wiki/Universally_unique_identifie...
- Waterluvian 4y ago“Moving Away From Misusing UUIDs”
- Alupis 4y agoThere's probably a non-trivial amount of folks that equate a UUID with "unguessable" given their appearance. They are, after all, not sequential and using them to obscure things like number of users (using a UUID in place of an incrementing number) seems like a natural fit. Given how easy it is to generate a UUID in most languages, and given the low likelihood of a collision within a system - it wouldn't be a huge leap to think UUID's could replace homebrewed random string generators for things like password reset tokens, etc.
- withinboredom 4y agoYou can generate sequential UUIDs, IIRC, that’s the best way to store them in a db and still have good partitioning/indexing. I don’t use UUIDs often, but I vaguely remember researching this problem space at some point.
- Alupis 4y agoI think most languages let you chose which version of UUID you want - with most defaulting to the random version (I think 4?) by default. There are other versions that are sequential/time-based though, but using these could open the door to de-obfuscating whatever data you wanted to protect via UUID's in the first place (like how many sales orders you receive per hour, etc).
- withinboredom 4y agoI don’t think uuids are designed for obfuscation, though they certainly help with that as a side effect. I could be wrong though, I’ve never looked into it.
- Alupis 4y agoThey (randomized type 4 UUID's) obfuscate as a side effect because they are much more difficult to guess due to their randomness. As the article points out though, they are not impossible to guess... but it will come down to your risk tolerance and what the UUID's are "protecting". People like to reach for UUID's when obfuscation is needed because inventing your own duplicate-aware random string algorithm isn't what most folks want to spend their time thinking about. Plus, these days, many databases come with UUID-aware data types that make using UUID's fairly straight forward.
- dheera 4y agoAlso UUID v3 and v5 produce IDs from identifiers such as URLs which can be quite useful if you want two different systems to generate the same exact UUID given knowledge of the same URL. For example, in a REST system that needs UUIDs I'd use the REST URL of the object as the UUID.
- human 4y agoSomething I don't understand: how are UUIDs not safe given that they are probably better than 99.9999% of passwords generated by users?
- rr808 4y agoUUIDs are nearly half the mac address of the server + a timestamp. They are in no way random.
- throwanem 4y agoThat's UUID v1. The random one that everyone uses is v4.
- kube-system 4y agoI have seem some common libraries that default to v1, so I can see why there’s some confusion in here.
- nordsieck 4y ago> Something I don't understand: how are UUIDs not safe given that they are probably better than 99.9999% of passwords generated by users? UUIDs are 128 bits. Which is beat by a 5 character a-z random string. It's certainly possible that they're better than the median password - especially if there isn't a check against a common password list. But it's pretty easy for user chosen passwords to be much, much better. I strongly doubt that your 6 9s estimate is accurate.
- prutschman 4y agoA 5 character a-z random string has log2(26^5) =~ 23.5 bits of entropy, way less than 128.
- deleted 4y ago[deleted]
- andreareina 4y ago
- ilyt 4y agoPretty much, my first reaction was "people use UUIDs for session tokens ? why? ? Seems like author made some bad choices in previous systems and now just figured out why tbh.
- epicureanideal 4y agoDepending on the UUID algorithm, some are cryptographically sufficient true random, then it would make sense..
- ksidudwbw 4y agowhat if you sha it?
- WirelessGigabit 4y agoThat actually reduces the usefulness as you're hashing the data into a smaller length.
- unlikelymordant 4y agoIt seems uuids are 128 bit, while sha is 160 bit. There is also sha256 and sha512 for longer hashed. So there shouldnt be any worries about the hash being shorter.
- markatto 4y agoYes, some hashes might not meaningfully hurt it, but they won’t add any entropy, which is the real problem.
- jchw 4y agoRereading I am guessing you're merely pointing out that the comment regarding shortening the length is untrue. If you already understand the entropy issue here, please treat my "you"s as royal you's. You have a 128 bit value. That's 128 binary digits. Each digit can be zero or one. That means you have 2^128 possible distinct values. (Ignoring the fixed bits in UUIDs since it's not important for sake of this argument.) Now you use a one-way cryptographic hash on top, like sha256. This will return a specific hash for any given input. It is always the same for a specific given input, and it is nearly always distinct. The output that a hash has may have more bits, but the number of distinct values can't increase; it can only ever decrease. That's because you could only ever give it 2^128 different values. How could it ever return more outputs if each input corresponds to one output? To make it more clear, let's say you have a database where you want to store a customer's zip code so you can use it as some kind of validation later on to ensure it matches, but you don't want to store it in plaintext, so you hash it. The hash is 160 bits. Secure, right? Wrong. There are less than 50,000 zip codes. It would be trivial to calculate the hash of every single one and use it as a simple hashmaps from hashed value to plaintext. You may be thinking this is impractical for an input domain as large as 2^128, but realistically it only adds a slight roadblock. Knowing the only valid values will be hashed UUIDs, instead of picking 160 random bits, you'd be much better off picking a random UUID, hashing it, and trying that for each attempt.
- Ptchd 4y ago> If you're using them for unguessable random strings then yeah, they're not ideal. Why? I like to use them for private/secret URLs ...
- echelon 4y agoThe best format: {opaqueTokenTypePrefix}_{crockfordEncodedEntropy} Also: pass token through a bad words and "credit card lookalike" filter. Optionally encode author cluster/region details in the low order bytes to resolve before eventual consistency in active-active systems.