7 ms·
Analyzing New Unique Identifier Formats (UUIDv6, UUIDv7, and UUIDv8)
- SturgeonsLaw 2y ago[flagged]
- ktm5j 2y agoWell, technically this is all about different versions of the same standard.
- lukev 2y agoThe cool thing about the various verions of UUID is that they're all compatible. The differences almost all come down to database locality (and therefore performance.) The exception is if you're extracting the time portion of a time-based UUID and using it for purposes other than as a unique key, but in my experience this is typically considered bad practice and time is usually stored in a separate column for cases where it matters for business purposes.
- refulgentis 2y agoIt's not necessarily that its dismissive, more so that its that a fuzzy pattern-matching comment, thats incorrect, and just a wordless link. Trivial to make, nontrivial to respond to: "Funny", in the way in-group cultural references usually are - responding means you're taking it too seriously. Yet, incorrect enough that'll misinform anyone who isn't diligently reading the full article and understands historical context. Noise thats likely to generate noise. Trolling, just missing active intent to derail.
- paulddraper 2y agoBut no one ever said UUID v_ replaces all the others. They aren't "versions" so much as variants.
- maxfurman 2y agoI'm having trouble understanding the use of v8. It can be pretty much any bits as long as it has 1000 in the right spot? It strikes me as too minimal to be useful. I must be missing something
- SigmundA 2y agoThe useful part is you can do anything you want with the other bits and have it still be a valid UUID.
- yardstick 2y agoBeing able to do anything with the remaining bits is very useful. You can do any scheme that suits your individual features needs, and it will be a valid UUID still. This also means future schemes can be implemented right now without having to get a formal UUID version. You could use the first few bits to indicate production vs qa vs dev data. Or a subtle hint to what it might be for (eg is this UUID a product identifier or a user identifier or a post or a comment etc). Similar to how AWS etc prefix IDs with their type.
- edflsafoiewq 2y agoBut it kind of defeats the purpose of encoding which of a fixed set of generation methods you are using in the ID, which is presumably to avoid having to check that none of the O(N^2) pair-wise combinations of N methods produce collisions.
- refulgentis 2y agov7 is really helpful for meaningful UX improvements. ex. I'm loading your documents on startup. Eventually, we're going to display them as a list on your home screen, newest to oldest. Now, instead of having to parse the entire document to get the modified date, or wait for the file system, I can just sort by the UUID v7 thats in the filename. Is it perfect? No, ex. we could have a really old doc thats the most recently modified, and the doc ID is a proxy for the creation date. But its much better than the status quo of "we're parsing 1000+ docs at ~random at startup, please wait 5 seconds for the list to stop updating over and over."
- canadiantim 2y agoTho presumably the uuid would give you the creation date but not the modified date. Still very useful.
- sedatk 2y agoOr just use the file date.
- refulgentis 2y agoThis is fine-ish till O(10^3)
- JadeNB 2y ago> Or just use the file date. Your parent says they don't want to wait for the file system: > Now, instead of having to parse the entire document to get the modified date, or wait for the file system, I can just sort by the UUID v7 thats in the filename.
- sedatk 2y ago> Your parent says they don't want to wait for the file system: That information comes for free when you're iterating files in a directory. There's no extra waiting than the file name itself because file dates are kept in the same structure that keeps the file names.
- oezi 2y agoI have recently wondered why Ruby on Rails is using a full-length SHA256 for their ETag fingerprinting (64 characters) when a UUID at 36 chars would probably be entirely enough to prevent collisions and be more readable at the same time. Esbuild on the other hand seems to use just 32bit (8 chars) for their content hash.
- nertzy 2y agoIsn’t it because you can generate the same content two different times and hash it and come to the same ETag value? Using UUID here wouldn’t help here because you don’t want different identifiers for the same content. Time-based UUID versions would negate the point of ETag, and otherwise if you use UUIDv8 and simply put a hash value in there, all you’re doing is reducing the bit depth of the hash and changing its formatting, for limited benefit.
- oezi 2y agoI would assume that you would only create a new UUID if the content of the tagged file changed serverside. Benefits are readability and reduced amount of data to be transferee. UUID is reasonably save to be unique for the ETag use case (I think 64 bits actually would be enough).
- ninkendo 2y agoThe point of the content hash is to make it trivial to verify that the content hasn’t changed from when its hash was made. If you just make a uuid that has nothing to do with the file’s contents, you could easily forget to update the UUID when you do change its content, leading to invalid caches (or generate a new UUID even though the content hasn’t changed, leading to wasteful invalidation.) Having the filename be a simple hash of the content guarantees that you don’t make the mistakes above, and makes it trivial to verify. For example, if my css files are compiled from a build script, and a caching proxy sits in front of my web server, I can set content-hashed files to infinite lifetime on the caching proxy and not worry about invalidating anything. Even if I clean my build output and rebuild, if the resulting css file is identical, it will get the same hash again, automatically. If I used UUID’s and blew away my output folder and rebuilt, suddenly all files have new UUID’s even though their contents are identical, which is wasteful.
- wood_spirit 2y agoI am a big fan of the new uuid v7 format. It has the advantage of being a drop in replacement most places everyone uses v4 today. It also has the advantage over other specs of ulid in that it can be parsed easily even in languages and databases with no libraries because you just need some obvious substr replace and from_hex to extract the timestamp. Other specs typically used some custom lexically sortable base64 or something that always needed a library. Early drafts of the spec included a few bits to increment if there were local ids generated in the same millisecond for sequencing. This was a good fit for lots of use cases like using the new ids for events generated in normal client apps. Even though it didn’t make the final spec I think it worth implementing as it doesn’t break compatibility
- sedatk 2y agoThere’s already a 72-bit random part. That should be sufficient to address conflicts. Incrementing a sequence completely kills the purpose of a UUID, and requires serialization/synchronization semantics. If you need that, just use a long integer.
- n42 2y agoWhat do you consider the purpose of a UUID?
- sedatk 2y agoAsynchronous unique ID generation.
- HexDecOctBin 2y agoYou can have both asynchrony and sequence by encoding thread ID in the UUID too, and make the sequence a thread local state.
- maxbond 2y agoMongoDB uses a similar approach. https://www.mongodb.com/docs/manual/reference/method/ObjectId/ https://www.mongodb.com/docs/manual/reference/method/ObjectI...
- pphysch 2y agoFor v7, the last chunk of bits (rand_b) can be "pseudorandom OR serial". There is no flag bit that must indicate which approach was used. Therefore, given a compliant UUIDv7 sample, it is impossible to interpret those bits. You can't say if they are random or serial without knowing the implementation, or stochastic analysis of consecutive samples. It's a black box. The standard would be improved if it just said those bits MUST be uniquely generated for a particular timestamp (e.g. with PRNG or atomic counter). Logically, that's what it already means, and it opens up interesting v8-style application-specific usages of those bits (like encoding type metadata in a small subset, leaving the rest random), while also complying with the otherwise excellent v7 standard.
- sedatk 2y agoSerial is just a terrible idea for UUID. UUIDs shouldn’t require synchronization to be generated.
- sedatk 2y agoI don’t understand the part where monotonicity of UUIDs is discussed. UUIDs should never be assumed monotonic, or in a specific format per se. If you strictly need monotonicity, just use an integer counter. Let UUIDs be black boxes, and assume that v7 is just a better black box that deals with DB indexes better.
- bongodongobob 2y agoInteger counters are a problem because they leak information. In most cases I've encountered that's not acceptable.
- sgarland 2y agoSo don’t expose them in the URL. Or have separate internal and external IDs. So many options that don’t destroy B+trees.
- sedatk 2y agoMonotonic UUIDs leak information too.
- switch007 2y agoUUIDv7, for example, leaks the timestamp I’ve met more than one architect who hands waves that fact away during a “leaking integers is bad!” campaign
- paulddraper 2y agoThe monotonicity can be useful in multiple contexts: colocating database data by time, providing "sooner than" comparisons. Integers are monotonic but can't be distributed like UUIDs. Unless you make them 128 bits ;) As usual, most people are not dumb most of the time, even if it seems that way.
- cm2187 2y agoAnd if you have a clustered index like in MS SQL Server, a non monotonic uuid results in inserting the data in the middle of the table (bad performance) rather than appending to the end.