14 ms·
UUID v7
- lordgilman 3y agoIs the headline true at all? It was punted to the next commitfest on 2/1. It could easily keep getting punted like it has been for a ~year now.
- coldtea 3y agoIs it that hard to implement? Supporting an additional UUID version in PostgreSQL sounds like the most trivial change to implement (compared to anything that touches core backend, table management, replication, query schedulling, and so on).
- lordgilman 3y agoThe patch is already written, it's on that page. The bottleneck in Postgres is reviewer bandwidth which is why it's been moved out of several commitfests.
- pgaddict 3y agoI don't think reviewer bandwidth is the main issue for this patch. It's a 200-line change (considering C code, there's more in docs/tests), and the code is not overly complicated / sensitive (in the sense that it's very isolated and unlikely to break random stuff). For me the main challenge was that it's still considered a draft (AFAIK). It may be unlikely to change, but if it does I'd rather not have to deal with persistent UUIDv7 data generated per some previous spec. Also, if I really want/need UUIDv7, it's not that hard to create an extension that generates UUID in arbitrary ways, including the proposed v7.
- x4m 3y ago+1 there's already well maintained extension https://github.com/fboulnois/pg_uuidv7 https://github.com/fboulnois/pg_uuidv7 It's slightly different from recommendations by draft RFC version (there's no counter), but fully within spec requirements. From practical point there's no difference at all.
- asabil 3y agoThe UUID spec update has not been finalized yet. It would be quite unfortunate to end up with a UUID v7 in PostgreSQL that’s not quite the standardized one because the patch got merged too quickly. EDIT: here is the IETF working group page https://datatracker.ietf.org/wg/uuidrev/about/ https://datatracker.ietf.org/wg/uuidrev/about/ Their milestone seems to submit the final proposal by March.
- Rafert 3y ago> It would be quite unfortunate to end up with a UUID v7 in PostgreSQL that’s not quite the standardized one because the patch got merged too quickly. The chances of that seem extremely low at this point. The contents of a version 7 UUID have not changed since work started on RFC 4122 bis in October 2022: https://author-tools.ietf.org/iddiff?url1=draft-ietf-uuidrev-rfc4122bis-00&url2=draft-ietf-uuidrev-rfc4122bis-14&difftype=--html#:~:text=implementation.-,5.7.%20%20UUID%20Version%207,-5.7.%20%20UUID%20Version https://author-tools.ietf.org/iddiff?url1=draft-ietf-uuidrev...
- pgaddict 3y agoThe chances may be low, but either it's a draft or a final version. There's clearly little pressure to rush this, considering it's not difficult to add a custom function generating UUIDv7 ...
- bebna 3y agolooks like v7 is basically v1 updated. v6 looks interesting with the db optimisation in mind. UUIDv7 spec: https://www.ietf.org/archive/id/draft-peabody-dispatch-new-uuid-format-04.html#name-uuid-version-7 https://www.ietf.org/archive/id/draft-peabody-dispatch-new-u...
- michaelmior 3y ago> Systems that do not involve legacy UUIDv1 SHOULD consider using UUIDv7 instead. The optimization relevant to v6 also applies to v7. The difference is that v7 UUIDs are not directly compatible with v1 UUIDs.
- yrro 3y agoFYI that's a link to a now very old internet draft. The current draft can be found here: https://datatracker.ietf.org/doc/draft-ietf-uuidrev-rfc4122bis/14/ https://datatracker.ietf.org/doc/draft-ietf-uuidrev-rfc4122b...
- naranha 3y agoDoes this stay compatible with the old uuid column type, as uuidv7() uses the same 128 bit format as gen_random_uuid()? That would mean it's easy to update old apps, as we would only have to change the default value for the columns.
- pgaddict 3y agoYes, internally it's the same as every other UUID (16 bytes, passed by reference). There's no reason to store / represent it differently.
- mplanchard 3y agoYes, we have already updated our in-app UUID generation to use v7 UUIDs and are storing them in regular postgres UUID columns (postgres 14). Works great!
- ghusbands 3y agoNote that, as jmull says in https://news.ycombinator.com/item?id=39262286 https://news.ycombinator.com/item?id=39262286 , embedding timestamps in every uuid can potentially expose private information.
- mplanchard 3y agoYes, you can check UUIDs to see if they’re v7, and extract the timestamp if so. This seems to me less problematic in most cases than being able to guess the next ID (as is the case with numeric IDs). At least for us, anybody with access to the ID also has access to the time the record was created, so there’s no new information being exposed. It’s a good thing to keep in mind though for sure.
- Merad 3y agoYes, you can use v7 today with uuid columns, you just need to add a custom function (or the more performant pg_uuidv7 extension) to generate them.
- throw0101a 3y agoRelated on the front page, "UUID Benchmark War": * https://ardentperf.com/2024/02/03/uuid-benchmark-war/ https://ardentperf.com/2024/02/03/uuid-benchmark-war/ * https://news.ycombinator.com/item?id=39254871 https://news.ycombinator.com/item?id=39254871
- deleted 3y ago[deleted]
- zukzuk 3y agoNo thread about UUID is complete without a plug for NanoID! https://github.com/ai/nanoid/blob/main/README.md https://github.com/ai/nanoid/blob/main/README.md
- __s 3y agoNot sure it's too relevant for postgres, where uuid is 16 bytes in db & can be generated by db
- swrobel 3y agoAlso not comparable to UUIDv7 because it isn't sortable
- sltkr 3y agoThat's a lot of words to say: “encode 126 random bits in base64 (url-safe variant)”. It seems the algorithm is equivalent to: function nanoid(alphabet="ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-_") { return Array.from(crypto.getRandomValues(new Uint8Array(21)), i => alphabet.charAt(i & 63)).join(''); } (Maybe String.fromCharCode() is slightly faster than Array.join(), but I doubt it matters much.) On node.js it's even easier, since url-safe base-64 encoding is supported natively: Buffer.from(crypto.getRandomValues(new Uint8Array(16))).toString('base64url').substring(0, 21) Do we really need an entire Github project dedicated to a 1-liner?
- didntcheck 3y agoAgreed, and I actually share the same view about "UUID" as a named concept itself. I wrote more in another comment, but in summary * Using uncoordinated random generation in a large address space for IDs is often a very useful idea * 128 bits is a good rule of thumb for "will never ever collide" * Making this into a heavyweight "UUID" concept, with it's own bespoke string format, and 7 different standard ways to generate them, feels like a ridiculous waste of cognitive effort that makes such a simple concept appear opaque and magic. If you want to encode other data (timestamp, node ID) in 16 bytes you can still do that of course. There's no need for anyone else to even know. Just do some quick calculations to ensure you haven't eliminated too much entropy
- hans_castorp 3y agoI wonder why the size of the table with uuid7 is smaller than the one with uuid4? Both are using the uuid data type.
- sparsely 3y agoThe main data is the same, it's the primary key which is smaller, presumably a result of the structure of the uuid7 being more regular.
- uhoh-itsmaciek 3y agoAs a sibling comment guessed, btree indexes in Postgres can store ordered data much more efficiently than random data. Inserting randomly ordered data leads to fragmentation, with lots of empty space on your index pages.
- nivertech 3y agoNice, but somewhat pointless, since the UUID (or any other type of ID representing an external entity) better be generated by the actor creating that entity, i.e. the FE client, or in the worst case, the BE/application server. If it is a database server generated ID, it could as well be a serial autoincrement bigint ;) EDIT: For those who misunderstood: I'm very much pro-UUID, and against serial autoincrement server-generated IDs. With the exception when you need to heavily optimize for speed and/or index storage space. And even then there are hybrid solutions like using UUID externally, and serial IDs internally. https://news.ycombinator.com/item?id=39261435 https://news.ycombinator.com/item?id=39261435
- arghwhat 3y agoThere are many reasons for not using a bigserial index, such as avoiding the data leak associated with publicly visible sequential database indexes. These reasons apply regardless of where an index is generated. Generating IDs away from the database is mainly done due to design restrictions demanding it, not because it is beneficial.
- beeboobaa 3y agoTrusting your clients to generate your database IDs... What could go wrong?
- piaste 3y agoWho said anything about trust? Server-side validation still applies; you don't just DELETE /user/{id} without verifying ownership, regardless of where the id comes from. But client-generated IDs make idempotency easier and remove whole classes of errors. They're typically a huge win.
- nivertech 3y ago1. you're not trusting anyone 2. the UUID by itself doesn't authenicate or authorize anything 3. there is a small chance of collision, and it can be handled on the backend/persistance/DB layer, i.e. return error to the client in case of collision and ask to generate a new UUID 4. many non-trivial and/or CQRS/ES apps work like this 5. if you are really paranoid, you can push down UUID generation logic to BFF (Backend-For-Frontend) layer 6. Lots of DistSys problems can be solved with client-generated IDs. But most people mistakenly think that DistSys applies to backend only, and exclude clients from the picture. 7. Scaling RDBMS (especially Postgres) is hard. UUID generation is slower than serial bigint, so it's best to keep it outside DB layer fo this reason also. 8. Client-generated UUIDs help to make client requests idempotent, and enable error handling with retries (although it's better to add an additional layer of request idempotency with IDEMPOTENCY_KEY HTTP header, or GraphQL Relay's clientMutationID).
- s4i 3y agoCan someone ELI5 what that "UUID v7 support" actually means in the title? I don't know how to navigate commitfest (nor would I probably understand the source code to begin with), but the reason I'm confused is that you can already use all the proposed draft UUID implementations in Postgres (as long as you generate the ID application-side). In fact, PG will happily accept any 128 bit ID to be inserted into a UUID column, as long as it's hex encoded – even the dashes are optional.
- Skinetio 3y agoYes my expectation is that Postgres can do that than for you. Its still easier to have an autogenerate on than doing it externally
- Scandiravian 3y agoPostgres can now auto-generate uuidv7, so you can define it in your schema and don't have to add the ID on the client side when doing an insert
- wongarsu 3y agoThe patch is adding functions for generating UUIDv7 in the database. Which you can already do, for example with the PL/pgSQL function from [1]. Or you can just generate them in your application code. As you mentioned everything else about UUIDv7 already works and didn't require any changes. It's really more about convenience, and maybe a bit of speed. 1: https://gist.github.com/fabiolimace/515a0440e3e40efeb234e12644a6a346 https://gist.github.com/fabiolimace/515a0440e3e40efeb234e126...
- s4i 3y agoThank you!
- didntcheck 3y agoI'm amazed at how "we" have managed to turn such a simple idea as "128 bits is a large enough address space for uncoordinated generation to be essentially collision free" into such a "heavy" concept with 7 different versions If you want to do something smart like encoding your node ID within the value, or prefixing a timestamp for sortability, then sure, do that in your application. No one else really needs to care how you produced your 16 bytes. Just do some napkin math to make sure you're keeping sufficient entropy I'm not sure "UUID" even needed to be a column type, versus a "INT16" and some string formatting/parsing functions for the conventional representation (should you choose to use that in your application). You could also put IPv6 addresses in the same type. Though I guess this depends on how much you think the database should encode intention versus raw storage in types
- sonium 3y agoYou probably should not use UUIDs to start with in your database at least not as an ID. UUIDv7 aims solve some of the issues of UUIDv4 that are even less suitable in for databases. 99% of times using BigInt for an ID is better.
- kkzz99 3y agoHuh, why would using uuidv4 be a problem? Collisions?
- davydog187 3y agoMy understanding is that they cause a lot of page fragmentation, which leads to excessive writes to the WAL
- brodo 3y agoYou can't use them as cursors, because they are not inherently ordered like integer ids.
- dventimi 3y agoI have never wanted to use database cursors and I predict that I never will.
- brodo 3y agoSure, if you don't offer pagination or only have small tables, you can get away with offsets. I tend to go for cursors as a default because I like to build applications with performance in mind and it’s the same effort.
- dventimi 3y agoWe may be talking about different things. I thought you were referring specifically to [database cursors](https://en.wikipedia.org/wiki/Cursor_(databases) https://en.wikipedia.org/wiki/Cursor_(databases)) so that's what I was talking about. If you're talking about something else, like the concept of so-called "cursor-based pagination" in general, then that is still an option even with even randomly-generated primary keys, so long as there are other attributes that can be used to establish an order (which attributes need not be visible to the a user)
- nodesocket 3y agoWhat’s a use-case for a v7 UUID vs using a v4 UUID which is entirely random based? Is it just that v7 includes a timestamp for better sorting?
- vlucas 3y agoYes. UUIDv7 can be sorted by default in order of creation, which helps reduce fragmentation problems of randomness over time.
- guptaneil 3y ago> Is it just that v7 includes a timestamp for better sorting? Correct. The sortable nature of UUIDv7 improves database performance and index locality by helping the index be more efficient since rows are inserted in a predictable order instead of scattered randomly.
- mrits 3y agoIf nothing else (and there is plenty of else) it saves a lot of time knowing the order something was created when debugging.
- mderazon 3y agoWould you recommend starting a new project with UUID v7 ?
- politelemon 3y agoIf you need some of its attributes. Some comments: https://news.ycombinator.com/item?id=39261469 https://news.ycombinator.com/item?id=39261469 https://news.ycombinator.com/item?id=36433481 https://news.ycombinator.com/item?id=36433481 And have a look at this post from last year: https://news.ycombinator.com/item?id=36438367 https://news.ycombinator.com/item?id=36438367
- cryptos 3y agoIf the included timestamp doesn't expose sensitive data, then using UUID v7 is a good default, because it has performance advantages on database operations and items might be sorted by ID in a meaningful way, what is sometimes desired (even if I never used it).
- jmull 3y agoWord to the wise: be very careful about adding semantics to unique ids that aren't inherent to the identity of the thing being identified. Over time conflicts between the id's primary job (uniquely identifying something) and the extra semantics can arise, and the solutions tend to get pretty messy. Here we have a unique id that embeds a timestamp. The classic conflict here is with privacy/security. A UUIDv7 user id tells you when the user was created. A UUIDv7 of a medical record tells you when some medical event occurred. There are things whose identity is inherently time-based and not private, so I'm not giving a blanket recommendation to not use these. Just understand what you are signing up for. For a database, you can use bigints for primary ids but only internally. Then you also have an external random (v4) uuid... and a timestamp if you want, for that matter -- now that it's a separate column, you can expose/hide it on a case-by-case basis, depending on need. So this gets you the benefits of a uuidv7 but maintains flexibility, though at the cost of some complexity and extra bytes/record. Other conflicts can arise too, and they can be hard to always foresee, so generally be careful about extra semantics in unique ids.
- personomas 3y agoIt looks like you're explaining this well, but I still don't understand what you're saying.
- nomercy400 3y agoDon't use UUIDv7 if you want to keep the creation time of the event/id/entry private.
- sgarland 3y agoThey're saying that in certain circumstances, if your API exposes the PK publicly, it may leak information you don't want leaked (the precise datetime something occurred, in the case of UUIDv7). If that's an issue for you, you can get around this in a variety of ways, as they mention: you could use an associative table that maps the externally-exposed random ID to an internal-only ID.
- remus 3y agoBe careful about combining 2 pieces of information in to 1 column. From the above examples, you may want something to uniquely identify a record in your db and you may want something that tells you when the record was created. If you combine these two things, you then have a problem if you want to give an untrusted party that unique reference without telling them when it was created.
- hardwaresofton 3y agoIf you like this (I do very much), you might also like pg_idkit[0] which is a little extension with a bunch of other kinds of IDs that you can generate inside PG, thanks to the seriously awesome pgrx[1] and Rust. [0]: https://github.com/VADOSWARE/pg_idkit https://github.com/VADOSWARE/pg_idkit [1]: https://github.com/pgcentralfoundation/pgrx https://github.com/pgcentralfoundation/pgrx
- personomas 3y agoWhen will this be available? (Thanks in advance)
- ashconnor 3y agoUseful explanation on UUIDv7 https://buildkite.com/blog/goodbye-integers-hello-uuids https://buildkite.com/blog/goodbye-integers-hello-uuids
- ComputerGuru 3y agoIt would be nice to have either fully integrated support for returning the UUIDv7 encoded in crockford's base32 or else ship a crockford's base32 encode/decode functionality in the same release so that this can be compatible with Ulid (given that it's one the primary open source existing works UUIDv7 was modeled after).
- simlevesque 3y agoIf you need a uuid v7 right now you can get one here: https://uuidv7.app/ https://uuidv7.app/
- qingcharles 3y agoThank the Gods. I was almost out.
- simlevesque 3y agoAlways glad to help.
- presentation 3y agoWhoosh
- samatman 3y agoI've said this before on HN, but we're talking about UUID v7 again, so it bears repeating: prefer UUID v4 for unique keying, of the sort that's transferable between databases. For timestamps, use timestamps, to sort by insertion, use an autoincrementing primary key. Disk space is not so expensive that we can't afford all three of these things. Conflating timestamps and uniqueness is a conflation. These concerns are best separate. If you put a timestamp in your UUID, your UUID now has a timestamp. You can't remove it if you don't want the timestamp to be a part of the uniqueness any more.
- oxfordmale 3y agoUUID7 offers several key advantages: - sortable by insertion - less vulnerable to sequence prediction attack - allows partitioning of tables in the future Sequence prediction attack is a problem when you want to use identifiers in a public API. For example, a user or competitor can iterate through your product catalogue by incrementing the key. Sorting by inserting is a common use case as well. You can achieve this with an auto-incrementing primary key, however, this will create issues when you need to partition the table. Of course, you may never need this functionality, but there is a reason UUID7 have been added, as they are very useful in certain scenarios.
- codr7 3y agoI get the public API argument, but it certainly feels backwards to structure the internals of your application for that specific use case. Generating and mapping an UUID to your identifier when you need it is pretty straight forward and allows you to avoid exposing the identifier at all. Which leaves partitioning, something that very few applications will ever need, even if plenty of developers hope they will. A slightly less weird use case, but perfectly doable up to a point using regular sequences. I see a solution looking for problems.
- oxfordmale 3y agoI have build many successful systems without ever using a GUID. However, at a certain scale they become very handy, and it then it definitely helps you can also sort them.