24 ms·
Goodbye integers, hello UUIDv7
- samatman 3y agoRelying on timestamps to be sortable, when clock skew and ntd guarantee that they won't always be, strikes me as poor design. If you need to sort by insert order, use an autoincrementing integer, if you need uniqueness, UUIDv4 is fine, if you need both use both. Use timestamps when you need to record the time, just don't commit the sin of presuming that clock time will never run backwards, I assure you, it does.
- jmmv 3y agoThe problem they describe is not about sorting: it’s about data locality. And for the latter, clock skew should not be a problem. Even with significant clock skew, data will end up clustered anyway, much better than with a random spread.
- samatman 3y agoYou get index locality with an autoincrement also, and it will actually reflect insert order. My point is that a timestamp won't do that, and worse, it will appear to most of the time. The failure can be fairly spectacular, a Unix clock can be set to any time at all, and it's good when the resulting bugs are limited to time. Having an actual insert order can be a real boon to figuring out what happened. I hold to the principle that relational data should be normal, and combining uniqueness with a timestamp doesn't do that. To do any of the calculations we use timestamps for, you have to strip off the entropy, this complicates pushing it down to the database level, where the libraries don't expect such conflation. You're going to have a bad time writing something like a join across tables with a restricted range of time if your time is embedded in UUIDv7. I maintain this is good advice: if you need index locality and insert order, use an autoincrement. If you need to record and work with time, use a timestamp. If you need global uniqueness, you can use any of the UUIDs, but v4 is the one that doesn't conflate uniqueness with unrelated properties, and should be preferred. If you think your need data locality but not insert order, think long and hard about what you're doing, because odds are you're wrong. If it turns out you're right, and the OP might be in that situation, sure, go ahead and use UUIDv7. Just, please, for the sake of your future self and everyone you work with, don't use a timestamp for insert order. Ever.
- deleted 3y ago[deleted]
- atonse 3y agoThat's still fine. Because even if there's skew and other such things going on, it's more likely to take advantage of cache locality since the page in the index that key would be stored in, is much more likely to still be in memory.
- Pxtl 3y agoFrustrating, I looked up MS/C#'s implementation and they don't get stored in a proper semisequential fashion in MS SQL Server because MS stores UUIDs in an odd binary format.
- wolverine876 3y ago> the random nature of standard non-time-ordered UUIDs (such as v4) can create database performance problems when used as primary keys. This problem is often referred to as poor database index locality. Couldn't that be solved with incremented serial numbers, rather than leaking time data?
- FPGAhacker 3y agoYes. But I think part of the design requirements is minimizing coordination required among distributed nodes. This solution attempts to solve the sort-ability issue of current uuids by moving the timestamp to the most significant bits.
- Ginden 3y agoIncremented serial numbers can leak things like eg. volume of sales in your shop.
- wolverine876 3y agoSo increasing serial numbers with random gaps?
- dkubb 3y agoIf it helps anyone, at work, I open sourced the UUID v7 postgresql function that I wrote: https://github.com/Betterment/postgresql-uuid-generate-v7 https://github.com/Betterment/postgresql-uuid-generate-v7 We've seen some amazing benefits, especially around improving the speed of batch inserts.
- perfmode 3y agoI chose ULIDs for a recent project. Hope it won’t bite me in the future.
- 0pteron 3y agoIs it not the case that having 128 bit primary keys take up 4 times as much memory as 32 bit integers when keeping the indices in RAM? I guess if you need the index to be clustered by time and also need a unique identifier in most queries then UUIDv7 fits your use-case but I still think having integer for the primary key will fit most use cases and be more efficient
- pyrolistical 3y agoAnd you can use it today with Postgres uuid type. Postgres doesn’t care what you store in it as long as it has the correct length. So you can generate a uuidv7 and store it natively
- jolux 3y agoWouldn’t the index types need to be updated to support ordering on UUIDs?
- perfmode 3y agoWhat are the benefits of using the Postgres uuid type (versus using TEXT or VARCHAR)?
- jvolkman 3y agoIt's 16 bytes versus 36 bytes.
- netcraft 3y agototally tangential to this but just have to share learned this last week that md5's are also 128 bits and so will fit perfectly in a uuid type in postgres, saving space in the physical table and indexes, and giving better index performance. So if youre storing md5 hashes, use UUID!
- thangngoc89 3y agoIt’s stored in binary format (16 bytes) instead of text (36 bytes)
- rockwotj 3y agoI find it interesting that it’s quoted random IDs are bad for performance, because it’s actually better for distributed storage systems because you don’t hotspot on a single node. For example see: https://stackoverflow.com/a/53901549 https://stackoverflow.com/a/53901549 and https://medium.com/google-cloud/cloud-spanner-choosing-the-right-primary-keys-cd2a47c7b52d https://medium.com/google-cloud/cloud-spanner-choosing-the-r...
- EGreg 3y agoThey likely mean it’s good for latency and not necessarily for throughput. I still think that graph databases are way better for this sort of thing.
- rockwotj 3y agoThey later note most of this traffic is going to a single postgres instance. Having all the keys go to the same range probably helps throughput because they can do a better job of grouping fsync. But that probably depends on the type of drives they are using (even fast NVMe benefit from locality).
- fnordpiglet 3y agoIn all the distributed systems I’ve built I hashed the keys to ensure good distribution. A nice thing of ordered keys is you can use part of the ordering to distribute keys with a tunable amount of key locality in each node for efficiency.
- akira2501 3y agoUUIDv7: Timestamp up front, random in the back.
- labster 3y agoTruly, the mullet of unique identifiers.
- wenc 3y agoI know HN doesn't like jokes, but this is really funny. And the subcomment about mullets too. (for folks who don't get it, mullets are a 1980s haircut (think MacGyver) with a short front but a long tail in the back. A funny description of them is "business in the front, party in the back")
- rockwotj 3y agoThe first time I heard about ordered string IDs was Firebase’s push IDs. They had an interesting solution to also address time skew to get better ordering for drivers: https://firebase.blog/posts/2015/02/the-2120-ways-to-ensure-unique_68/ https://firebase.blog/posts/2015/02/the-2120-ways-to-ensure-...
- user3939382 3y agoIt’s nice for front end state. You post the new entity, the front provides the ID, and as long as you get a 200 you can update your state, or update optimistically and roll it back. You don’t need to wait for the API to figure out what your ID is.
- mooreed 3y agoFeels like a spiritual successor to the ksuid [1] lib which I first heard of used in conjunction with DynamoDB [1]: https://github.com/segmentio/ksuid https://github.com/segmentio/ksuid which has very similar use cases.
- andersa 3y ago> We use sequential primary keys for efficient indexing, and UUID secondary keys for external use. The upcoming UUIDv7 standard offers the best of both worlds Unless you consider users being able to extract the generation time from the id to be an issue, of course.
- giancarlostoro 3y agoI've been seeing a few different vendors do this already. MongoDB's ObjectIds are inherently timestamps (so you can actually generate generic MongoDB IDs to query based on time). There's also Discord's Snowflakes as well. I'm sure there's loads of others. All it tells you is when something was generated, not much else. I do love how MongoDB has it stored in such a way that it is easy to query against. I wonder if any RDBMS' will allow you to query these timestamps as well.
- andersa 3y agoThere are definitely many cases where it isn't an issue since you were going to tell the user the time anyway (like sent time on a message)
- giancarlostoro 3y agoI think Twitter also does it as well. I think its really nice honestly.
- madeofpalk 3y agoFWIW, Snowflake came from Twitter https://blog.twitter.com/engineering/en_us/a/2010/announcing-snowflake https://blog.twitter.com/engineering/en_us/a/2010/announcing... Discord uses Twitter's Snowflake.
- afavour 3y agoCan’t agree with that logic. Unless it’s specifically documented leaking timestamp data is going to get totally forgotten. So when you add (e.g.) the ability to change the sent timestamp on a message you’re going to inadvertently leak when a timestamp has been changed. Could cause embarrassment in a lot of scenarios.
- amanzi 3y agoCan you take the first portion of the UUIDv7 string, and decode it to figure out the exact date and time that record was created? I'm wondering if there might be security/privacy concerns in some situations if the UUID codes are visible in your app?
- jonhohle 3y agoI just commented the same thing. I can't imagine most applications would want to leak time information in their identifiers but these undoubtedly will be used most placed out of convenience. In a year or so we'll read about an attack and everyone will migrate back to v4 or have to maintain a cryptographic identifier in addition to their temporal identifier.
- riffraff 3y agoMost application may not want it, but will it hurt them? I mean this is a similar concern to sequential IDs: many apps do not want to leak them, and in some cases it might cause issues, but in general it doesn't matter.
- afavour 3y agoYeah, years from now we’re going to see some story about how a company fudged their timestamps in order to get away with X, only to be given away by the timestamp hidden in public UUIDs.
- declan_roberts 3y agoIt’s 2023. Why aren’t we using more characters from the utf-8 keyspace to make things like UUIDs use less characters?
- chewbacha 3y agoUUIDs are 128-bits, not characters. The string representation is just for humans.
- shepherdjerred 3y agoI think the parent is saying that we can make UUIDs more human-readable by displaying the underlying 128 bits with a larger set of characters.
- dragonwriter 3y agoWe could make the string representation more compact with more characters, but I’m not sure compactness and readability are the same, especially any compactness that takes more than the ASCII character set (sure, just the hexadecimal digits may be fewer than ideal.)
- chungy 3y agoBase58 and Base64 exist in pure-ASCII space. Expanding into non-ASCII characters would probably just be confusing.
- treve 3y agoI think it's a fair question, because yes you should store them as numbers, but they are still often sent in text formats and urls. Wanting a shorter representation is reasonable. The easiest is probably to just base64 the binary representation of the 128 bit number, which results in a 128/6=22 character string, which is a bit smaller. If glyph-length and not byte-length is more important you could go even smaller but I'm less sure if that's a good idea.
- gabereiser 3y agoThis is the way. Look not at the characters but at the hex.
- deleted 3y ago[deleted]
- jonhohle 3y agoThis is great for internal distributed systems where having ordered keys is useful, however, it should probably be noted that these probably shouldn't be used as public identifiers (even though this will probably be the defacto standard and used publicly without thought). Having any information, specifically time information, leaking from your systems may or may not have unanticipated security or business implications. (e.g. knowing when session tokens or accounts are created).
- 0xEFF 3y agoEvery access and id token issued by oidc already has an issued at (iat) and expiration (exp) fields.
- rtsil 3y agoDon't create them each time a record is created, create a batch in advance in sufficient number, and do the same every time the previous batch has ran out. UUIDv7 is 128 bits, you can store a large number of them without major penalty.
- Macha 3y agoBut then you need to have the client communicate with the server to identify it's newly created object or complicate your logic to have incomplete objects in a pending state, which is one of the things people were using UUIDs to avoid.
- rtsil 3y agoYou don't store the UUIDs in the database as incomplete records. You can put them in a unused_uuids table and store some of the values in memory to minimize the round-trip. You can even store them in a simple file, and remove each used UUID from that file. When the file is empty, you create a million more of them.
- Macha 3y agoThe incomplete objects refers to when someone clicks "new" in your UI. Until it's saved back to the server, that "new" object has no ID, since you need to communicate with the server somehow to get that UUID in this approach. So now the client creates objects without IDs, so now all your models need to assume IDs are optional, and you can't create your object references on unsaved objects.
- jiggawatts 3y agoIt seems insane to me to “validate” GUIDs/UUIDs. Half the point of these things is that they’re treated as opaque identifiers.
- masklinn 3y agoI’d assume the validation is parsing the uuid with a uuid library (to decode it), and the library eagerly validates the version field, either to check for garbage or because it wants to yield a different subtype for each version.
- philsnow 3y agobut why decode it at all, if it's meant to be opaque?
- 8organicbits 3y agoI think there are a couple minor problems: - if the ID is intended to be opaque then the vendor shouldn't document it as a UUID, as this changing to a different format would be a breaking change - if the customer isn't going to process the subcomponents of the UUID then they should process it as an opaque string - if the UUID library encounters a version number in a UUID it doesn't understand, it shouldn't reject the UUID but present it as an unstructured string. After this blog post it seems likely that even Kite more customer will parse the IDs to extract time, since this has been documented.
- masklinn 3y agoProbably because a typed UUID avoids treating it like a random string, and decoding the UUID means you have 16 bytes per in memory rather than 36 (assuming usual 8-4-4-4-12 representation over the wire).
- kijin 3y agoIf UUIDv4 was all that ever existed, there would be no need to validate anything apart of the fact that it's supposed to contain 32 hexadecimal characters. All other versions, including the new v7, attach meaning to certain bits of the identifier. That cat has been out of the bag for a long time, so now everyone needs to maintain code to ensure that some rogue node doesn't spew back-dated identifiers belonging to the wrong department.
- Lazare 3y agoUUIDv7 is a nice idea, and should probably be what people use by default instead of UUIDv4 for internal facing uses. For the curious: * UUIDv4 are 128 bits long, 122 bits of which are random, with 6 bits used for the version. Traditionally displayed as 32 hex characters with 4 dashes, so 36 alphanumeric characters, and compatible with anything that expects a UUID. * UUIDv7 are 128 bits long, 48 bits encode a unix timestamp with millisecond precision, 6 bits are for the version, and 74 bits are random. You're expected to display them the same as other UUIDs, and should be compatible with basically anything that expects a UUID. (Would be a very odd system that parses a UUID and throws an error because it doesn't recognise v7, but I guess it could happen, in theory?) * ULIDs (https://github.com/ulid/spec https://github.com/ulid/spec) are 128 bits long, 48 bits encode a unix timestamp with millisecond precision, 80 bits are random. You're expected to display them in Crockford's base32, so 26 alphanumeric characters. Compatible with almost everything that expects a UUID (since they're the right length). Spec has some dumb quirks if followed literally but thankfully they mostly don't hurt things. * KSUIDs (https://github.com/segmentio/ksuid https://github.com/segmentio/ksuid) are 160 bits long, 32 bits encode a timestamp with second precision and a custom epoch of May 13th, 2014, and 128 bits are random. You're expected to display them in base62, so 27 alphanumeric characters. Since they're a different length, they're not compatible with UUIDs. I quite like KSUIDs; I think base62 is a smart choice. And while the timestamp portion is a trickier question, KSUIDs use 32 bits which, with second precision (more than good enough), means they won't overflow for well over a century. Whereas UUIDv7s use 48 bits, so even with millisecond precision (not needed) they won't overflow for something like 8000 years. We can argue whether 100 years is future proof enough (I'd argue it is), but 8000 years is just silly. Nobody will ever generate a compliant UUIDv7 with any of the first several bits aren't 0. The only downside to KSUIDs is the length isn't UUID compatible (and arguably, that they don't devote 6 bits to a compliant UUID version). Still feels like there's room for improvement, but for now I think I'd always pick UUIDv7 over UUIDv4 unless there's an very specific reason not to. Which would be, mostly, if there's a concern over potentially leaking the time the UUID was generated. Although if you weren't worrying about leaking an integer sequence ID, you likely won't care here either.
- kiitos 3y agoSecond precision is too coarse for many (most?) use cases.
- erik_seaberg 3y ago> first component (prefix) of the identifier is a sortable timestamp > values generated are practically sequential These statements aren’t strict enough to be relied on. Maybe you have engineered the hell out of your distributed clock scheme, and your IDs actually are completely monotonic, which is great. But you probably haven’t done that, which means conflicts will surely happen and you must handle them gracefully.
- QuadrupleA 3y agoBut the timestamp is less than half the bits. The rest are random. So timestamp conflicts don't matter.
- erik_seaberg 3y agoBy “conflict” I don’t mean a UUID collision, I agree with the logic that 2^128 is so much entropy that memory corruption is the more likely culprit. I mean that you can’t rely for correctness on time(X) < time(Y) when X happened before Y. It’s damn hard to keep two commodity server clocks within ±1 ms of each other even within a single LAN, and across production you’re more likely to see ±10 ms, or worse if your sysadmins don’t realize you intend to bet the farm on no clock skew.
- jpgvm 3y agoIt's not designed for this use case. It's for cases where events in the same epoch may as well of happened concurrently. ULID on the other hand does address this case by providing monotonicity within an epoch for a given producer. So would still need to treat each producer as a separate partition of the key space but you could order X > Y as long as both were produced by same producer (which effectively acts like a sequencer in this case). EDIT: nvm, the UUIDv7 spec -allows- for arbitrary allocation of the remaining 62 bits which can be used as a counter: https://www.ietf.org/archive/id/draft-peabody-dispatch-new-uuid-format-01.html#name-uuidv7-layout-and-bit-order https://www.ietf.org/archive/id/draft-peabody-dispatch-new-u... All points for ULID can also apply to UUIDv7 depending on generation algorithm.
- 3y ago
- LAC-Tech 3y agoWhy use UUIDv7 over ULIDs? As Lazare points out in this thread they're basically the same thing, except with ULIDs you get those 6 extra bits of randomness back that UUIDs have to use for metadata.
- oittaa 3y agohttps://datatracker.ietf.org/doc/html/draft-ietf-uuidrev-rfc4122bis-11#name-motivation https://datatracker.ietf.org/doc/html/draft-ietf-uuidrev-rfc... ULID isn't an "official" standard like UUID. Having a real standard usually promotes interoperability and makes it easier to use. Additionally as others have pointed out you can already use UUIDv7 with some databases since it's just 16 opaque bytes and the database doesn't care what's actually in the UUID field.
- jhealy 3y agoI'm not the author, but I work at the same company. ULIDs are nice, but we're a 10 year old company with many TBs of data across multiple logical databases and most rows had UUIDv4 ids. Maybe if we were starting from scratch ULIDs would have been an option, but given where we were UUIDv7 was a much easier transition.
- dataangel 3y agowhy bother with any version of the uuid standard? just generate a random 128-bit number and use it. that's all the newer ones are anyway
- lelanthran 3y ago> why bother with any version of the uuid standard? just generate a random 128-bit number and use it. that's all the newer ones are anyway Good question. Won't random 128-bit numbers actually be superior to UUIDs in every way except predictability?
- mholt 3y agoSorry to be this guy, but did you read the article? Only UUIDv4 is "just generating a random 128-bit number" (almost) -- and there are valid reasons that's not a good choice. Which the article explains. :) Like database insertion performance. Sequential IDs have benefits.
- Traubenfuchs 3y agoAm I the only one instinctively upset by the communication/bandwith/storage overhead of the dashes as well as the version and variant bits of UUIDs? It might be insignificant, but to me it makes UUID feel tainted, dirty. 11.1% of a UUID are dashes. 15.3% of a UUID are wasted bits if you count version and variant bits. Anecdote: I worked for a company that used numeric primary ids internally and externally and increased the primary key by TWO to THREE for each new customer to make it appear to the outside world we had twice to three times the rate of customer growth.
- throwawayb4dc65 3y agoThe dashes are just there to help separate the groups visually. You can write code to remove/add the dashes if you want to shorten the URL. If you use the uuid data type when storing it in a database, it only uses 128 bits.
- rossant 3y agoThe dashes do not take up any space, they are not encoded. I think it's fairly common to reserve some space to versions and ECC in uuids, packets etc. That's a price to pay to avoid many bugs and compatibility issues.
- Traubenfuchs 3y agoUUIDs are often pushed around in JSON and as string types, including the none random parts.
- BeFlatXIII 3y agoDo those dashes cause any meaningful performance impact, or is this just a microoptimization obsession?
- coolgoose 3y agoI am confused how this is new. UUIDv1 is time based, you just need to be careful about entropy, and in MySQL 8 you can for a longish time use it as an ordered field.
- 8organicbits 3y agoThe use of a MAC address and fine grained timestamp are challenges of UUIDv1. https://blog.devgenius.io/analyzing-new-unique-identifier-formats-uuidv6-uuidv7-and-uuidv8-d6cc5cd7391a?gi=e90b435d8f84 https://blog.devgenius.io/analyzing-new-unique-identifier-fo...
- hknmtt 3y agoI have been using ULID for years. Using digits now would feel very strange.
- tzahifadida 3y agoTo me it sounds like a corner case. Example: a) UUID4, CreatedTime/UpdatedTime. b) Bigint, CreatedTime/UpdatedTime. c) UUID7 internal (which also includes time badly), UUID4 external/whatever short ID. How exactly this helps if you need external ids (which you usually do today)? It doesn't even make it a short ID. Even if there is a corner case, are we just saving a few bytes while adding more complication? Clustered Index is a myth in PostgreSQL, not practical since you have to run a special program to reorder. So, a regular index might suffer but not really. Why? Because I am not ordering by the ID most of the time, I am ordering by "Created Date/Updated Date" or Name or whatever. Who cares about ordering IDs? WAIT!!! But what about Next Tokens? ok, these are painful, but easily solved: Next can be (>=Created Date,>ID). Same result. Pagination, stays the same since it is sorted by Created Date.
- deleted 3y ago[deleted]
- otherme123 3y agoI understood it as c) only UUID7, no secondary external UUID. The external Id is used instead of Bigint because you don't want your external users to query 1, then 2, then 3 (IDOR)... But the random part of the Uuid7 makes this impossible. Uuid7 isn't a substitute for Created/Updated, but a substitute for the dual field Uuid4/Bigint.
- miiiiiike 3y agoThis is neat. I've been using a custom snowflake cluster for years. Having this in the language/DB would be great for smaller projects. For bigger/public projects I'd like to be able to add a sequence, node, and data center id to the UUID too.
- dajonker 3y agoSimilar to the old situation in the article, we are using sequential 64 bit primary keys, but we use an additional random 64 bit key for external usage (instead of 128 bit). The external key is base64 encoded for use in URLs which results in an 11 byte string. This hides any information about the size of the data, the creation date of customer accounts (which would be sort of visible with UUIDv7) and prevents anyone from attempting to enumerate data by changing the integer in URLs. I thought about using UUIDs as external keys but the only compelling use case seems to be the ability to generate keys from many decoupled sources that have to be merged later. 64 bit should be enough for most things https://youtu.be/gocwRvLhDf8?si=QBheJCG21bAAV0Z7 https://youtu.be/gocwRvLhDf8?si=QBheJCG21bAAV0Z7
- Waterluvian 3y agoIt sounds like you basically just made your own 64 bit UUID. If you’re exposing this ID for manual use by a human (like URLs) then that sounds pretty helpful to be shorter!
- Kunix 3y agoI am using a variant of SnowflakeId [^1] in order to have 64 bit keys too. It's similar to UUIDv7 (it leaks the creation time), but it's not an issue for me. So I am able to have a single 64 bit key, which can easily be formatted into a small string for user-facing urls. [^1]: https://instagram-engineering.com/sharding-ids-at-instagram-1cf5a71e5a5c https://instagram-engineering.com/sharding-ids-at-instagram-...
- RhodesianHunter 3y agoWhat's the risk of collisions with your external ID in this scenario?
- sealeck 3y agoI would imagine that they enforce this using e.g. a unique constraint in their database.
- 3y ago
- JCharante 3y agoI wonder who this article is written for. Who would be reading about UUIDs but not know about cache hit rates? > As a result, retrieving the most recent data from a large dataset will require traversing a large number of database index pages, leading to a poor cache hit ratio (how many requests a cache is able to fill successfully, compared to how many requests it receives).
- markcollin 3y agoInteresting - have beem using uuidv4 for a long time. Will explore further on uuidv7
- jsf01 3y agoHow long will it be before the “milliseconds since epoch” part of the uuid overflows or repeats?
- jolmg 3y ago$ date -ud @$(( 256 ** 6 / 1000 )) Tue Aug 2 05:31:50 AM UTC 10889
- birracerveza 3y agoWell, at least it's not a Friday.
- jug 3y agoHaha this is what we came up with for our home brewn unique ID's in a GIS application since decades ago. For the same reasons.
- jimmySixDOF 3y agoDiscussion here a couple months ago : Analyzing New Unique Identifier Formats (UUIDv6, UUIDv7, and UUIDv8) (2022) https://news.ycombinator.com/item?id=36438367 https://news.ycombinator.com/item?id=36438367
- zooFox 3y agoOne benefit of an epoch is that it's easily readable (or comparable, at the very least). I am not sure I can read epoch in hexadecimal format though.
- okl 3y agoNeed a new clock? https://retr0.id/stuff/2038/ https://retr0.id/stuff/2038/
- danwee 3y agoSo, how do you guys use UUIDs for real? I worked in a company in which they were using UUIDs in Mongo, and of the most painful things were implementing API endpoints that filter resources. Imagine you have an endpoint in which you are filtering by resources A, B, C and D. Ideally you would end up with something like this: GET /filter?a_id=X&b_id=Y&c_id=Z&d_id=w But in practice we were using POST and passing the ids in the body payload. Why Because my old team said "the UUIDs are long, so we may reach the maximum URL length if we pass them as parameters". I didn't like it, and I still don't like it at all.
- dgb23 3y agoThe maximum url length is typically quite long. Another thing is that you don’t necessarily need to encode uuids canonically. They are just u128’s. It’s relatively straightforward to find a url friendly string representation that is shorter.
- danwee 3y ago> It’s relatively straightforward to find a url friendly string representation that is shorter. Are we talking about shortening the whole URL or shortening specific UUIDs? If the latter then I imagine one would still need to keep track of the mapping UUID <-> shorten version, somewhere, right? If so, why not just add yet another field/column for an old good numeric integer that can be used for filtering? Would that work?
- Nevermark 3y agoI think the point is for URLs you can consistently represent the 128 bits of UUIDs with shorter strings by using a higher base. I.e. more characters. So a nonstandard but isomorphic shorter string representation.
- RobIII 3y agoYou could simply base32, base36 or base64 encode the UUID's (as an example). Other than that; in IE6 times URL's had a limitation of 1K or 4K or something around that lenght IIRC, so unless you're using dozens of UUID's in a URL this hasn't been a problem for ages.
- wvh 3y agoA few years back, I wrote some code that generates a sortable 128-bit UUID-like identifier starting with a milliseconds-since-epoch timestamp, a node number and a random byte tail. It has been working fine in Postgresql, using its builtin UUID type. I suppose downstream system have been using the string representation though. The main reason for going such an identifier was being able to generate them from different, non-centralised places. A nice side effect is that you can't accidentally get an erroneous ID that happens to work the way you can with a sequential integer primary key. For another project, I've also used sortable 64-bit snowflake-like identifiers; they have the added benefit of being able to use 64-bit integer representation in code and database identifiers, even if you might want to externally represent them in base58 or similar encoding. The original UUID types aren't as useful as they once were, so it'd be worth writing a new RFC and extending those original types.
- ahoka 3y agoIsn’t that almost the same as v1?
- xarope 3y agoI am just about to wrap up some prototyping comparing snowflake, typeids, uuidv4 and ulid. Why did I not bump into uuidv7 earlier?!?
- RobIII 3y agoDon't know, because uuidv7 has been coming for ages... https://www.ietf.org/archive/id/draft-peabody-dispatch-new-uuid-format-01.html https://www.ietf.org/archive/id/draft-peabody-dispatch-new-u...
- dgb23 3y agoAs a beginner I treated and understood (SQL) databases as something I have to use in order to store stuff. Later I was excited about the power and expressiveness of SQL and its extensions. There is a ton of leverage and you can make it so that interfacing with it directly becomes much more useful. However now I’m in a different phase. I see it as a durable data structure. I think in terms of “what does it provide to make the overall system better?” The issues around indexing and uuids that is discussed in the article fits nicely into this line of thinking. In web development, database access and performance often dominates and infects the whole system.
- foreigner 3y agoAs a beginner I thought of the database as a backend for the app. Now I think of the app as a frontend for the database. :-D
- kuchenbecker 3y agoCrud, you're right!
- eviks 3y agoAlways wondered what the point of dash-separating uuid if the separated parts are unreadable anyway just like in this version, just makes it harder to select as a single blob of text
- T3RMINATED 3y ago[dead]
- toni88x 3y agoIMO the benefit of UUIDs over integers is that they can be generated client-side without clashing. But you cannot trust timestamps generated by clients and therefore the order. So what is the benefit over UUID4?
- XCSme 3y agoWould using the timestamp in the UUID be possible for date-range queries?
- frederikb 3y agoFor me the central benefit is that you can create them in a distributed manner and are not reliant on a central system as a single source of truth for creating your identifiers. I can therefore easily generate a new UUID in a trusted backend service which just accepts the command received from the untrusted client and then forwards the request for asynchronous processing while returning the UUID to the client. This is a typical architecture and the only change is that I can now create UUIDs which may have performance benefits, depending on the data storage technology of my read models. If you need to create the UUIDs on the client side to support specific requirements such as offline-first, then I would indeed consider adding some reconciliation which replaces the IDs provided by the client-side by new ones generated by a trusted component as soon as synchronizing takes place.
- frederikb 3y agoIn any case regardless of UUIDv4, v7 or any other format you should not allow the untrusted client to determine the real ID - as long as there is at least one trusted component in the architecture which would take over this role. This should help eliminate a whole set of possible security issues.
- jakewins 3y agoA useful/horrifying pattern on this topic: you can use UUIDv1 as a prefixed id, giving you a way to generate tagged IDs in a system that uses UUIDs. You set the node field to a broadcast MAC address, and use that as a namespace/prefix. This inches close to the boundary of the RFC, but is arguably compliant. As an example, you may generate demo or “canary” data items that are UUIDv1s with a well known node field, which then lets you do distributed “isDemoData()” checks by just looking at the UUID.
- maxsupport 3y ago[dead]
- tzahifadida 3y agoNot specifically the topic, but I looked for a library for golang and it is not that common, there is a library in <20 stars, too experimental for me. Also, not sure the postgresql extension is in the main distribution, couldn't find it if it does. For example, GCP only supports this one IIUC https://www.postgresql.org/docs/current/uuid-ossp.html https://www.postgresql.org/docs/current/uuid-ossp.html Java has something, but again not really clear how tested. So using this is a bit iffy...
- insanitybit 3y ago> The nature of Buildkite's products mean recent data is accessed more frequently than old data. With non-sequential identifiers, the most recent data will be randomly dispersed within an index and lack clustering I would assume that `serial` would solve this problem too.
- phkahler 3y agoIs there some reason new versions of UUID keep appearing? It seems like the desired properties are never quite achieved so new ones appear later. Is there a table with UUID version across the top and characteristics down the side, so I can see the differences and pick one that fits my needs? That might also help to explain why there are so many variants.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- bhouston 3y agohttps://en.wikipedia.org/wiki/Universally_unique_identifier#Versions https://en.wikipedia.org/wiki/Universally_unique_identifier#...
- ricardobeat 3y agoUnless you have specific needs, the only type of UUID you should care about is v4. v1: mac address + time + random v4: completely random v5: input + seed (consistent, derived from input) v7: time + random (distributed sortable ids)
- kozak 3y agoAs someone who only cares about v4, I periodically wonder why don't I just use fully random 128-bit identifiers instead (without the version information).
- klysm 3y agoBecause the rest of the world uses uuidv4 and the extra couple bits doesn’t really buy you anything
- hansvm 3y agoThat'd be perfectly fine. You need to take some care to avoid duplicates or patterns though. Databases might not come with crypto-random out of the box but usually do have UUID support. Similarly, you need to use the crypto-random routines in your favorite standard library. Depending on the implementation you still might have to worry about seeding issues. That's probably moot though since the UUID library would probably be compromised by something like that under the hood too.
- gwbas1c 3y agoAnyone ever try encrypting a database ID (IE, a sequential int,) and use that as a public key? IE, take a 32 or 64 bit int that's the primary key, encrypt it, and then use that as the public ID in a web application, URL, API, ect.
- pknerd 3y agoSpeaking of RDBMs, how good are UUIDs when making joins and fetching a certain record?