5 ms·
When do you use UUIDs and why? I use traditional auto increment integer IDs as primary keys in SQL databases. I use those SQL query results in application code
by castell 11y ago
When do you use UUIDs and why?
I use traditional auto increment integer IDs as primary keys in SQL databases. I use those SQL query results in application code. Debugging seems a lot easier with smaller integer values rather than UUIDs that I saw in several enterprise software and SQL-DBs.
- the_mitsuhiko 11y ago> When do you use UUIDs and why? When you need the client to determine the PK before an operation happens. In particularly necessary when working with distributed environments. There are many different forms of UUID and when you know what you are doing you can build very powerful systems with them that you cannot do with auto incrementing integer keys (even with holes).
- vox_mollis 11y agoIn distributed systems, you can easily deal with this by assigning each client node an ordinal, and the client generates a sequential id that is a multiple of the highest ordinal of the ring added to the client's local ordinal. You can synchronize ordinal values between nodes with something like zookeeper or etcd.
- lostcolony 11y agoIf you're having to synchronize it using a strongly consistent mechanism, couldn't you just use an auto-incrementing integer anyway (not meaning to be snarky here - I legitimately don't know)? The benefit of a UUID is you don't need strong consistency; even on a partition of < n+1 nodes you can still generate IDs, and even if the network is unpartitioned, you can generate an ID without making a roundtrip. I would also debate your use of the word 'easily'; what's easier, setting up zookeeper/etcd, or just generating a UUID? But yes, if your use case is "I want to make sure I have non-conflicting IDs across my cluster", synchronizing them is a possible solution. And the right one, if your requirements stipulate absolute ordering based on the generation of each ID.
- jerven 11y agoYou would need to synchronize less often and that can be big benefit performance wise. i.e. you can tolerate more partitions without effects. If done right and in some cases you can have the same benefits as UUIDs with smaller keysizes. After all UUIDs only make it unlikely that you have conflicting IDs not impossible. Especially if your random source is not that random.
- jacquesm 11y agoAuto-incrementing using some kind of centralized store is synchronous, UUID generation is asynchronous.
- jacques_chester 11y agoIn which case, you could equally use v1 or v5 UUIDs. Plus a UUID can be generated by a source outside your network -- eg, client side, with the ability to treat them as quasi-nonces for replay detection. Or inside your network, during long-running partitions. The bigger question is: why would I install, manage, update, monitor zookeeper/etcd and write my own custom ID generator and all the storage mechanisms around it, when I can just use UUIDs?
- vox_mollis 11y agoRight, but I'm addressing the parent case in which ordered id is required (for whatever use case)
- jacques_chester 11y agov1 UUIDs are ordered.
- the_mitsuhiko 11y agoOur clients come and go and are largely unknown to us. Even in that case what you are describing is UUID1.
- lostcolony 11y agoIf your APIs have user specific endpoints, use a UUID rather than an integer in the API (what your DB uses is a separate question, see below). While it should not matter (because your authentication and authorization mechanisms are good, right?), merely making it so a user can't easily guess what endpoints correspond to another user is an additional protection. For the database itself, the benefits are simply those that come from the fact that every ID, regardless of where it was generated, should be unique. This means you can, per the_mitsuhiko, generate keys in a distributed fashion, but it means more than that. Even rows from disparate databases have unique identifiers, which has benefits all the way down the line, as everyone has a way to refer to a particular entry that uniquely identifies it, regardless of where it came from. It lets you separate, and recombine the data, without worrying (depending on the type of UUID, at least) about collisions (separating being something like a NoSQL database, where based on the hashing key it gets written to one nodeset, allowing you to scale out; recombining being something like that, or you had multiple regions/databases in a SQL solution, but you need to run queries against both and merge the results for insertion into a single database for various metrics or similar). Whether you need any of that is entirely dependent on your use cases; an incrementing ID in the database is perfectly valid for many needs.
- TimJYoung 11y agoYou are correct, UUIDs are a pain because they're large and hard to remember/compare when debugging. But, they can be useful for internal functionality in DB engines also. One of our database engine products uses them with the manifests attached to replicated updates. A UUID uniquely identifies each table being replicated (assigned when a table is "published"), and the manifest contains the list of UUIDs that have loaded the update. That way the engine knows whether a particular table can safely ignore the update, which is important with bi-directional replication.
- lmm 11y agoI use UUIDs for pretty much any "primary key"-like scenario: * They're natively supported in most databases and languages * They're strongly typed, meaning there's no risk of accidentally doing things that make no semantic sense with them (e.g. adding or multiplying) * They ensure that any API user will use the correct datatype, rather than some clients breaking when your ids go above 2^31 * They avoid exposing information about how many entries they are, e.g. you can't tell how many users I have by signing up and checking your user ID * The client can generate them without roundtripping to the database; this can save on roundtrips when you're saving several related pieces of data, and makes it easier to have circular datastructures if you need them * As others have said, they're usable in an AP datastore None of this is impossible to do with integers, but UUIDs make it very easy.
- sinatra 11y agoDoesn't the random nature of uuids make them extremely inefficient candidates for primary key as they can't be indexed well?
- lmm 11y agoIt's not like an auto increment key is meaningful, so the read performance will be the same either way. As for writes, a new key could go anywhere in the index, but OTOH you can insert multiple values concurrently (unlike with auto increment keys with MVCC where you'll have collisions and one insert will have to be rolled back and redone). As always, benchmark your use case and see what gives acceptable performance.
- bza 11y ago> As for writes, a new key could go anywhere in the index This forces index tree rebalancing to occur on many (even most) writes, which is hugely detrimental to performance.
- lmm 11y agoWhich tree structure is this for? Many tree structures (e.g. the classic red-black tree) perform much better (doing less rebalancing) for randomized inserts than for ordered ones.
- elchief 11y agoI use Spring Security and it uses UUIDs for CSRF tokens. Luckily, Java does this securely.
- perlgeek 11y agoAnother use case for UUIDs are message IDs. Not necessarily for emails, but for other messaging systems that must be able to refer to other messages (like in-reply-to), but without assuming there's a central database that stores every message.
- ddebernardy 11y agoAnd you're correct for most cases. But if you need to keep a master DB and local DBs on mobile devices without internet connectivity in sync, you'll be thankful to be able to generate v4 UUIDs and assume that the collision rate is nil for all practical purposes.
- Lazare 11y agoInteger IDs are fine for a non-distributed app where you have a single DB and every app server is directly connected to it. The second that's no longer the case, integer IDs become a nightmare. On the other hand, a well chosen UUID, can be generated anywhere and be relied on to be globally unique. So maybe your client wants to create some records on the client, sync to a local db, and then do a batched update the next time the laptop it's running on has a wifi connection. Or maybe you want to scale your databases horizontally, or distribute your app across multiple data centers, or basically make any sort of AP (ie, Available and Partition Tolerant) system. In which case, integer IDs are asking for a world of pain, but a well chosen token is great, because you can keep working efficiently when talking to your Single Source Of Integer Truth is slow or impossible. (And yes, there's other solutions to the issue. But UUIDs can solve it with a minimal amount of design and coding.) > Debugging seems a lot easier with smaller integer values rather than UUIDs My experience has been that both are just as easy, and leaking integer IDs can be a (mild) security vulnerability.