3 ms·
is there something shorter than UUID i hate how long it is something like youtube URLs but guaranteed to be without duplicates
by pajeets 2y ago
is there something shorter than UUID
i hate how long it is
something like youtube URLs but guaranteed to be without duplicates
- asperous 2y agoOne advantage of uuids is they can be generated on several distributed systems without having to check with each other that they are unique. Only long ids make this reliable. Youtube ids are random and short, but youtube has to check they are unique when generating them. Maybe one way is to split up a random assignment space and assign to each distributed node, but that would be more complex.
- imron 2y agoAnd then there’s uuid5 which you can use to generate identical unique identifiers across multiple systems without having to check on each other. Very very useful to have in some circumstances.
- wtetzner 2y agoEven UUIDs are not guaranteed to not have duplicates. It's just extremely unlikely, largely due to their length.
- bigiain 2y agoDifferent use cases have differing requirements for uniqueness though. A lot of stuff doesn't need "you'd have to generate 1 billion v4 UUIDs per second for 85 years to have a 50% chance of a single collision." sort of guarantee. You might think your "Uber, but for short term giraffe rental" startup needs that sort of guarantee in the investor demo prototype, but it doesn't. Just use an auto increment in Postgres or MySQL (or an integer primary key column in SQLite). If you fool those investors into pouring the money pipe over you, your first real technical/senior engineer hire is gonna throw all the code you have running on your laptop away anyway. Maybe someone at Meta sat down once and figured "we have 4 billion users, each of who on average has 3 cats that they take a dozen picture of every day, so about 150 billion cat pictures per day. So if we name them using uuids we're still good for almost 50 thousand years before we have a 50% chance of displaying a pic of Fluffy when we should have displayed a pic of Mr Whiskers". Then they promptly ignored the problem (or fixed Zack's code that was using php's hash("md2", $query['catname']) ).
- wtetzner 2y agoIt's not just a scale thing though. It also matters how problematic a collision will be. But yes, if you're just using them as primary keys in a database you're probably fine with auto increment for most use cases.
- wongarsu 2y agoIf you are fine with creating IDs in a centralized way (as you would do in 99% of cases anyways) you can just use a normal incrementing integer primary key. Then encrypt it with XTEA (either at your API boundary or in the database) to get non-sequential unguessable 64 bit keys. [1] has example code for postgres. If the original key don't have duplicates then the XTEA encrypted keys don't have duplicates either. Then just encode it in a format of your choosing. Youtube uses a modified base64 encoding (no padding, and + and / are replaced by - and _). And youtube video ids seem to also be 64 bits, just like xtea output. 1: https://wiki.postgresql.org/wiki/XTEA_(crypt_64_bits) https://wiki.postgresql.org/wiki/XTEA_(crypt_64_bits)
- quibono 2y agoAre there any edge cases or things to be aware of when using this, or is it pretty much plug&play? I'm thinking of using this in one of my projects.
- deleted 2y ago[deleted]
- morepork 2y agoThe risk with an incrementing integer is that your database falls over then you can lose some IDs depending on how often you take backups. After you restore you need to make sure that whatever integer you start at is greater than the highest integer issued beforehand, which may be different to the highest integer in your restored DB. Otherwise you can have clients with the same ID.
- pclmulqdq 2y agoAn atomic counter of some sort solves the problem of UUIDs with the cost of a synchronization step (although this can sort of be minimized). UUID and its variants are long to avoid duplicates without having to synchronize.
- syncsynchalt 2y agoIf you're not distributed, use an incrementing integer. If you're distributed, look into vector clocks[1] or snowflake[2][3] [1] https://en.wikipedia.org/wiki/Vector_clock https://en.wikipedia.org/wiki/Vector_clock [2] https://github.com/twitter-archive/snowflake/tree/snowflake-2010 https://github.com/twitter-archive/snowflake/tree/snowflake-... [3] https://en.wikipedia.org/wiki/Snowflake_ID https://en.wikipedia.org/wiki/Snowflake_ID