5 ms·
I wish there's a standard for short UUID, like `73WakrfVbNJBaAmhQtEeDv` or `bK7nP9xM`. I mean, it's not UUID cause it can be duplicated somewhere, I just want a
by hamasho 2y ago
I wish there's a standard for short UUID, like `73WakrfVbNJBaAmhQtEeDv` or `bK7nP9xM`. I mean, it's not UUID cause it can be duplicated somewhere, I just want an ID standart combination of random and short enough to remember.
- kyzer-davis 2y agoWe are working on standardizing an alternate encoding technique for the 128 bit UUID so it can have a shorter text form. You can the discussions here: https://github.com/uuid6/new-uuid-encoding-techniques-ietf-draft https://github.com/uuid6/new-uuid-encoding-techniques-ietf-d...
- jfdjkfdhjds 2y agojust use creation timedate plus auto increment int. and then a small hash with base64 or 37 or whatever is in vogue these days. thats what old timers used before uuid 1. guess we should guerilla standardize something like this as uuid-0 or uuid-deprecated-2.0 for keeping up with the spirit.
- selcuka 2y ago> just use creation timedate plus auto increment int. The problem with auto-increment integer ids is they are not always possible with distributed systems.
- sgarland 2y agoThey are, actually, you just have to coordinate the ranges each node has.
- selcuka 2y agoWhat if you don't know how many nodes you have? UUIDs can also be generated on the client side (in cases where you can trust the client).
- sgarland 2y ago> UUIDs can also be generated on the client side (in cases where you can trust the client). I'm fairly certain the first rule of websec is you never trust the client. I definitely would not trust a user's browser to directly insert a value into a DB. > What if you don't know how many nodes you have? Shouldn't matter; you have a centralized system that hands out chunks of IDs on-demand (and has its own mechanism to ensure no repeats). This is similar to what Vitess [0] does. [0]: https://vitess.io/docs/20.0/reference/features/vitess-sequences/ https://vitess.io/docs/20.0/reference/features/vitess-sequen...
- selcuka 2y ago> I'm fairly certain the first rule of websec is you never trust the client. Not every piece of information is confidential in every system. Sometimes a UUID is just that, a UUID. > you have a centralized system that hands out chunks of IDs on-demand I don't follow. If your system requires a central node that can reliably generate unique auto-incrementing integer IDs, why bother with UUIDs at all? Just base-64 encode the integer ID, or hash it with a salt to protect against enumeration attacks, if you want. If you don't want the dependency to a centralised system, just use UUIDv7, which is just a timestamp plus random bits, or implement a shorter version of it. There is no need to overengineer.
- sgarland 2y ago> I don't follow. If your system requires a central node that can reliably generate unique auto-incrementing integer IDs, why bother with UUIDs at all? I also don’t follow. I thought your initial assertion was that auto-incrementing integer IDs weren’t always possible, thus the need for UUIDs. Monotonic ints, or more broadly anything k-sortable, are generally optimal for RDBMS indices due to most indices being B+trees. That’s why there’s such enormous effort towards NOT using UUIDv4. > just use UUIDv7 Indeed; this is my recommendation when devs insist they can’t possibly use integers. Personally, I maintain that most places can use ints, it’s just that they’ve hideously over-complicated things to the point that it would be far too much work.
- deleted 2y ago[deleted]
- bigiain 2y agoI've used schemes like "concatenate a shared secret, millisecond resolution times, local autoinc ID, and some sort of distributed machine identifier (like ip address or MAC address), then taken a truncated hash of that with as many bits as needed for the desired uniqueness guarantees." I wouldn't use it for assigning bank account numbers, but for most web or app stuff it's fine.
- gregmac 2y agoThe closest that comes to minds is ULID[0]. It is short (26 character base32), 128 bit and lexicographically sortable. I think the reason there's no other popular standard is you give up something. 128 bit gives a pretty low risk of collisions in almost all uses, but as you go smaller you start having to consider the specific scenario and impact, etc, which doesn't work well for a standard. You could use another encoding (eg base64 or base85) to get it shorter, but you start sacrificing other things (case sensitivity, url-safeness) - again, not great for a standard. [0] https://github.com/ulid/spec https://github.com/ulid/spec
- djbusby 2y agoAnd ULID is translatable to UUID. So, eg, use ULID on display, links, etc and UUID data type in the language/DB
- notpushkin 2y agoFor my own project, I went with base32-encoded UUIDv7, prefixed with type name (a la Stripe ids). Compared to ULID, it's still lexicopraphically sortable, but is backed by an actual standard, so a bit more sound IMO. UUIDv7 will only work until year 4147 (compared with ULID's 10889AD), but by then I think we'll have another UUID version we can switch to. Here's my implementation in Python: https://codeberg.org/prettyid/python https://codeberg.org/prettyid/python, https://pypi.org/project/prettyid https://pypi.org/project/prettyid And a rudimentary TypeScript library: https://codeberg.org/prettyid/js https://codeberg.org/prettyid/js, https://npm.im/prettyid https://npm.im/prettyid
- tommy_axle 2y agoNot a standard per se but nanoid seems to fit the bill. Widely implemented.
- geitir 2y agoGit uses SHA and then dynamically set the number of characters to use based on repository size. You could do something like this.
- pants2 2y agoSqids[1] might fit the bill for you - the IDs it produces are much shorter than UUIDs, however they're not universally unique - they're generated from an integer sequence. 1. https://sqids.org/ https://sqids.org/
- NetOpWibby 2y agoSqids looks fantastic, thanks for sharing!
- slivanes 2y agoA feature of a Sqid library I've used is that it can pad the value out to a minimum set of characters, so even an internal id of 1 can look substantial. https://github.com/sqids/sqids-php https://github.com/sqids/sqids-php
- andrewstuart 2y agoI was just today wanting shorter UUIDs so if you like a more compact/short UUID you can convert them like so to url safe base64. It's the same UUID just in 22 character form and can be converted back. It's n ot really a conversion because a UUID is just a 128 bit value so its an alternative representation. 483971cf-aad7-4c84-abf1-4a94c9d72f99 -> SDlxz6rXTISr8UqUydcvmQ (length: 22) fb67926f-3cfb-486c-a7da-30662147a20b -> A2eSbzz7SGyn2jBmIUeiCw (length: 22) 799069a9-b32a-415f-b689-a8cc3f51bfa4 -> eZBpqbMqQVA2iajMP1GBpA (length: 22) 8161ee0b-f7a5-4b32-95ea-9b9efe94e5f2 -> gWHuCBelSzKV6pueBpTl8g (length: 22) b1ea416c-f209-43cb-bfaf-d9cf6229459e -> sepBbPIJQ8uBr9nPYilFng (length: 22) ee70989a-b614-4665-9881-41054544c313 -> 7nCYmrYURmWYgUEFRUTDEw (length: 22) cce06fe2-b64f-47bc-a91a-d3dfd343e1e5 -> zOBv4rZPR7ypGtPf00Ph5Q (length: 22) aea3de6e-e769-4c8d-ba2d-77922d227176 -> rqPebudpTI26LXeSLSJxdg (length: 22) import uuid import base64 def make_short_uuid(data): encoded = base64.urlsafe_b64encode(data).rstrip(b'=').decode('utf-8') return encoded.replace('-', 'A').replace('_', 'B') def generate_and_print_uuids(): for _ in range(8): uuid_obj = uuid.uuid4() uuid_bytes = uuid_obj.bytes print(f'{uuid_obj} -> {make_short_uuid(uuid_bytes)} (length: {len(make_short_uuid(uuid_bytes))})') generate_and_print_uuids()
- andrewstuart 2y agoThinking about it, the above is wrong because it replaces - and _ with A and B which makes it not reversible. This is reversible. You could aolso come up with a solution that has no - or _ by doing a custom base64 encode with only AZaz and 0 to 9. If you really want to not have _ or - in your short form UUIDs you could just discard the UUID when you create it if the short form includes those characters and try again. c22c1dcf-ea74-470e-acbf-b1722e243025 -> wiwdz-p0Rw6sv7FyLiQwJQ (length: 22) Reversed: c22c1dcf-ea74-470e-acbf-b1722e243025 8702aecb-6d09-4a5e-8cc8-621aada6ed96 -> hwKuy20JSl6MyGIarabtlg (length: 22) Reversed: 8702aecb-6d09-4a5e-8cc8-621aada6ed96 643a9829-9f91-4b88-80a6-db2c0eb83e8b -> ZDqYKZ-RS4iAptssDrg-iw (length: 22) Reversed: 643a9829-9f91-4b88-80a6-db2c0eb83e8b 8e7f3c1a-3d19-425e-8803-fe55a296688e -> jn88Gj0ZQl6IA_5VopZojg (length: 22) Reversed: 8e7f3c1a-3d19-425e-8803-fe55a296688e 1859f017-a5f3-4875-825a-fdd531dfac1a -> GFnwF6XzSHWCWv3VMd-sGg (length: 22) Reversed: 1859f017-a5f3-4875-825a-fdd531dfac1a 6a153b44-7fca-45b2-b13e-7f45790be7bf -> ahU7RH_KRbKxPn9FeQvnvw (length: 22) Reversed: 6a153b44-7fca-45b2-b13e-7f45790be7bf fd6bad83-a0f8-4c7f-baf1-10374be3e8e9 -> _Wutg6D4TH-68RA3S-Po6Q (length: 22) Reversed: fd6bad83-a0f8-4c7f-baf1-10374be3e8e9 cf2452d4-947b-4b92-a280-ff869e77ba65 -> zyRS1JR7S5KigP-Gnne6ZQ (length: 22) Reversed: cf2452d4-947b-4b92-a280-ff869e77ba65 import uuid import base64 def make_short_uuid(data): return base64.urlsafe_b64encode(data).rstrip(b'=').decode('utf-8') def reverse_short_uuid(short_uuid): # Add padding back to make it Base64 decodable restored = short_uuid + '=' * (-len(short_uuid) % 4) # Decode the Base64 string back to bytes return base64.urlsafe_b64decode(restored) def generate_and_print_uuids(): for _ in range(8): uuid_obj = uuid.uuid4() uuid_bytes = uuid_obj.bytes short_uuid = make_short_uuid(uuid_bytes) reversed_uuid_bytes = reverse_short_uuid(short_uuid) print(f'{uuid_obj} -> {short_uuid} (length: {len(short_uuid)})') print(f'Reversed: {uuid.UUID(bytes=reversed_uuid_bytes)}\n') generate_and_print_uuids()
- wereHamster 2y agoI usually generate N bits of randomness and base58 encode it. Choose N to your liking. You loose the benefits of monotonic sorting that is present in some UUID versions. Base58 is url safe and does not contain any special characters. And you can still store values as binary (eg. bytea in Postgres instead of a text column).
- physicles 2y agoI've also taken UUIDs and re-encoded them as base58. Works fine.
- Terr_ 2y ago> combination of random and short IMO we need to be clear on the distinction between (A) the UUID bit-generation scheme versus (B) the way it is encoded for human use/reading/transcription. They are mostly-separate problems. For example, you could have a very secure mathematical scheme, but it gets ruined by a horrible representation where each bit is written as either a capital-I, a lowercase-l, or the number 1. Conversely, could have a deeply insecure scheme that uses a nice compact serialization where everything is grouped into chunks and "1Il" confusion is not possible and there's a check-digit, etc.
- candiddevmike 2y agoNanoid? https://github.com/ai/nanoid https://github.com/ai/nanoid