4 ms·
Forgive my ignorance. I'm more of a social scientist than a programmer. Questions: 1. Why not go for 16-character strings (instead of 26 or 36), with each char
by Solar19 8y ago
Forgive my ignorance. I'm more of a social scientist than a programmer. Questions:
1. Why not go for 16-character strings (instead of 26 or 36), with each character representing 8 bits?
Sure, you'd need 256 possible characters, but it's almost 2019 and Unicode has been with us for decades now. Surely we could be more cosmopolitan than Americentric ASCII and curate 256 characters for an 8-bit encoding?
With a 16-byte string, we could compare and process strings much faster, particularly with SIMD instructions like Intel/AMD's SSE 4.2 string comparison instructions. They're optimized for 16-byte strings and were introduced many years ago in the Nehalem architecture. That's a couple of generations before Sandy Bridge, so any server today is going to support it.
2. What does it mean to be "user-friendly" when it comes to these sorts of IDs? What are some scenarios where users interact with them or communicate or share them with someone or some authority? Crockford wanted his 32 character set to be easy to convey on a telephone, which seems like an expiring use case today. It seems like we should be able to use all sorts of non-ASCII characters now, without resorting to the Unicode Klingon or Tengwar blocks. Do we really need to be able to pronounce them all like Crockford anticipated?
NOTE: Unicode characters beyond the Basic Latin block take two or more bytes each, so we wouldn't be able to use them encoded as Unicode. What I'm advocating is a 256 character set with each character encoded in one byte, strictly for the purposes of generating these sorts of unique IDs represented by compact 16-character strings. Call it Duarte's Base256. All these other BaseN systems seem orthogonal to character encodings, or they just assume ASCII. I guess my idea would require both a character set and an encoding scheme. The latter would be similar to ISO/IEC 8859-15 and Windows 1252, but more complete with 256 printable characters. A lot of them could probably be emoji.
How good or terrible is this idea?
- zeroimpl 8y agoThe string representation is for display purposes/information exchange only. I'm sure most implementations would internally store the data in a 16-byte form (eg the C implementation uses __uint128_t), where the data is essentially a 128-bit number. Given that, inventing a new character set seems pointless, since you'd compare the data using 128-bit binary operations already anyways (as opposed to lexicographical string comparisons). Which leads to the question - how is ULID different from UUID in practice?
- dragonwriter 8y agoThe difference is that time is at the front, so if you need to sort by milliseconds the ID was created, you can.
- Solar19 8y agoThanks, that helps. By the way, why are they time stamped at the beginning, or at all?