4 ms·
This removes '1' but not 'I'. But if you want to avoid ambiguity, you need to remove _both_ characters, not just one. Users don't know your IDs never contain a
by jamesfisher 4y ago
This removes '1' but not 'I'. But if you want to avoid ambiguity, you need to remove _both_ characters, not just one. Users don't know your IDs never contain a '1', so when they see 'I', they'll still think, "hmm, is that a 1?"
- dahfizz 4y agoIf you are taking base32 as user input, you can translate '1' to 'I' so that the input is correct regardless of which the user types. But, needing users to read / type a base32 ID is always bad UX anyway.
- iainmerrick 4y agoYeah, best to avoid having to type in a big nonsense string in the first place. But if you do need to do it, base32 is at least better than the obvious alternatives (decimal, hex, base64)
- iainmerrick 4y agoIt’s much clearer with lower-case “i” (which this article uses). Same situation with “o”.
- KMag 4y agoCrockford's version of base32 [0] (sometimes called crockford32) treats 0, o, and O the same when decoding, likewise for 1, i, I, l and L. It also doesn't use the letter U, to reduce the chances of including English obscenities. So, you don't need to remove ambiguities if you can collapse the ambiguities into equivalence classes. If you're going to use some flavor of base32, crockford32 is hard to argue against. I could bikeshed some alternatives (I would use the letter U and include Q in the 0 equivalence class, and not worry too much about obscenities), but of the somewhat common base32 standards, crockford32 is the best, IMNSHO. [0] https://www.crockford.com/base32.html https://www.crockford.com/base32.html
- layer8 4y ago> you don't need to remove ambiguities if you can collapse the ambiguities into equivalence classes. It still leaves users who don’t know or can’t assume that normalization is applied worrying about which character exactly is being displayed.
- yellowapple 4y agoThe users don't need to know or assume anything if the decoder accepts both characters as aliases for one another. The issue is further addressable by being explicit about the normalization being in effect and/or by being explicit about which one the encoder is emitting ("if it looks like a zero, then it's always a zero").
- layer8 4y agoMy point is that most developers implementing this won’t consider that. This is just one example of a more general phenomenon of making some UI behavior foolproof, but not considering that the user doesn’t know that it’s foolproof, so the user isn’t saved from wondering about what will happen and what is the correct input. In the present case, instead of having to inform the user about the normalization, it would be simpler to refrain from using potentially ambiguous characters in the first place.
- yellowapple 4y agoOn that note, it's probably worth shamlessly plugging Base32H, which is my own take on base-32 inspired heavily by Crockford's: https://base32h.github.io/ https://base32h.github.io/ Base32H has some advantages and disadvantages compared to Crockford's (see: https://base32h.github.io/comparisons#crockfords-base32 https://base32h.github.io/comparisons#crockfords-base32 ); long story short: L/l is its own digit, 5/S/s are merged, and U/u/V/v are merged. In my totally-not-biased-at-all opinion, the advantages outweigh the disadvantages enough to have warranted creating Base32H instead of just using Crockford's. If I was willing to break from duotrigesimal, I'd probably merge 0/O/o/Q/q, 1/I/i/L/l, 2/Z/z, 3/E/e, 6/G/g, 7/T/t, 8/B/b, and 9/P/p. The resulting Base24H would hinder readability of encoded words, and would break the alignment to whole bits, but it's probably close to the optimal intersection of information density, unambiguity, and convenience.