13 ms·
Understanding and avoiding visually ambiguous characters in IDs
- jheriko 2y agojust use numbers and crossbar your 7s - problem gone. if someone's writing is incompetent tell them. if you can't then they ruined it for themselves by being shit at writing the number 7.
- eternityforest 2y agoI suppose the first line of defense is a QR code URL. I don't think anyone really enjoys typing long codes. After that there's ECC. A few extra bytes for a reed-solomon code will fix a lot of issues.
- octopoc 2y agoUuidExtensions[1], a C# library, has a way of generating / encoding IDs that has several useful properties: 1. IDs can be generated anywhere (client-side, server-side, etc.) and are still unique 2. IDs are ordered by time 3. IDs don't use L and O because those can be confused for other characters I've found it very handy in my travels. [1] https://github.com/stevesimmons/uuid7-csharp?tab=readme-ov-file#alternative-id25-string-format https://github.com/stevesimmons/uuid7-csharp?tab=readme-ov-f...
- TwoNineFive 2y agoOn linux you can use Theodore Ts'o pwgen tool with the -B arg. -B, --ambiguous Don't use characters that could be confused by the user when printed, such as 'l' and '1', or '0' or 'O'. This reduces the number of possible passwords significantly, and as such reduces the quality of the passwords. It may be useful for users who have bad vision, but in general use of this option is not recommended. >pwgen -B 32 oos9upoVieghuew7aeb3iev3jiequeiw acohthahpie7ae4aeboshahWiengieth yahW3qua3atheeP9jo4aiY3zeepoosh3 Noh4ooth4ohzeec4zug3ephoo7meich7 oozae9Eireix4Chaiboz9dofie4Xunof Mohj3uupee9ahngahh9on9sujee9ehae weimah9aiXeis3owaexei4uh3ibeecai PaeV7eeChaezahruNgeequoh7zok7thi eeJieyah4exiephaiPootei4dokoojoh fohhah3Eec3bah7aeR9iedah7Ve3ea7o vahs4eich4pheisoug9aiR3ohChoh7Ch eth9KaeLahdie7ahy9ohCiebohphuse9 ieye3udumaengai9ies7kae4geeque9T iesoh9eosohthoongaeroo4ehiishohY mee4ohjei4ohmika3taijei3Yaixosei ohWoo4eapid7miebee9pooKai3oofeis Eechook9quohp7se7ees9thaefahb9an aht3quooV4eiph9ap7aiw4wee7oi7eij ishep3weeh7Eero9ohdohth9MietooJ4 Kai9aich9Jee9Angeihee9eehei9esie toonaix4xe3Moob3zaic3Eesahs9ahy3 gaey9doozee7sei9quuPae3vohph4Huo ouYaephahcog3peiw7iecoo7eetheeph eeNgiezae7oongi7uena7eenaezuT7co tai9vuace9eV7Paih7ieN3Ahghiegh3v VaeteeMoobeixai9ingeyahYuzaipaht eeng7vei7pho4Ahpoa4kahgheethahz7 phas4theiThu4uqu7iCh3Aepha3shae3 ieRep3kaideeHeekiNgequieng9raeYo eegahsh9aizooshee9too9oojiox4Lei ovohcaePahM9thaebajuChoo3pipheej oowaimeiWahf4Neighoo3Eeyah3uvi4v vi4choiThei3eisohw4iP9huehohs4oe ukuchiethaquax3hieChouMahpooy4ee aegheeyeemeNeevehud9ohng3dai4jai eth3iedah9Tee3wohneisoo4aicuToos iecap7EeJ7raixiuseesiNou9ooT9fie ied3ooveingu7fu7dahdaaYe9tai7ien eijee7iKighaingaiChei7giemu4chi3 Thie3faih3ahshooRunohwoaghoh4Aev
- ceving 2y agoRecently I came up with something similar: https://gist.github.com/ceving/cb68c8f2392255c5ed4ea65a6a1993e1 https://gist.github.com/ceving/cb68c8f2392255c5ed4ea65a6a199... But I use a alphabet with 32 characters: abcdefghikmnopqrstuvwxyz23456789 I prefer 32 characters, because that makes it possible to pack 5 random bytes into a token with 8 characters.
- Terr_ 2y ago> In some cases, you might also want to avoid characters that sound similar when spoken. For example, b and p can sound similar when spoken out loud. This can be especially important in situations where IDs are communicated verbally. In many cases these kinds of IDs are just an encoding of a ground-truth that is a big integer or a sequence of bytes, and that mean we don't have to use ASCII-character granularity, we can also use words. True, that creates a certain cultural bias for wherever you get the words from, but it opens up new possibilities for error correction and detection, both by the computer and also by the humans transcribing things.
- gajus 2y agoSomewhat related, I always liked the concept of https://what3words.com/ https://what3words.com/
- simonw 2y agoThey have some pretty bad flaws in their design relating to this topic: https://twitter.com/jonty/status/1570062564523917312 https://twitter.com/jonty/status/1570062564523917312 > the actual address should be "keen.lifted.fired" instead of "keen.listed.fired" and someone clearly misheard over the phone
- Terr_ 2y agoYeah, ideally the dictionary first would undergo rather rigorous pruning based on things like phonetic similarity or how easily a typo might move between two valid words. That scoring/clustering process makes for interesting problems in their own right, especially if one throws accents into the mix.
- TheDong 2y agowhat3words has a proprietary implementation and has sent fairly silly legal threats: https://news.ycombinator.com/item?id=27020810 https://news.ycombinator.com/item?id=27020810 I'll happily boycott that for-profit company which is masquerading as a public utility, but charging money and going after anyone who reverse engineers what words are what locations. See also the comments in https://news.ycombinator.com/item?id=27058271 https://news.ycombinator.com/item?id=27058271 This is exactly the sort of thing that shouldn't be a private company, just like Lat/Lon coordinates and street addresses are effectively public domain, any suitable replacement for lat/lon should also be public domain.
- iblaine 2y agoMy OCD approves of this idea. Let’s also add, IDs cannot start with 0 or O.
- grantmnz 2y agoThis post has some overlap with work I did a while back on a "coupon code" system that is optimised for users taking a code printed on paper and entering it into a web form. A number of measures were employed to avoid/correct transcription errors. Example, docs and links here: https://www.mclean.net.nz/cpan/couponcode/ https://www.mclean.net.nz/cpan/couponcode/
- shkkmo 2y agoThis seems slightly flawed in that it completely removes all members of a similar set rather than normalizing to a single element per similar set. Thus after normalization, '1lI' would become '111'. This allows you to add seven characters back to the author's code generation alphabet without re-introducing any ambiguity.
- dools 2y agoIt only reduces the ambiguity if everyone does the same and everyone knows that you've done it.
- shkkmo 2y agoIf you control the system for generating the codes and the system for verifying the codes (which is generally the case for these kinds of codes), then nobody needs to know you've done anything. It's the same normalizing to upper/lowercase characters when you parse a non-case sensitive code.
- hananova 2y agoWhy not include '1', but make it so '|Il1' all map to the same internal value? That way you have no ambiguity while minimizing alphabet reduction.
- shkkmo 2y agoI'm not following how your suggestion. It seems like we're saying the same thing?
- wccrawford 2y agoIf you need more possible values, I agree. However, if you don't need them, I would remove them so that the user doesn't have to spend any time wondering which character it is. Even though you're processing them all after they type them and fixing them, the user has spent time and effort that they didn't need to, just picking which one it is. IIRC, I chose to keep them when I did something like this, but I don't think I thought to accept the others and convert them automatically. That project is sunset now, so it's not an issue.
- dools 2y agoI wish my parents had access to this when they chose to call me Iain Dooley. The world has almost unanimously decided my name is now Lain.
- dfc 2y agoI think that Iain is the Scottish version of Ian? Is it unacceptable to choose the alternate spelling, Ian?
- koolba 2y agoI’d considered it grossly unacceptable to change the first thing gifted to you by your parents.
- kibwen 2y agoYour name is your own first and foremost. You can honor your parents in other ways.
- quesera 2y agoEh, there's nothing magical about parental preferences. A loving parent would not want their child to live with a name that they didn't like. Fortunately with names, there are no returns, but exchanges are accepted (with a low restocking fee) in perpetuity.
- stevekemp 2y agoFunny story, I was named "Steven" and yet I've been called Steve my whole life, at my preference. Recently I went through the process of changing my name legally, because I'd fallen into a bad habit of writing "Steve" when asked for my name on some documents, but then remembering my "official" name was "Steven" on others. Having multiple IDs with different names, especially after moving to a new country, was just too much of a pain - for example my official residence permit name didn't match my passport name, which caused some fun at airports.
- digging 2y agoThe first thing gifted was life, and though that was not bestowed with consent, it's one thing I'd argue for retaining as long as possible. Everything else is fair game to discard in service of making that life a good one.
- deleted 2y ago[deleted]
- jonplackett 2y agoHow come neither v nor u are in the final set? They’re not even mentioned and don’t look like a thing else, except maybe each other in some typefaces.
- leovander 2y agovv ~ w
- jonplackett 2y agoAha! Of course!
- oh-the-irony 2y agoWorks on words and special characters, too. I just skimmed the comments and had to scroll back up to verify that I had NOT just read "Anal Of course!"
- jonplackett 2y agoHaha. You see what you want to see I guess ;)
- re 2y agoSee also Douglas Crockford's Base 32: https://www.crockford.com/base32.html https://www.crockford.com/base32.html This takes the approach of allowing ambiguous characters by decoding them to the same value, and also considers the problem of accidental obscenities.
- 38 2y agoI believe an implementation is here: https://godocs.io/encoding/base32 https://godocs.io/encoding/base32
- re 2y agoThat uses a different alphabet. https://www.rfc-editor.org/rfc/rfc4648.html#section-6 https://www.rfc-editor.org/rfc/rfc4648.html#section-6
- spintin 2y agoInteresting, I did different choices: 5-bit base-32 oi23456789 abcdefghkl mnpqrstuvw y o = 0 i = 1 j, x and z removed. I like that you can fit 6 characters in an 32-bit integer and still have to bits to spare... makes for compact usernames and network bandwidth.
- robocat 2y agoAlso avoid lowercase rn which can be mistaken for m. And avoiding vowels can help avoid offensive words within a generated code: FUKFUK9 - https://www.replacements.com/china-fukagawa-fuk9/c/27446 https://www.replacements.com/china-fukagawa-fuk9/c/27446 KUNT1 - https://id.made-in-china.com/co_gzberlin/product_Power-Steering-Gear-Rack-and-Pinion-for-Toyota-Hilux-Vigo-2WD-Kun1-2004-2008-44200-0K010-44200-0K360-44200-0K050-442000K010-442000K360-442000K050_uoenhheyny.html https://id.made-in-china.com/co_gzberlin/product_Power-Steer... base32 removes the I,O,U but other words with A,E need to be avoided too - no vowels helps avoid words in English.
- bckr 2y agoAn approach we are trying is speakable IDs. Three characters for the type of thing, then four random words from a list of clean words with 5 characters: xxx_flown-moons-deary-flake
- kibwen 2y agoYou'll want to be careful to consider homophones while also taking accents into account. E.g. if your dictionary contains "deary", it probably shouldn't also contain "dairy".
- ahazred8ta 2y agoSeveral hard-to-mess-up wordlists have been standardized. -- https://en.wikipedia.org/wiki/PGP_word_list https://en.wikipedia.org/wiki/PGP_word_list
- Izkata 2y agoThis introduces a new type of risk - if it can be interpreted as a sentence, "moons" as a verb isn't really a clean word.
- jgbmlg 2y agoI guess I better stop using Bozos_Gismos
- arp242 2y agoIt would be helpful to also add a screenshot for that font overview, because: https://imgur.com/a/h7Ks1Qj https://imgur.com/a/h7Ks1Qj And even on systems which do have these fonts, they may not always be exactly the same.
- kibwen 2y agoHonestly, stuff like this is why I stick with (case-insensitive) hexadecimal for user-facing IDs. I find hex to be the sweet spot between "decently sized alphabet to keep ID lengths down" and "easy to read, communicate, and enter manually". It's also fairly resistant to accidentally generating IDs which will offend your users (unless your users are 1337-speaking time-traveling pre-teens from 2002 who are going to snicker at "b00b5"), which is a nice perk.
- geor9e 2y agoI had this exact situation at work when they shipped millions of devices with serial numbers, and didn't leave out any letter or number. Customers had so much trouble reading them accurately, I had to make a regex script that generated every possible typ0 permutation of what the customer said, and then it would list only matches from the factory database. From there, folks would try to correlate other info like dates to figure out what their real serial number probably was. It was a nightmare. Ironically several of the digits never changed, and some were just 0 1 or 2 to represent which factory made it, so there was no need for the entire character set in the first place. They seem to have been convinced we'd produce 8 quadrillion devices.
- swores 2y ago> They seem to have been convinced we'd produce 8 quadrillion devices. While I'm not arguing that their decisions were wise nor that they shouldn't have been able to foresee and prevent the issues they caused you and your colleagues, I would add this one thought in response to the line quoted: It's often either beneficial or at least considered beneficial to prevent business information leaking through serial numbers, the simplest example being that if you start labelling your products with 1, 2, 3.. and never deviate, then it's fairly easy to take a sample of not many serial numbers and estimate how high they go and therefore how many have been sold. Sometimes it can also be beneficial to make it harder to guess a valid serial number (eg it prevent customers from pretending to have a valid one to get a refund, or whatever). Of course, even if you have these concerns and want to mitigate them, it doesn't prevent you from also taking steps to prevent difficulty reading the correct characters. If anything it should make them more aware of the potential issues you faced since it means someone is already actually thinking specifically about what system to use, as opposed to what likely happened in your case of someone spending 30 seconds going "we need serial numbers, using X digits means we'll never run out, job done".
- bluenose69 2y agoI bought some software many years back. The serial number had 6 or so digits in it. At one point, I contacted the developer for some other purpose, and pointed out that I had made my purchase as soon as I heard about the product. He told me I was the first customer, and that he had decided to make up long serial numbers to avoid this counting problem.
- btilly 2y agoThis brings up memories. One day while sick, I distracted myself from being sick by writing up a silly module to do arithmetic in arbitrary bases. And, because it was easy I stuck it on CPAN. https://metacpan.org/pod/Math::Fleximal https://metacpan.org/pod/Math::Fleximal is the module. Of all of the silly things I'd done, I would have sworn that this is the one that should never generate a support request. But it did! Why? Well I'd included a demonstration of how to turn hexadecimal into an alphanumeric code. And someone had the bright idea of using the same thing to turn long numbers into readable codes! My module worked, but I was still a bit flabbergasted that THIS wound up in production somewhere!!
- myself248 2y agoTelephone equipment avoids the letters i and o in the alphabetical designation sequence for this reason, they look like numerals 1 and 0.
- gajus 2y agoOut of curiosity, anyone knows why would this post be removed from the front page? I was excited see that the post is getting engagement. I saw it in 3 position. Then checked an hour later and it is nowhere to be seen. I am assuming this is some sort of opportunistic algorithm at play that gives a chance to a post, but removes it if it is not performing, but curious if anyone has more details.
- lifthrasiir 2y agoHN submissions tend to be in the front page when they receive a bit of early votes within (roughly) the first hour, but they disappear rather quickly without further votes. Given that this submission was only 3 hours old when you posted this comment, it is quite expectable. (For the record it's now in the fifth place, suggesting that it has eventually received enough votes to stay in the front page later.)
- gajus 2y agoMakes sense. I am just curious about the logic behind the algorithm, more than anything else.
- donavanm 2y agoEncoding should also depend on the user. Base 32 (crockford & rfc 4648) has a nice unambiguous alphabet for compact representation and explanation of why. However if your users are speaking aloud you might want a word list representation, “TIDE ITCH SLOW REIN RULE MOT”, like s/key rfc 1751. DO NOT invent your own word lists; there are an infinite number of dragons lying in wait for idioms, homophones, dialects, etc. Dont be like me and unintentionally create a major incident like “wet clam butterfly.”
- dmurray 2y ago> However if your users are speaking aloud you might want a word list representation, “TIDE ITCH SLOW REIN RULE MOT”, like s/key rfc 1751. DO NOT invent your own word lists; there are an infinite number of dragons lying in wait for idioms, homophones An unfortunate example. That's TIED HITCH SLOE REIGN RULE MOW? With only two parity bits, you can't even be sure this decoding is invalid. RFC 1751 [0], from which this example comes, doesn't envisage the encoding being used in oral communication. Instead, it makes codes easier the user to "read, remember, and type in". For oral transmission among professionals, sticking to the 26 upper case letters and relying on the NATO alphabet for encoding is a reasonable choice. Getting codes from untrained users in a lossy oral environment is still an unsolved problem. [0] https://datatracker.ietf.org/doc/html/rfc1751 https://datatracker.ietf.org/doc/html/rfc1751
- hinkley 2y agoNot to mention accents. Some people are going to sound like they’re saying Todd for Tide, and have you heard how Baltimore pronounces Iron?
- Muromec 2y agoIt would help if nato alphabet was universally known thing. Typing something letter by letter in Latin when neither party is a native speaker of English is very much painful almost half the time it happens
- yencabulator 2y ago
- albert_e 2y agoI love conversations like this. These are arguably not the most cutting edge or exciting topics but hold a lot of significance and power to make life easier for humans (and machines too). Some of these are areas of best practices that, when done really well -- people may not even notice it. That's an unfortunate fact of life that comes up often -- where the attention to detail and sincerity that people bring to the table often gets lumped under "obviously it should be that way, nothing special to see or applaud here".
- jadengeller 2y ago> When it matters? This applies to usernames too! It's easy to phish if platforms render capital I and lowercase l the same
- eviks 2y ago> not only to avoid visually ambiguous characters, but also to avoid spelling words in common languages. Or you should do the opposite - use real dates/words in ID and your visual confusion almost disappears (though there is a bunch of ambiguity here as well in similar pronunciation, so also not perfect). Humans aren't robots, so shouldn't be forced to read meaningless list of random letters (example of geospatial system of coordinates based on that is what3words)
- raspyberr 2y agoImagine having a coordinate system be owned by a private company.
- bingbingbing777 2y agoYou're free to create your own, or not use theirs.
- Akronymus 2y agoDont they have a patent on it?
- sakjur 2y agoYeah. https://patents.google.com/patent/US9883333B2/en https://patents.google.com/patent/US9883333B2/en I think it’s also a good example of increasing computer dependency by ‘human centric’ design: I can quickly and manually sort through a bunch of packages with coordinates or pluscodes written on them with some sense of locality. What3Words is designed to give a sense of familiarity but require an API lookup for every single address. Letters and numbers also translate directly in most languages, words don’t (take bow as an example. Is it when someone leans over, an archer’s weapon of choice, or a cutesy headpiece?), so the familiarity aspect is limited to people with a good grasp of English. Its main feature is that it can be commercialized, unlike regular coordinate systems.
- 2y ago
- vesinisa 2y agoThe author makes a point of avoiding letters that are hard to distinguish even when spelled out in handwriting, but the example table includes the number 7. I can not count the number of times I have found it hard to distinguish between someone's 7 and 1. It helps if you draw a horizontal bar on the 7 but many don't, so you can never really be sure if a 7 is in fact a 1 with the serif or vice versa.
- gajus 2y agoI never ran into into this situation, but I plan to update the article based on aggregated feedback. A few good suggestions have been made.
- deanishe 2y agoWhen it comes to handwritten numbers, Brits frequently mistake German ones for sevens, and Germans British sevens for ones.
- bithaze 2y agoA small typo I noticed - "Case-sensitive: 53^5 = 62,259,690,411,360" should be to the eighth power, not the fifth.
- gajus 2y agoThanks. Fixed
- silvestrov 2y agoSuggestion: after "a longer ID with a lower chance of visual ambiguity" show how many characters that will be needed to have the same number of IDs as 53^8 using the 22 encoding. I.e. for a given number of IDs, how many characters are needed in the 53 versus 22 encoding (people who are not good at math might assume it is more than twice as many).
- jonp 2y ago
- jzwinck 2y agoIf you use both upper and lower case, you are likely to eventually be surprised by some third party system or protocol that is case insensitive. I even found a commercial system which allowed users to choose IDs with case sensitivity (iD and id being distinct) but if you query it for one which does not exist they do case insensitive matching and return the wrong data. When I reported this bug they said it was for convenience!
- gajus 2y agoThanks for the anecdote. I've included it in the article.
- NetOpWibby 2y agoWhat a nuts system
- atoav 2y agoKeepassXC (open source password manager application) uses colour to make passwords more readable. They use one color for each "class" of character: uppercase, lowercase, numbers, symbols, ... This is a extremely simple idea, but especially with random passwords this helps a lot even if the font is already hyperlegible.
- AndyMcConachie 2y agoAs a colorblind person I hate this idea.
- emsixteen 2y agoI advocate for accessibility and inclusivity constantly, but not implementing additional measures which are helpful to most due to some not being able to make use of that one aid is not the way to go. Direct your hate elsewhere.
- atoav 2y agoYeah, why? Because the additional information layer benefits some people? Depending one your type of color-blindness and the choice of colors this might even be an improvement that works for color-blind people. We are not talking about encoding information only in color (= bad idea), we are talking about encoding information that is already present additionally in the color. And if your app has accessability settings (it should) this would be a thing that you could switch on and off.
- Cthulhu_ 2y agoIt's an additional layer on top of other ones like using a non-ambiguous font, large size display, alternating background shades, character index numbers under each character, etc.
- pbhjpbhj 2y agoYou can also add a list of exclusions easily in the KeepassXC password generator. I do because when you type in a long password on a TV remote, or similar interface, and then realise the l1|I were confused it's soooo0 infuriating.
- kuon 2y agoI came up with base24[1] for this. There are some letter that can be ambiguous but I kept them to make it case insensitive. [1]: https://www.kuon.ch/post/2020-02-27-base24/ https://www.kuon.ch/post/2020-02-27-base24/
- ThePallas 2y agoA few years ago, I created a system that generates a serial number from a prefix and a 32-bit unsigned integer and fixes up this kind of input error when passing the serial. https://github.com/pallas/gubbins https://github.com/pallas/gubbins
- deleted 2y ago[deleted]
- clan 2y agoYears ago I worked support at an ISP who had usernames which was a 12 digit number. Most regular users and 1st level support do no know the NATO phonetic alphabet. An easy trick is trick is then to read back the number for confirmation but use another grouping of digits. Most users read 1 digit at a time so I would read back 2. One-Two becomes twelve. If they used 2 digits I would for ease use 3 rather then 1. This is a very easy way to do a fake "checksumming" regular people. Tangent: All number started with 12 which in effect made them 10 digits. They worked together with a banking system and the bank folks thought 10 digits was not secure enough so they complied and added 12 in front of everything.
- deleted 2y ago[deleted]
- account42 2y ago> Tangent: All number started with 12 which in effect made them 10 digits. They worked together with a banking system and the bank folks thought 10 digits was not secure enough so they complied and added 12 in front of everything. Delicious malicious compliance - I like it.
- NickHoff 2y agoI'm an American living in Germany. When I first arrived, the way Germans write the digit 1 surprised me. They write it with the upper hook thing very long, almost like a capital lambda (Λ), which sometimes makes 1 and A visually ambiguous. This isn't really a problem, just something funny about moving to a new country.
- froh 2y agomy us colleagues regularly mistook the ones for sevens. that's btw why we cross the sevens, like tees and effs
- lynguist 2y agoI use 1 with a long hook except when I write binary numbers where I use just a | for 1. I have some other context dependent characters/letters. I write small z like that in normal writing, but as a mathematical variable I write it as ƶ. (To disambiguate from 2.) I write small t like † in normal writing, but as a mathematical variable I write it as t. (To disambiguate from + (plus).) I write q like that in normal writing, but as a mathematical variable I write it with a stroke, which does not display on the iPhone, a ꝗ, a bit similar to a ɋ. (To disambiguate from a (ɑ).) It’s all about disambiguation, and sometimes having different letter shapes for isolated characters.
- matthewtse 2y agoSo cool to read an article discussing a problem I run into on a regular basis. Whenever I'm creating a 2FA backup on a piece of paper, anxiety hits me every time I cross over certain characters, o/0, v/u, 5/S, etc. I've come to add some fanciness to how I write these characters for this exact reason. On "Phonetic similarity", reminds me of how I chose my wifi password. I wanted a common word with multiple consonants that a 3rd grader could spell, so I could share the password with a single phrase and have it be unambiguous. Ended up choosing "vacation".
- 2024throwaway 2y agoI can’t believe people out there write these things down by hand on paper. It’s mind bottling.
- matthewtse 2y agoI do that out of paranoia/mistrust for my wifi network, printer, printer software, etc. It's probably fine to just print it out, but for more sensitive items I definitely write it down by hand.
- Piskvorrr 2y agoIt's not as if the printer keeps a hidden cache of printed pages. Except maybe it does...even if the feature was created for entirely benign reasons.
- ant6n 2y agoIt’s not as if photocopiers could randomly replace letters or numbers, right? …right? Or perhaps they could: https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...
- Piskvorrr 2y ago
- maggit 2y agoI have realized that there is a big design space here, as I recently did a write-up of my take, Id30. 30 bits of information encoded base 32 into six chars, eg bpv3uq, zvaec2 or rfmbyz, with some handling of ambiguous chars on decoding. https://magnushoff.com/blog/id30/ https://magnushoff.com/blog/id30/
- nmstoker 2y agoA friend told me about how his work had some senior IT mgrs, who'd clearly been playing with their iPhones too long, decide that the firm shouldn't use Ids at all any more, and started pushing this without consulting the business, even though it was totally inappropriate given how widely they were needed... Caused mayhem and needles arguments!
- EnigmaFlare 2y agoDoesn't help when you have to match the person's name and they have these characters in them. My name contains the letter "o" and I once had a lot of trouble getting something done at the bank. Multiple staff had to crowd around the computer to figure it out. Eventually somebody discovered that when I had opened my account, that o had been entered as a 0 for some reason and the font they were using, also for printing, showed them looking almost identical.
- 8organicbits 2y agoSimilar story here, but a "Q" instead of an "O". The tail of the Q looked like dust on the screen. Somehow I haven't run into issues...
- junga 2y agoAlso do not use the same character repeated in a "long" sequence. I hate this with IBANs. Too often there's something like '000000' right in the middle of an IBAN and in case copy and paste is not possible I end up counting the number of zeroes at least thrice. Groups of four characters separated by spaces would help in this case but that's another topic.
- thih9 2y agoWe could always use 1s and 0s, maybe group them in eights. Tongue in cheek, but I guess that would be a valid (even if extreme) solution.
- dusted 2y agoThis is why I only ever use xterm with the default bitmap font, it's literally the only one where I'm absolutely sure which character is which.
- jeroen 2y agoAn alternative would be to print IDs using https://en.wikipedia.org/wiki/FE-Schrift https://en.wikipedia.org/wiki/FE-Schrift, which was specifically designed to make normally similar characters to look different.
- serial_dev 2y agoGood luck distinguishing 0 and O with that font in a random sequence of characters. The type face you linked is not optimized for humans. > Its monospaced letters and numbers are slightly disproportionate to prevent easy modification and to improve machine readability. It's a slightly different issue than what was described in the article (e.g it can't address the cases where IDs are written down).
- namibj 2y agoAFAIK it's designed for non-automatic number plate reading in Germany.
- bdjsiqoocwk 2y ago"visually unambiguous dictionary" to the author. It's well known that some people have a hard time distinguishing p/b/d/q.
- chiefalchemist 2y agoFour quick thoughts: - We haven't solved this already? Who hasn't tried to read some code and couldn't tell O from 0 or l from 1, etc.? - Aside from ambiguous characters you have to be aware of spelling and leet spelling. e.g., 53X, S3X, 5EX, etc. - FFS stop with the 10+ character strings without spaces or hyphens. There's no reason for that. - Not everyone has perfect vision. Ambiguous characters *and* less than perfect vision (often with not spaces / hyphens) is a mortal UX sin. We've all been on the wrong end of these, and yet they are common enough - in 2024??!!? - that they need to be mentioned here.
- nemoniac 2y agoOther prior art is the use of a modified base 58 encoding in Bitcoin addresses. https://en.bitcoin.it/wiki/Base58Check_encoding https://en.bitcoin.it/wiki/Base58Check_encoding
- loloquwowndueo 2y agoAs long as we are pointing out mistakes in the article: 9qg6G8B2Z5SIl170O (ariel) The name of the font is Arial, not Ariel. (No mermaids here, move along)
- rob74 2y agoYup... also, a screenshot (or using webfonts) would have probably worked better there. On Linux, most of the lines look the same...
- gajus 2y agoHeads up that the article is open source in case you wanted to contribute an edit. https://github.com/gajus/gajus-com/blob/main/src/blogPosts/2024-04-22-avoiding-visually-ambiguous-characters-in-ids/blogPost.mdx https://github.com/gajus/gajus-com/blob/main/src/blogPosts/2... I fixed the typo though. Thanks!
- croes 2y agoIf we include handwriting then lowercase n and u get be hard to distinguish if written in cursive
- croes 2y ago>Avoiding Confusion With Alphanumeric Characters https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3541865/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3541865/
- branon 2y agocl looks like d in some fonts or with bad kerning
- 8organicbits 2y agoThe Latin/English alphabet is common but not universal. I believe this challenge is why TOTP codes use Arabic numerals. The user's keyboard can type these reasonably. Spoken is always a challenge. Even an English speaking audience will pronounce "0" as zero, oh, or zed.
- bloak 2y ago0123456789 are best called "European", I think, as Arabic numerals would be: ٠١٢٣٤٥٦٧٨٩
- 8organicbits 2y agohttps://en.m.wikipedia.org/wiki/Arabic_numerals https://en.m.wikipedia.org/wiki/Arabic_numerals
- arp242 2y ago"They are also called [..] European digits"
- shagie 2y agoThe reference chases through to https://www.unicode.org/terminology/digits.html https://www.unicode.org/terminology/digits.html Term: ASCII digits Example: 0123456789 U+0030..U+0039 Explanation/Description: Commonly used with Latin, Greek, Cyrillic and many other scripts, including some non-European scripts. Used in alternation with native digits in scripts that have them. (Some scripts with native digits make only limited use of ASCII digits.) Infrequently used in many of the remaining scripts. Synonyms: Western digits, Latin digits, European digits Which then links on to: https://www.unicode.org/glossary/#european_digits https://www.unicode.org/glossary/#european_digits > European Digits. Forms of decimal digits first used in Europe and now used worldwide. Historically, these digits were derived from the Arabic digits; they are sometimes called “Arabic numerals,” but this nomenclature leads to confusion with the real Arabic-Indic digits. Also called "Western digits" and "Latin digits." See Terminology for Digits for additional information on terminology related to digits.
- criddell 2y ago> I would be wary of excluding characters just because they look like other characters when combined I wish the author would have said more about this. Why be wary?
- digging 2y agoThe implied reason is that it shortens the list of available IDs substantially.
- criddell 2y agoThat was my first thought, but the section on case sensitivity already discussed the impact of a reduced alphabet and pointed out adding more characters takes care of that quickly. So I assume the reason is something else.
- p0w3n3d 2y agoLetters l and I are visually indistinguishable when written in Arial.
- geoffreysimpson 2y agoI did my PhD on (malicious) visual impersonation of domain names using many of the techniques described here. There are many references to other visual doppelganger techniques included in my paper here: https://par.nsf.gov/servlets/purl/10256904 https://par.nsf.gov/servlets/purl/10256904 My research focused solely on the .com domain name space, so our character set was limited.
- mistrial9 2y agothat research paper only considers ascii characters in domain names?
- geoffreysimpson 2y agoThe paper only considers .com domain names, which have a limited character set support, discussed in RFC 1034 https://www.ietf.org/rfc/rfc1034.txt https://www.ietf.org/rfc/rfc1034.txt Essentially A-Z, 0-9, and the - character, and domain names can not start with the dash character.
- kuboble 2y agoIn handwriting there is a difference between European and American. In Europe we don't really have problem with 1 vs 7 or g vs 9. But our nines and ones do look like gs and sevens to Americans. I heard an American making a joke that "I have gg problems but European handwriting ain't 7 of them."
- rsync 2y ago“Oh By”[1], The universal shortener, has had protections for this built in from the very beginning. Since the whole point is the ability to convey a message in the physical world end with chalk or pencil or whatever – we needed to make sure that characters were unambiguous. So there are no zeros or ‘o’ characters or ones or ‘l’ characters… I think there were one or two other rules that govern this but I can’t think of them right now… [1] https://0x.co https://0x.co
- yencabulator 2y agoI'm a fan of z-base-32 for this. https://philzimmermann.com/docs/human-oriented-base-32-encoding.txt https://philzimmermann.com/docs/human-oriented-base-32-encod... Command line tool at https://github.com/tv42/zbase32 https://github.com/tv42/zbase32 $ echo hello, world | zbase32-encode pb1sa5dxfoo8q551pt1yw $ entropy 16 | zbase32-encode y64s31aq6cgjoko9fwbuasf4ce
- cryptonector 2y agoAnd TFA doesn't even mention Unicode, scripts, ASCII, Latin, nothing. As you can imagine it all gets much worse with Unicode (though through no fault of the Unicode Consortium). See Unicode TR#39 [0]. [0] https://unicode.org/reports/tr39/
- waltbosz 2y agoI thought this was good neat UX: on the Nintendo Switch I was entering a serial number for some DLC, and the on-screen keyboard had all the ambiguous character keys disabled, which means that the serial numbers are generated without any ambiguous characters. I'm not sure if this UX was built into the OS, or just part of the game I was playing (Mario + Rabbids Sparks of Hope).
- benaubin 2y ago> However, as the number of members in the set increases, the number of possible IDs increases exponentially. Case-sensitive: 53^8 = 62,259,690,411,361 Case-insensitive: 22^8 = 54,875,873,536 Nitpick, but isn't this polynomial to the members of the set?
- pxx 2y agoaⁿ grows (a/b)ⁿ as quickly as bⁿ. The multiplicative difference still grows exponentially in n.
- afiori 2y agoa^n is polynomial in a and exponential in n. This is why longer password are more efficient than complex passwords: to gain the same security effect of doubling the password length you would need to square the alphabet
- pxx 2y agoYou've the proper definitions but are missing the context. An exponential with larger base still has an exponential multiplicative difference compared to an exponential with a smaller base. We're comparing the growth rate of of two exponentials representing variable-length identifiers. We're not looking at a constant-length identifier (which is what you're doing with only looking at a^n). Notice the context of where exponential is used in the article: we are changing n from 5 to 8.
- denimnerd42 2y agomy work id has a 0 and a O in it and it drives me crazy. i only remember it due to muscle memory on the keyboard
- nullc 2y agoModern bitcoin addresses use a base-32 character set that leaves out some of the most ambiguous pairs and also permutes the address ordering so that the most visually similar remaining characters produce single bit errors which are better handled by the addresses error detecting (and potentially correcting) code. https://github.com/bitcoin/bips/blob/master/bip-0173.mediawiki#user-content-Specification https://github.com/bitcoin/bips/blob/master/bip-0173.mediawi...
- svat 2y agoRelated reading, from the font designer's side: “Oh, oh, zero!” by Charles Bigelow (of Bigelow and Holmes, makers of typefaces like Lucida and Wingdings), published in TUGboat the journal of the TeX users group: https://tug.org/TUGboat/tb34-2/tb107bigelow-zero.pdf https://tug.org/TUGboat/tb34-2/tb107bigelow-zero.pdf (There's also a “footnote” by Donald Knuth: https://www.tug.org/TUGboat/tb35-3/tb111knut-zero.pdf https://www.tug.org/TUGboat/tb35-3/tb111knut-zero.pdf, and follow-up by Bigelow: https://tug.org/TUGboat/tb36-3/tb114bigelow.pdf https://tug.org/TUGboat/tb36-3/tb114bigelow.pdf)
- TacticalCoder 2y ago> Related reading, from the font designer's side: “Oh, oh, zero!” by Charles Bigelow I don't know. People tend to use the letter 'O' a lot. And people tend to use zero '0' a lot too. Who gives a fuck about "Oh"? I mean, seriously, which percentage of articles, blog, PDFs, webpages, products etc. throughout the world have have 'O' and '0' that can be mistaken one for another? And which percentage have "Oh"? When was the last time a user had to read a product ID over the phone and did misread big O / "Oh" for 0? I don't even think there was a last time, because nobody is using "Oh" in identifiers. While, on the other hand, it's perfectly fine to use a slashed-zero for zero, to be sure nobody mistakes it for the letter 'O'. So basically: your link and TFA aren't that related.
- svat 2y agoI'm not sure I understand your comment, because at first glance it seems to be making a distinction between "Oh" and "O", when Bigelow's article is using "Oh" as the name/vocalization of the letter 'O' (as should be clear from the very first sentence, even if not the title). So, assuming (still not clear from your comment) that you do understand "oh" to mean the letter 'O', as intended, still your comment is surprising, because some of your own other comments talk about O/0, and the submitted post here too starts with that very example: > What are visually ambiguous characters? > O / 0 - The letter O and the number 0 can look very similar So surely the article is relevant to (at least the first example of) the post? I admit it goes much deeper into just this one example, and only a bit into other examples like 1/l/I and 2/Z or 5/S, but still it's relevant and of value as a representative example I think.
- TacticalCoder 2y agoAnother confusing thing is doing this: xxxxx-xxxxx-xxxxx-xxxxx Instead of something like this: xxxxx-xx-xxxxx-xxx-xxxxx Something could also be said about such scheme lacking the embedding of a checksum. Here's an IBAN (bank account number) in the EU (which thankfully are using a checksum as part of the account number): LU29 0022 1712 5582 7000 ^^ || two checkdigits Also some companies think they're "smart" because they pick numbers like this: LU29 002 0000 0001 8000 Repeating the same digit, usually a zero, a shitload of time ain't smart. It's fucking dumb.