4 ms·
I would argue that even if you decide that you are using some other language and not English, there is only a well-defined subset of Unicode characters that sho
by amenod 5y ago
I would argue that even if you decide that you are using some other language and not English, there is only a well-defined subset of Unicode characters that should ever be allowed in the codebase. Bidi override control characters are clearly not among them, whichever language you choose.
- rbanffy 5y ago> Bidi override control characters are clearly not among them, whichever language you choose. Not sure how would you write a comment in an RTL human language in the middle of LTR code without it. Lots of people write learn RTL languages well before writing any code. What compilers can do is to process those characters and assign them semantic value that makes the code equivalent to what is expected to be rendered. Now, bidi overrides in identifier names is a nightmare I’d prefer to avoid.
- amenod 5y agoThe same way as you write a comment in a LTR human language in the middle of RTL code - you don't. You stick to either LTR or RTL. This is code, not prose.
- rbanffy 5y agoCode is meant to be read and, occasionally, executed. Comments are usually ignored by compilers and are targeted towards humans.
- jrochkind1 5y agoYou do not actually need the bidi override control character to put a comment in an RTL language in the middle of LTR code. You only need it if you are doing this, and the default Unicode algorithm for guessing LTR/RTL boundaries gets it wrong, so you need to override with an explicit bidi override control. I'm not even sure how feasible that is to do in current editor/IDE environments developers who have this use case might use. I am genuinely curious how often these sorts of situations come up in actual development. > What compilers can do is to process those characters and assign them semantic value that makes the code equivalent to what is expected to be rendered. I don't understand what you mean or how that's even possible, for the kinds of attacks discussed in OP.
- jrochkind1 5y agoBtw here's proof. Here is ltr text and rtl עִברִית text عربي interspersed with no bidi override control characters to be found. Unicode can handle this, it has a heuristic algorithm for it. Note how if you try to select the text character-by-character, your selection does funny things at the rtl to ltr boundaries, because the byte order doesn't match the order on the screen. It really is handling the directionality changes, with the letters entered in "order" across changes, there is no funny entry or ordering going on, this is plain old normal unicode handling interspersed directionality changes just fine, with no bidi overrides. It just sometimes gets it wrong for the intent of the author. Especially when there are characters at the boundaries that are themselves not strongly associated as rtl or ltr (like ordinary "western arabic numerals" or punctuation). That's what the bidi override control char is for.
- WalterBright 5y ago> Not sure how would you write a comment in an RTL human language Siht ekil.
- chmod775 5y ago> there is only a well-defined subset of Unicode characters that should ever be allowed in the codebase It's not even remotely well-defined, and probably never will be. Also, as long as we keep adding to unicode, you will need to keep your whitelist of code points updated. You can however find a well-defined subset of characters that can be allowed. In either case you'd be essentially excluding entire languages.
- amenod 5y agoYou misunderstood my point: >> There is only ... that should ever be allowed... What I am saying is someone decides to code in a non-english language (which is completely reasonable) they should define a subset of unicode characters that is acceptable. Additionally, the allowed characters should not permit tricks like these. As for excluding entire languages... well, yes. This is already the case today. But OTOH it's not like understanding what "if" means gives you any special advantage in programming.