3 ms·
We have not yet developed writing systems for all of the world's languages. So there is no practical bound you can put, because it is possible that in a future
by lambda 11y ago
We have not yet developed writing systems for all of the world's languages. So there is no practical bound you can put, because it is possible that in a future development of a writing system, there will either be entirely new characters invented, or new combinations of existing characters and diacritics.
In addition, since the existing Unicode standard does not impose a bound, and implementations support stacking of large numbers of diacritics, people have used combinations of extremely large numbers of diacritics as a stylistic device, for example: http://stackoverflow.com/a/1732454/69755 http://stackoverflow.com/a/1732454/69755. Since examples like this exist in the wild, they need to be supported properly, and so the fundamental string types and operations in string handling libraries should be able to handle it.
- protomyth 11y agoYes, we have to live with what was created, but given 128-bits is boil the ocean territory and someone has to type this stuff, I don't see a compelling argument for unbounded.
- lambda 11y agoBut why bother with a fixed width 128 bit encoding? We have perfectly fine variable-width encodings, that take up 8 bits per character for the most common set of characters encoded, the ASCII range (even for CJK text, the fact that most markup languages like HTML and XML use ASCII means that actually, the bulk of documents fall within the ASCII range), and can handle arbitrarily complex stacks of diacritics. How would it be helpful to waste a huge amount of space for the vast majority of text, impose additional limitations on the number of combining characters, just for the benefit of fixed-width encoding of graphemes, which even in itself is not all that helpful since for most applications, you are either going to need to scan through the text linearly, or you could use byte offsets in the variable width encoding just as well as you could use grapheme offsets in a fixed width encoding?
- protomyth 11y agoI used 128-bit as a max. I just don't like some unbounded thing when we know that there is a practical max. I wouldn't call something perfectly fine when getting the number of characters actually displayed on the screen is a chore.
- lambda 11y agoWhat's wrong with unbounded? The amount of text you're rendering is unbounded (except by resource constraints). Combining diacritics are just another kind of text, that happen to lay out vertically rather than horizontally even in horizontal layout. What do you mean by "getting the number of characters actually displayed on the screen"? Do you mean number of glyphs? Number of grapheme clusters? The number of glyphs is dependent on your text rendering system and fonts; even with plain ASCII, you may have multiple codepoints rendered as a single glyph, such as in ligatures like "ffi" in certain fonts. The number of grapheme clusters can be computed by the given algorithm along with some tables of character properties; it's not particularly simple, but trying to represent languages that were originally hand-written on a computer is not simple: http://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries http://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries
- protomyth 11y agoThe number of times I have to hit the backspace to delete a line of text.
- nitrogen 11y agoInstead of backspace, shift+home, delete is a fast way to delete lines of any length.
- TazeTSchnitzel 11y agoThat's application-dependant.