5 ms·
Further reading: * https://hoytech.github.io/truncate-presentation/ https://hoytech.github.io/truncate-presentation/ / https://metacpan.org/pod/Unicode::Trunca
by re 2y ago
Further reading:
* https://hoytech.github.io/truncate-presentation/ https://hoytech.github.io/truncate-presentation/ / https://metacpan.org/pod/Unicode::Truncate https://metacpan.org/pod/Unicode::Truncate
* https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundarie...
Truncating at codepoint boundaries at least avoids generating invalid (non-UTF-8) strings, but can still result in confusing or incorrect displays for human readers, so for best results the truncation algorithm should take extended grapheme clusters into account, which are probably the closest thing that Unicode has to what most people think of as "characters".
- iforgotpassword 2y agoTo avoid this and a bunch of other confusion, when accepting user input I recommend normalizing it to the composed form before writing to a DB or file. While unicode-aware tools and software should handle either form just fine, you probably want to avoid that there's something in the pipeline somewhere that treats the decomposed and composed form of the same string as different.
- arp242 2y agoIt's good advice to normalise to pre-composed form, but that doesn't solve the problem the previous poster mentioned as not everything exists as a composed form. That said: most things do have a composed form, so you can probably get away with it – right up to when you can't.
- quink 2y agoYeah, working an a library system our path was to compose everything (taking into account of course that the octet sizes specified in the directory may or may not actually be accurate depending on whatever system produced the record) and around the same time deprecate any pretense we had of supporting MARC-8.
- magicalhippo 2y agoHow does that affect filenames? IIRC the lower levels of Windows will happily work with filenames that are not valid Unicode strings, for example if you use the kernel API rather than Win32. But what about Win32? If you create a file before normalization and then open it using the normalized form, will it open the same file or return file not found? What about other systems? For example AWS' S3 allows UTF-8 keys, with no mention of normalization[1]. On the phone so can't try myself right now. Anyway for general text I agree, but for identifiers, filenames and such I prefer to treat them as opaquely as possible. [1]: https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-keys.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/object...
- heinrich5991 2y agoNo need to go to the kernel API to create filenames that are invalid UTF-16. The Win32 API will happily let you do it.
- magicalhippo 2y agoI was AFK so couldn't check, but yeah you're right. Just made two files named ä.txt which happily sat next to each other, one being NFC and other NFD. So yeah, don't mess with the normalization of filenames.
- panzi 2y agoThey're both valid UTF-16, though. Can you create a filename with only half of a surrogate pair in it? I don't use Windows, so I can't check. Linux literally allows any arbitrary byte except for 0x00 and 0x2F ('/' in ASCII/UTF-8). It's a problem for programming languages that want to only use valid Unicode strings, like Python. Rust has a separate type "OsString" to handle that, with either lossy conversion to "String" or a conversion method that can fail. Python uses the custom use Unicode range to represent invalid byte sequences in filenames. It's all a mess. JavaScript doesn't give a damn about the validity of their UTF-16 strings. (Note that Rust's OsString is different from it's CString type. Well, I guess under Unix they're the same, but under Windows OsString is UTF-16 (or "WTF-16", because it isn't actually valid UTF-16 in all cases).)
- account42 2y agoNot all grapheme clusters have composed forms so normalization doesn't actually gain you anything here.
- nordsieck 2y ago> Not all grapheme clusters have composed forms so normalization doesn't actually gain you anything here. Just because the worst case can't improve doesn't mean that making the average case better is worthless.
- panzi 2y agoI saw something about Arabic text, where that naive truncation at codepoint boundaries turns one word into a different word! Like the sequence of codepoints generate something that is represented as a single glyph in fonts, but truncated its totally different glyphs. I don't remember more details, I don't know any Arabic, but grapheme clusters aren't just about adding diacritics to latin characters. In other languages it all might work quite differently. So truncating at word boundaries (at breakable white-space or punctuation) is probably best. Though of course that way you might truncate the string by a lot. shrug-emoji (I don't think the talk where the stuff about Arabic was mentioned was Plain Text by Dylan Beattie, but I haven't re-watched it to confirm. So maybe it is. Can't remember the name of any other talk about the subject right now.)
- gmueckl 2y agoRandomly truncating words can have the same effect in any language. It's outright trivial to find examples in English or German. I don't understand why one has to invoke Arab script for a good example.
- asabil 2y agoYes, but you don’t end up with different glyphs. Arabic script has letter shaping, that means a letter can have up to 4 shapes based on its position within the word. If you chop off the last letter, the previous one which used to have a “middle” position shape suddenly changes into “terminal” position shape.
- mort96 2y agoThe emoji "" can't be normalized further -- it's a "" followed by a "". If you just split on code points rather than grapheme clusters, even after normalizing, your naïve truncation algorithm will have accidentally changed the skin colors of emoji. Or turned the flag of Norway into an "". Or turned the rainbow flag into a white flag . EDIT: oh lord Hacker News strips emoji. You get the idea even though HN ruined the illustrations. Not my fault HN is broken.
- recursive 2y agoPresumably referring to Fitzpatrick modifiers.
- tingletech 2y agoHN is not broken, it's working as designed.
- postmodest 2y ago[poop]
- mort96 2y agoIt makes technical conversations about Unicode ridiculously annoying. It's working as designed and the design is broken.
- lisper 2y agoWhether the absence of emojis on HN is a feature or a bug is arguable. But if you can't figure out a way to work around this constraint (e.g. put your example literally anywhere else on the web and post a link here) HN is probably not a good fit for you.
- mort96 2y agoReading a discussion thread where each message is just a link to some pastebin with the actual message isn't very nice. Besides, I wasn't going to write the message again after HN removed arbitrary parts of it, hence the edit; I think people got the gist. You may feel that discussion about Unicode doesn't belong on HN but I feel otherwise.
- Sharlin 2y agoEven if you only support scripts for which Unicode has composed codepoints, these days you likely can’t get away without properly handling emoji, and there are no precomposed versions of all the numerous emojis that are made of multiple code points (eg. skin color and gender variants as well as flags).
- masklinn 2y agoFor actual best results you’d probably want to truncate at the word or syllable boundary, and it should likely be language specific.