4 ms·
As an alternative to ICU, there is suckless's libgrapheme (https://libs.suckless.org/libgrapheme/ https://libs.suckless.org/libgrapheme/) which is more than a 1
by simjnd 2y ago
As an alternative to ICU, there is suckless's libgrapheme (https://libs.suckless.org/libgrapheme/ https://libs.suckless.org/libgrapheme/) which is more than a 100x smaller and provides full Unicode compatibility.
- wwalexander 2y agoHeh, I maintain the MacPorts port for this lib! Like all suckless projects, it’s written in very simple and portable C which makes it a breeze to package. https://ports.macports.org/port/libgrapheme/ https://ports.macports.org/port/libgrapheme/
- frign 2y agoThanks for your work on packaging my library. Please let me know if I can make the process simpler for you; I take great care to make packaging as simple as possible for the packagers. Likewise I have no sympathy for those writing software that is almost deliberately hard to package.
- wwalexander 2y agoThank you for your work on the library! It really couldn’t be easier to package but thank you <3
- HexDecOctBin 2y agoAre there any similar options that handle collation and normalisation? No small Unicode library seems to implement them, unfortunately.
- frign 2y agoAuthor of libgrapheme here: Both collation and normalisation are non-trivial and have many gotchas thanks to the way the Unicode consortium likes to write their specifications. I sometimes get the feeling that they don't even care about implementers and just document what is done in the reference implementation ICU. The only sensible normalisation one can implement is the full decomposition (NFD), and maybe the full composition (NFC). You rely on the full decomposition if you want to collate correctly, which is a problem because the amount of memory needed to store the decomposition is unbounded in general. I don't want to make the libgrapheme users jump through hoops, and I also don't want to do any memory allocations in libgrapheme either. There is an idea floating in my head on how to solve this, but I'm currently busy finalising Unicode 15.1 support (Unicode 16.0, released on the 10th, will be trivial to upgrade to) and releasing my already fully-compliant implementation of the Unicode bidirectional algorithm.
- HexDecOctBin 2y agoI see, thanks for replying. I agree, Unicode specs are hell to work with (I tried doing a auto-codegen thing based on them and just gave up due to the size of tables generated and the seemingly-arbitrary edge cases). libgrapheme looks pretty good otherwise, I'll keep an eye on it for whenever I have to wrangle with Unicode on a low level again (hopefully not for a long time).
- velorek 2y agoInteresting. Cool library, I'll check it out. :) Thanks for sharing
- frign 2y agoThanks for recommending libgrapheme. I am honoured, being the author of this library.