3 ms·
5th bullet of "Who is this library for?" at the top of the documentation [1] had me excited: > you want to work with higher-level primitives (code points, grap
by boxfire 5y ago
5th bullet of "Who is this library for?" at the top of the documentation [1] had me excited:
> you want to work with higher-level primitives (code points, graphames) when iterating text that do not break your text apart;
In reality grapheme clusters are implemented nowhere and are only in the glossary. I wanted to see how this handled the sprawling complexity that can beget.
Someone call me when a library can iterate in that form, which isn't written by committee.
[1]: https://ztdtext.readthedocs.io/en/latest/ https://ztdtext.readthedocs.io/en/latest/
- remexre 5y agoIf you're fine with an extra copy, I think the UTF8PROC_CHARBOUND flag does this in https://github.com/JuliaStrings/utf8proc https://github.com/JuliaStrings/utf8proc ?
- boxfire 5y agoNice this does the job, probably could be done more efficiently, but I like that its not a kitchen-sink library. Unfortunate that the data generator is certainly not "clean C"... perl, ruby, and julia, all to generate character width table? Ugh. Otherwise this lib does a pretty good job. Thanks for the reference! edit* Yikes https://github.com/JuliaStrings/utf8proc/blob/master/utf8proc.h#L554 https://github.com/JuliaStrings/utf8proc/blob/master/utf8pro... Still this is a good template to maybe cut a (safter) state-based iterator out of their logic...