3 ms·
Fun fact: when treated with unicode Normalization Form Canonical Decomposition, 8 out of 9 polish letters (ż,ó,ć,ę,ś,ą,ź,ń) break down into base letter + combin
by notathrowaway51 3mo ago
Fun fact: when treated with unicode Normalization Form Canonical Decomposition, 8 out of 9 polish letters (ż,ó,ć,ę,ś,ą,ź,ń) break down into base letter + combining diacritical mark, but ł stays intact. That means you can't use sqlite's unicode61 remove_diacritics tokenizer to normalize polish text for FTS.
- ks2048 3mo agoWhen a Polish speaker searches for something with “ł”, do they expect to also see “l”?
- kuboble 3mo agoNo. But the other way around sometimes yes.
- dhosek 3mo agoI remember discovering that while writing some code for a job interview. The reason for it is simple, even though in many input systems (like the ABC International I use on my Macs) it’s a two-character sequence to enter ł, there is not actually a combining character for that line through the l. I’m not sure, but I think sqlite’s remove_diacritics works the way that I’ve implemented that functionality in some of my own software: convert to NCD then remove combining characters from the string. I would expect that a few other special cases also behave the same way, such as ħ or ø which also will not decompose.