4 ms·
Well, teddyh might have a point here, nonetheless: By now, I understand that ftfy is about fixing mixed up encodings between UTF8, latin-1, CP437, CP125[12] and
by fnl 12y ago
Well, teddyh might have a point here, nonetheless: By now, I understand that ftfy is about fixing mixed up encodings between UTF8, latin-1, CP437, CP125[12] and MacRoman (only). But by claiming you are fixing "Unicode" in general as the first thing on the GitHub page, you might be misleading first-time visitors. Maybe you should try to place the "warning" about the encodings your library does handle right at the start somewhere? And make it clear that "moji-un-baking" is the library's central and main use-case, not just an "interesting thing" it can do. Despite being quite aware of Unicode and string encoding, I had exactly the same thoughts as teddyh as I read the first few paragraphs ("Oh, now we will see those encoding illiterates converting all those beautiful bytes in some highly informative character encoding to all-too-boring-ASCII.")
Which leads me to my other concern: Why do you use NFKC compatibility as the default normalization? Given you are a text mining company, you of all guys should know you loose valuable information - particularly about numbers, super- and subscript characters - with this normalization strategy. Doing NFKC on stuff like all kinds of articles, books, patents, etc. would lead to potentially disastrous results (e.g., NFKC "decomposes" the string 'O\u2082\u00B9' to 'O21' instead of 'O_2^1' - "oxygen, reference 1"). In general, I think NFC is what Python and many other libraries do, while I believe NFKC should only be used when you know what you are doing (and why you need it). Maybe it is useful for some strange, geeky tweets, but I would argue that its the corner case, not the default.
- rspeer 12y agoI wonder if I could change the default to NFC in the next version without breaking people's expectations. It is a safer default. When it comes to text analytics, the underlying tagger and stuff won't know what O21 is any more than it knows what O_2^1 is anyway. And NFKC is useful for mixed Latin and Japanese text, which I wouldn't entirely dismiss as strange and geeky. But it's true that the default could be more conservative.