5 ms·
... and in reverse with this very cool tool (found on HN I think) https://ftfy.vercel.app/?s=tr%C3%83%C2%A8s+int%C3%83%C2%A9ressant https://ftfy.vercel.app/?s=t
by murkle 5y ago
... and in reverse with this very cool tool (found on HN I think) https://ftfy.vercel.app/?s=tr%C3%83%C2%A8s+int%C3%83%C2%A9ressant https://ftfy.vercel.app/?s=tr%C3%83%C2%A8s+int%C3%83%C2%A9re...
- capitainenemo 5y agoWell... you can do it in reverse with iconv too... $ echo très intéressant | iconv -f utf-8 -t iso-8859-1 très intéressant Admittedly no autodetection. Luckily EU mangling is usually just one or two encodings.
- rndgermandude 5y ago> Luckily EU mangling is usually just one or two encodings. Just to list the iso-8859 parts concerning EU member states: - iso-8859-1 (Latin-1, Western European, including German umlauts, French accents, etc) - iso-8859-2 (Latin-2, Central European, including characters to support Polish, Czech, Slovakian, Hungarian and other) - iso-8859-3 (Latin-3, South European, including characters to support Maltese) - iso-8859-4 (Latin-4, North European, including characters to support the Baltic states) - iso-8859-5 (Latin/Cyrillic, including characters to support Bulgarian) - iso-8859-7 (Latin/Greek, including characters to support Greek) - iso-8859-10 (Latin-6, Nordic, refinement of Latin-4, popular in Baltic states) - iso-8859-13 (Latin-7, Baltic Rim, because -10 was not enough) - iso-8859-15 (Latin-9, basically Latin-1 with the €-sign and some commonly used characters missing in Latin-1) - iso-8859-16 (Latin-10, South-Eastern European, "Intended for Albanian, Croatian, Hungarian, Italian, Polish, Romanian and Slovene, but also Finnish, French, German and Irish Gaelic (new orthography)") And they are all still in use. ;) It seems to me -15 is now more popular than -1, probably because it supports the Euro currency sign.
- capitainenemo 5y ago$ for i in {1..16};do echo -n "ISO-8559-$i: ";echo très intéressant | iconv -f utf-8 -t "iso-8859-$i" 2>&1;done | grep -Pv "illegal input|failed" ISO-8559-1: très intéressant ISO-8559-9: très intéressant ;)
- garaetjjte 5y agoOh, but that's not all! :) Microsoft had its own Windows-125x codepages, which were not always compatible with ISO ones.
- josephcsible 5y agoGP did say "usually", and isn't it usually either -1 or -15?
- rndgermandude 5y agoI have also seen a lot of -3, -5, -7. And recently too ;) The others not that much, but not never either. At least I believe I did, because often times you get that stuff without any hint what it is and then you can make a somewhat educated guess only. There is still a lot of software out there that doesn't default to unicode but either some of the iso-8859's or some windows code page. E.g. WordPad in Windows 10 uses the an ANSI code page when saving an .rtf (WordPad default format)[0] or as plain .txt[1]. But fret not if you're on macOS, my TextEdit just saved this when I created a new document and inserted a few ä-s (the TextEdit default is .rtf as well) >Ohne Titel.rtf: Rich Text Format data, version 1, ANSI, code page 1252 At least .rtf has some embedded metadata specifying the code page of the document and more! Including mixing in some unicode[2]. Why do I mention .rtf? Because there is a lot of .rtf out there, being saved and emailed around each day. All bets are off for plain text formats that do not come with such useful metadata (except when it's some valid UTF (with BOM)). There is also a ton of (old) email software/webmailers deployed that do not produce unicode. Some old and/or shitty enough to produce beautiful HTML soup that does not specify any character set at all. And all kinds of tools that produce text or csv and the like in all kinds of funky encodings that aren't anything unicode and quite often are selected based on the system locale or ANSI code page (on Windows). Or just hardcoded. Or yet better: mixing different encodings in the same file. Do you want to read the exif metadata your camera embedded into those jpegs it spit out? The standard says everything is 7-bit ASCII, but of course that won't fly, so camera vendors and image editor vendors started using all kinds of encodings. And then there are vendor-specific additional tags. Windows XP's photo editor e.g. managed to smuggle in a creator tag that embeds the name of the user account which created/edited a file and it's using UCS-2/WTF-16, of course! [0] To be fair, the format was specified long before there was unicode. To be even fairer, it's all 7-bit ASCII with escaped 8-bit whatever-codepage. [1] There is an "Unicode text file" format option, which, you guessed it, saves it as... maybe WTF-16, maybe UCS-2, maybe UTF-16, not entirely sure, didn't check. [2] If you need unicode, then the .rtf format lets you embed escaped unicode sequences inside the ANSI code page document :P