6 ms·
With or without a BOM?
by codeulike 10y ago
With or without a BOM?
- mcpherrinm 10y agoPutting a BOM in UTF-8 is just silly. Unlike -16, there's no option for which order you put the bytes in. The only time you'll see a BOM in UTF-8 is in poorly converted UTF-16.
- slavik81 10y agoApparently, Powershell requires a BOM to recognize UTF-8 scripts. https://github.com/chocolatey/choco/wiki/CreatePackages#character-encoding https://github.com/chocolatey/choco/wiki/CreatePackages#char...
- codeulike 10y agoYep, but a lot of MS software will only read UTF-8 correctly if a BOM is present.
- damienkatz 10y agoJoking? BOM is completely unnecessary in UTF8, only useful to losslessly preserve UTF16 text when converting back and forth.
- jcranmer 10y agoWithout. UTF-8 is such a distinctive pattern that if text with high bits set matches UTF-8, it's almost certainly UTF-8. There's no need for a BOM to tell you it's UTF-8 (looking at you, Windows), and it can easily confuse software instead.
- umanwizard 10y agoHuh? What would a BOM in UTF-8 even do? 1-byte objects can't have an internal byte ordering.
- the_mitsuhiko 10y agoIn UTF-8 a BOM can be placed to support round tripping the information with UTF-16.
- umanwizard 10y agoSo what order do you put the BOM in? Does it not even matter?
- Avernar 10y agoThe unicode BOM is code point U+FEFF. The process of encoding it determines the order. Encoded to UTF-8 it becomes EF BB BF. Encoding to UTF-16 big endian it will become FE FF. Encoding it to UTF-16 little endian it becomes FF FE. Converting it back from UTF-8 always gives you U+FEFF since UTF-8 doesn't care about endianess. Converting it back from UTF-16 using the correct endianess gives you U+FEFF. Converting it using the wrong endianess gives you U+FFFE which is defined by unicode as a "non character" that should never appear in text.
- umanwizard 10y agoMakes sense, thanks :)
- jcranmer 10y agoMS popularized the idea of adding the UTF-16 BOM into UTF-8 to distinguish between UTF-8 text files and Windows code page files, or what they called "Unicode" and "ANSI." There's (nearly?) unanimous agreement among everyone else that BOMs in UTF-8 text are really stupid. Note that the "BOM" in this case means storing the U+FEFF character in UTF-8 form (just as UTF-16 stores it in the appropriate endianness). This means that the result would be EF BB BF.
- Const-me 10y agoNot everything is a web request or response that have a “content-encoding” header transmitted somewhere out of band. The BOM allows to distinguish a byte stream between non-Unicode, UTF8, UTF16 and UTF32. Like it or not, but it's part of the standard: http://unicode.org/faq/utf_bom.html#BOM http://unicode.org/faq/utf_bom.html#BOM