3 ms·
The BOM is in-band signalling. It breaks the ASCII-backwards-compatability of UTF8. BOMs are only necessary where the provenance of data is not known. Normally
by jbert 14y ago
The BOM is in-band signalling. It breaks the ASCII-backwards-compatability of UTF8.
BOMs are only necessary where the provenance of data is not known. Normally, there is a context provided which will determine the encoding.
Basically, yes - text should be tagged (unless 'utf8 everywhere' wins) but imho the tagging should be external to the content.
- makecheck 14y agoI agree with the use of context, yes; if you already know your input is UTF-8 (e.g. C strings in a program API or a protocol or whatever) there's no point in adding an extra specifier. If a program requires ASCII compatibility in order to work then by all means make the input files ASCII (no BOM), just make sure the files have no true UTF-8 dependencies in them. Once a program supports UTF-8 "properly" however, the BOM is useful as a signal that the input is somewhat complicated. At some point in the future when UTF-8 really is everywhere and programs may no longer even try to sniff encodings, etc. then yes, the BOM has no real reason to exist.
- jbert 14y agoBut you have the same problem with text or binary distinctions. On some platforms, no distinction between them needs to be made (quite usefully, adding to tool simplicity and conceptual simplicity). If you use BOM, a tool which can operate on text or binary must be told which it is operating on. This would have to be done via some external context (e.g. cmdline switch). And that would never go away, even in a utf8 everywhere world. Basically, in BOM-world: - You either need to tag each fragment of text or you still need to use external context (e.g. which encoding do I get text columns from my database) - You perpetually need to differentiate between binary and text data for all tools which do nothing more complicated than read and write and in non-BOM-world you: - add to the contextual clues you need anyway something like "any files which you are going to interpret as text on the system should be interpreted as utf8" - when moving data on or off the local the local system, use a network protocol which supports tagging the text payload (e.g. email, http). The problems arise mostly with file shares (or their equivalent, version control systems) where text files are exchanged without an accompanying protocol. That is where "BOM world" or "UTF8 world" will ultimately have to settle their differences. BOM-world would like all systems, everywhere, to make a text/binary distinction for ever. UTF8-world would like to say that textual data lacking a context should be interpreted as UTF8. But feel free to use UTF16/UTF32 for specific purposes or systems.