4 ms·
> On encoding the utf-8-sig codec will write 0xef, 0xbb, 0xbf as the first three bytes to the file https://docs.python.org/3/library/codecs.html https://docs.p
by shpx 2y ago
> On encoding the utf-8-sig codec will write 0xef, 0xbb, 0xbf as the first three bytes to the file
https://docs.python.org/3/library/codecs.html https://docs.python.org/3/library/codecs.html
The codec you're imagining would also make reading a file and writing it back change the file if it contains a BOM.
- int_19h 2y agoIndeed it would, but since codecs are only used for files that are semantically text, and in such files BOM is basically a legacy no-op marker, it's not actually a problem. Naive code using text I/O APIs would also have this issue with line endings, for example, so there's precedent for not providing the perfect roundtrip experience (that's what bytes I/O is for).