3 ms·
SHOULD include a zero byte I guess it is expected to be at the end of the magic number to act as a null-termibated string? MUST include a byte sequence tha
by conaclos 2y ago
SHOULD include a zero byte
I guess it is expected to be at the end of the magic number to act as a null-termibated string?
MUST include a byte sequence that is invalid UTF-8
I guess it is to differentiate a text file from a specific format?
MUST include at least one byte with the high bit set
Any reason?
- wjholden 2y agoThe author explains their reasoning in the next post: https://hackers.town/@zwol/114155807716413069 https://hackers.town/@zwol/114155807716413069
- robinhouston 2y agoI think the idea of all these is to make the file not be recognised as text (which doesn't allow nulls), ASCII (which doesn't use the high bit), UTF-8 (which doesn't allow invalid UTF-8 sequences). Basically so that no valid file in this binary format will be incorrectly misidentified as a text file.
- CrossVR 2y agoI think preventing the opposite is more pressing. Imagine creating a text file and it just so happens that the first 8 characters match a magic number of an image format. Now when you go back to edit your text file it is suddenly recognized as an image file by your file browser.
- rcxdude 2y agoWell, the 3rd point follows from the second: all sequences without the high bit set are valid ASCII, and all valid ASCII sequences are valid UTF-8.
- dark-star 2y agothe high bit one is pretty ancient by now. I don't think we have transmission methods that are not 8-bit-clean anymore. And if your file detector detects "generic text" before any more specialized detections (like "GIF87a"), and thus treats everything that starts with ASCII bytes as "generic text", then sorry, but your detector is badly broken There's no reason for the high-bit "rule" in 2025. I would argue the same goes for the 0-byte rule. If you use strcmp() in your magic byte detector, then you're doing it wrong
- 7jjjjjjj 2y agoThe zero byte rule has nothing to do with strcmp(). Text files never contain 0-bytes, so having one is a strong sign the file is binary. Many detectors check for this.
- dark-star 2y ago> Text files never contain 0-bytes that might be true for ASCII but there are other text encodings out there And again, if a detector doesn't check for the more specific matches first, before falling back to "ah, that seems to be text", then the detector is broken
- Joker_vD 2y ago> I don't think we have transmission methods that are not 8-bit-clean anymore. I've just dealt with a 7N1 serial link yesterday, they still exist. Granted, nobody really uses them for truly arbitrary data exchange, but still.
- layer8 2y agoGit interprets a zero byte as an unconditional sign that a file is a binary file [0]. With other “nonprintable” characters (including the high-bit ones) it depends on their frequency. Other tools look for high bits, or whether it’s valid UTF-8. PDF files usually have a comment with high-bit characters on the second line for similar reasons. These recommended rules cover various common ways to check for text vs. binary, while also aiming to ensure that no genuine text file would ever accidentally match the magic number. The zero-byte recommendation largely achieves the latter (if one ignores double/quad-byte encodings like UTF-16/32). [0] https://github.com/git/git/blob/683c54c999c301c2cd6f715c411407c413b1d84e/convert.c#L98 https://github.com/git/git/blob/683c54c999c301c2cd6f715c4114...
- Retr0id 2y agoHaving a high-bit set allows you to immediately detect if a file has been mangled through transmission over a 7-bit-only medium.