4 ms·
And you trust a message from an unknown source? You can't simply memcpy according to some length indicator, that's just not safe. You still have to parse and va
by cptwunderlich 7y ago
And you trust a message from an unknown source? You can't simply memcpy according to some length indicator, that's just not safe. You still have to parse and validate.
- kingofpandora 7y agoGenuine question - is it dangerous to memcpy X bytes that we know must be interpreted as, say, an integer?
- deleted 7y ago[deleted]
- tangent128 7y agoPotentially. Most network protocols are big-endian while x86 is little-endian.
- byte1918 7y agoNo. Everything is 0s and 1s after all. Take for eg. a byte. It has 8 bits and by permutating all the 0s and 1s you end up with all the possible values of a signed byte; all the numbers -128 to 127. So now, if you were to copy a byte from a random memory location that byte will just contain a permutation of 0s and 1s which when interpreted as a signed int, will simply contain a number between -128 and 127.
- kingofpandora 7y agoThat's what I thought ...
- bmn__ 7y agohttps://capnproto.org https://capnproto.org serialisation scheme skips the decoding. Does that make it not safe?
- tebeka 7y agoNo serizalization is safe - https://docs.microsoft.com/en-us/security-updates/securitybulletins/2004/ms04-028 https://docs.microsoft.com/en-us/security-updates/securitybu... - https://en.wikipedia.org/wiki/Billion_laughs_attack https://en.wikipedia.org/wiki/Billion_laughs_attack - https://en.wikipedia.org/wiki/Zip_bomb https://en.wikipedia.org/wiki/Zip_bomb - ...
- strbean 7y agoNone of those are serialization schemes. XML can be used for serialization, but if you look at the whole ecosystem it is a Turing-complete complexity monster, so of course it isn't safe.
- DougBTX 7y agoIt depends on what constraints apply to the data. Any bit pattern could be used for an int, but to guarantee a UTF-8 string it would need to be validated.
- lwf 7y agoAs long as you know the length of the entire buffer, you just ensure that: current_addr + message_len - start_addr < buffer_len Or am I missing something?
- diabeetusman 7y agobuffer_len could be larger than the message, copying some incorrect things into memory. Similar to HeartBleed, where there wasn't validation on the heartbeat message, and the server would echo back buffer_len instead of just what was sent.
- theamk 7y agoI believe author intended buffer_len to be the length of incoming buffer (size of HTTP payload, number of bytes read from file, length of the database entry, etc...). So the worst that can happen is that entire input message is consumed -- like a JS payload which missed closing quote. I can think of a very contrived situation where this can be a problem, but in most cases this will be perfectly safe.
- maxwindiff 7y agoInvalid unicode sequences?
- kentonv 7y agoObviously you need to check that the length doesn't go past the end of the message, but that's a trivial O(1) check. You don't have to scan the bytes of the string first to decide if they are safe to memcpy.
- laumars 7y agoYou might want to validate those byte sequences are valid character encodings.
- kelnos 7y agoYou should be doing that with JSON as well, so this isn't a pro/con of either format.
- laumars 7y agoThat was obviously my point. ie just because MessagePack is a binary format it doesn't mean you can skip the same string checks that JSON requires; which means parsing MessagePack strings is unlikely to be any faster than JSON strings (contrary to the suggestions others have implied with the "just memcpy" comments). It's just with JSON that validation is done as part of the parser (remember JSON only technically supports a subset of ASCII and any extended characters or unicode is encoded via escape codes) where as with MessagePack you'd need to do that validation as an additional step. Integers, on the hand, might differ since JSON would need additional validation (again, backed into the parser) which MessagePack would not because MessagePack encodes the integers as binary integers where as JSON encodes them as ASCII values that would need converting back to binary integers. (hint: read the message I'm replying to).
- kentonv 7y agoMany (most?) applications do not actually care whether a byte blob of text is structurally valid UTF-8. They are either passing it around as an opaque byte blob, or already applying much stricter application-specific validation. Validating UTF-8 automatically at the serialization layer is a huge waste of cycles, especially in a big distributed system.
- hinkley 7y agoKeep fighting the good fight.