4 ms·
Little-endian is slightly more confusing for humans I've heard this before, but the reason is that you view hex data and list numbers left-to-right as if they
by jfries 9y ago
Little-endian is slightly more confusing for humans
I've heard this before, but the reason is that you view hex data and list numbers left-to-right as if they were letters. They are not.
0x12345678 stored big-endian, numbering bytes left-to-right:
12 34 56 78
Looks good, but I think that this is actually more confusing, because when you number the bytes and bits you will see that the bytes are written left-to-right, while at the same time the bits are written right-to-left.
The solution is to show dumps with bytes numbered right-to-left. This is coherent with how we number bits, and also how we relatively position digits in any other number.
0x12345678 in little-endian, but written right-to-left:
12 34 56 78
Now the numbering of the bytes is consistent with the numbering of bits. You can easy see that bits 0-7 of the bits interpreted as a 32-bit word belong in byte 0, while bits 16-23 belong in byte 2.
- loeg 9y agoYour scheme to byteswap little endian in hex displays falls over in the face of differing word sizes.
- jfries 9y agoNot at all, you just list the whole dump right-to-left. Then it works out regardless of word size.
- loeg 9y agoSo the dump is ordered backwards, from top of memory to bottom? That seems harder for humans than little endian integers.
- jfries 9y agoI don't understand what you mean. Line breaks can be inserted where ever suitable. Whether the bytes are listed left-to-right or right-to-left makes no difference.
- kazinator 9y agoA dump of memory (several kilobytes, megabytes or whatever) is inherently big endian: it proceeds from the base address and goes up. It is counterintuitive to swap pieces of it into some locally opposite order. Yet, that's what has to be done so that numbers are readable. That's why "od" has modes for that. $ od -tx1 /bin/ls | head -1 0000000 7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00 $ od -tx2 /bin/ls | head -1 0000000 457f 464c 0102 0001 0000 0000 0000 0000 $ od --endian=big -tx2 /bin/ls | head -1 # GNU extension, probably. 0000000 7f45 4c46 0201 0100 0000 0000 0000 0000
- jfries 9y agoA dump of memory (several kilobytes, megabytes or whatever) is inherently big endian: it proceeds from the base address and goes up. It only looks "inherently big endian" when you print bytes on each line starting from the left. This way of printing numbers makes little sense as you end up with some kind of mixed-endian where the bytes are ordered one way and the bits another way. Start each line with bytes from the right instead to make the bytes and bits numbering consistent, and you can print large little-endian dumps with all bits and bytes are where you expect them to be.
- kazinator 9y agoEverything is consistent in the big (down to the nybble and bit) endian view. The number 0x12345678 is actually the nybbles 1 2 3 4 ... which are the bits 0001 0010 0011 0100 and so on. In the hex dump 12 34 56 78 we just understand the bytes to be big-endian also at the nybble level (and bit also). That is to say the "1" can be understood to be at the lower "nybble address", relative to the "2". If that buffer is sent over a serial communication channel or network, the bits actually go in that order 0001 0010 0011. The 1 nybble goes out first, as 0001 (three zeros out the door, then a one), then the 2 nybble and so on.
- spc476 9y agoIt depends. with RS-232 (the old serial port standard) bits were transmitted least-significant-bit first, with the most significant bit sent last [1]. Getting back to hex dumps, here's one of some data: 00000000: 6C 6F 77 09 30 0A 66 72 65 65 09 31 32 32 0A 65 low.0.free.122.e You can see it's ASCII. A little endian dump of that would be e.221.eerf.0.wol 65 0A 32 32 31 09 65 65 72 66 0A 30 09 77 6F 6C :00000000 It makes reading any text in binary data a bit difficult. [1] My first computer was a Tandy Color Computer, which had a serial port driven directly by the CPU. I learned pretty quickly which bit goes first.
- phs2501 9y agoAt least in Freescale's PowerPC documentation, it's convention to number the bits left-to-right in big-endian. So the most-significant bit is bit 0, which matches up with the most-significant big-endian byte being 0. See, for a random example, page 1101 of https://www.nxp.com/docs/en/reference-manual/MPC8379ERM.pdf https://www.nxp.com/docs/en/reference-manual/MPC8379ERM.pdf. Personally I prefer the little-endian representation.
- kazinator 9y agoThe documentation might have that numbering, but that makes no difference in programming on the PowerPC. When we take the value 1 and shift left by 1 bit, we get 2. That the documentation thinks this is bit 7 going to bit 6 is immaterial. Calling the MSB "bit 1" is a tip of the hat to serial communications. In serial communication and networking, it is predominant to transmit the MSB first. If the documentation is about a wire format, then using that numbering is correct down to the data link layer and (modulo framing considerations and such), physical.
- phs2501 9y agoOh, I'm aware it makes no difference internally; I picked that documentation because I am somewhat intimately familiar with it as I worked on a product with that CPU and did a lot of driver work. My only point was to the poster I was replying to, who wrote "because when you number the bytes and bits you will see that the bytes are written left-to-right, while at the same time the bits are written right-to-left". That assumes that bits are universally numbered starting with 0 at the LSB, which isn't true.
- ajdlinux 9y agoMSB0 is convention in basically all IBM documentation (including PowerPC/Power Architecture stuff), hence all the Freescale/NXP PPC manuals follow. The main practical problem with MSB0 is that you need to consider the width of whatever field or register you're looking at to work out the correct bit shift.
- kazinator 9y agoWhen 0x12345678 is stored in little endian, it looks like 78 56 34 12 in a byte dump, which is stupid because hex digits are grouped as pairs and then revered. The endianness at the bit level is irrelevant because the bits are chunked into bytes, and are usually not even addressable. A storage format exhibits endianness only when it is addressable. In the C language, bits are only "addressable" via the shift operators. These are rooted in pure arithmetic, so that 1<<1 is always 2, regardless of whether you're on a big or little endian platform. Whether you call the value 1 "bit 7", "bit 8", "bit 1" or "bit 0" is just, pardon the pun, word semantics. The only time you deal with "bit endianness" is with certain compressed data formats, and with bitfield layout rules.
- tetrep 9y ago> Whether you call the value 1 "bit 7", "bit 8", "bit 1" or "bit 0" is just, pardon the pun, word semantics. It's extremely important to your mental model to understand how the bits are arranged, or left shift (<<) is going to produce different results in your head. If you think the value 1 is bit 7/8, you get: 1 0 0 0 0 0 0 0 Which would left shift to: 0 0 0 0 0 0 0 0 Which would not be equal to 2.
- kazinator 9y ago> It's extremely important to your mental model to understand how the bits are arranged And the shortcut for that is simply regard bytes as big-endian. On all platforms. Whether you're on a PPC or x86, the byte 0x80 is going to go out on the wire as 1 first, followed by 7 zeros. So if you're writing a data compressor and the spec says that the variable-length bit strings (huffman or whatever) are stuffed into bytes in network order, that means you fill bytes from the left down. That is done with code that works the same way on BE or LE platforms.