9 ms·
On Endianness (2021)
- thedanbob 4y agoLittle-endian makes more sense for computers where calculation is most important. Big-endian makes more sense for humans where, I would argue, comparison is most important. If I want to know e.g. if I can afford something, I'd prefer hearing the price as "four hundred and ..." to instantly get a ballpark rather than "five and ninety and four hundred".
- Taniwha 4y agoNo - I think that big-endian only makes sense because you grew up with it. Think about who used numbers back when this decision was made - small business people who mostly did math by hand - addition in particular - which we do from small digits to large - if we'd done the sane thing when we borrowed arabic numbers we'd write them in the order that numbers come out of an addition operation, rather than having to guess at how much space to leave for the result and write them backwards from the order we normally write them
- mathematicaster 4y ago
- cowtools 4y agoYou can make a program to convert it to some human-readable number then. You have to do this because it's binary anyways, so you're either representing it as octal, hexadecimal or binary. I don't know about you, but if I am trying to reason about hexadecimal numbers then I just separate it into 0xDEADBEEF = D*16^7 + E*16^6 + A*16^5 + D*16^4 + B*16^3 + E*16^2 + E*16^1 + F*16^0. The endianness only changes the order of the bytes I start reading at. What we ought to do is make a new prefix for reading the hexadecimal numbers in little-endian order like 0xDEADBEEF = 0yEFBEEDDE. Of course, this doesn't really fix the problem (wanting to read the number with smaller symbols first) as bytes are still considered to be in big endian if you consider the semantics about left/right shifts, which play on our preconceived notions of big endianness in everyday decimal math. You would want a system where everything is treated little-endian (bits within bytes, bytes within arrays)
- bshanks 4y agoI like the 0y format idea. Your point about bit shifts is good, too. If the rightmost bit is the least-significant bit, then for consistency perhaps memory should be written with the lowest memory address on the right. Then 0xDEADBEEF makes sense both in terms of bytes within the value and bits within the byte; the first 4 bits within the lowest byte in memory are F, then next 4 bits within the lowest byte are E, the first four bytes within the second-to-lowest byte are E, the next 4 bits within the second-to-lowest byte are B, etc. So if 0xDEADBEEF were stored in little-endian, a hex editor could display DE AD BE EF
- userbinator 4y agoI've always thought "little-endian is logical, big endian is backwards". In LE, bit n has value 2^n and byte n has value 256^n. In BE, bit n has value 2^(k-n) and byte n has value 256^(k-n) where k is the maximum length; it causes increase of ordinal position to not correspond with increase in value, and makes it length-dependent.
- nybble41 4y ago> In BE, bit n has value 2^(k-n) … Bit numbering does not necessarily follow byte numbering. Personally I favor BE byte-order—if only because it means a standard hex dump shows the bytes in normal reading order regardless of grouping, which is especially helpful when larger integers are not naturally aligned, and it also matches the order of digits in (English) text strings—but I would agree that bit 0 should always be the least-significant bit.
- bshanks 4y agoWith LE, you could either: - list memory right-to-left so the lowest memory address is on the right and the highest is on the left (i guess i prefer this solution because it's compatible with the convention of the least-significant-bit being the rightmost bit, which is enshrined in the name of "right shift", and with the convention of the least-significant-bit being bit 0, which makes sense in formulas like "value = sum_digitindex digit[digitindex]*radix^digitindex") or you could - write hexadecimal values with the most-significant-digit on the right, so the byte 2 followed by the byte 10 (in decimal notation) would be written 20 A0 instead of 02 0A. I don't know if popular hex dump programs have switches for either of those, though?
- nybble41 4y agoApplying the first solution consistently causes text strings to appear like "dlrow olleH", which IMHO is not very practical. In general we expect array elements, including byte and character arrays, to be listed in the locale's normal writing direction, which would generally be in ascending order by index from left to right. The second solution is more internally consistent, but breaks with the expected big-endian writing direction for Hindu-Arabic numerals in English text. I can't say I've encountered any hex editors which support either option.
- guidoism 4y agoA few weeks ago I read the original paper where the names big endian and little endian came from and was really surprised that it was published in 1980! So, before then we didn't have a common way of describing this I guess. It's crazy to me if true. Good paper BTW. Worth reading.
- drfuchs 4y agoBecause pretty much no machine in the universe used little-endian before Intel decided to save a cycle somewhere in the 8008 design, so there was no reason to have a way to refer to it. And even then, nobody paid much attention until IBM put an 8086 into the PC, rather than a Motorola 68X0, right about the time of the article you mention. It's been downhill ever since.
- userbinator 4y agoMultiple-precision arithmetic routines were almost always little-endian.
- guidoism 4y agoI agree with the sentiment but the RFC even mentions the PDP 11 as the first computer to claim being little endian. I wonder if the ascent of C and Unix at that time also contributed to it.
- guidoism 4y agoI’ve been writing a lot of MMIX lately and I really like big endian. It’s so much nicer to watch memory and registers change when ordered this way.
- throw0101a 4y agoThe root of the conflict lies much deeper than that. It is the question of which bit should travel first, the bit from the little end of the word, or the bit from the big end of the word? The followers of the former approach are called the Little-Endians, and the followers of the latter are called the Big-Endians. The details of the holy war between the Little-Endians and the Big-Endians are documented in [6] and described, in brief, in the Appendix. I recommend that you read it at this point. * Danny Cohen, IEN 137 (1 April 1980), https://www.rfc-editor.org/ien/ien137.txt https://www.rfc-editor.org/ien/ien137.txt
- drfuchs 4y agoOK, temporarily-victorious little-endians: Explain how you find it perfectly natural that the bits in a byte are big-endian? [Various replies (and votes) indicate I had better edit to clarify:] How come the single byte value written as 0x12 means "eighteen" and not "thirty-three"? Shouldn't the two four-bit nibbles also be considered as little-endian? Or it could even mean "seventy-two", if the bits are little-endian as a whole?
- TazeTSchnitzel 4y agoThe bits in a byte do not have a defined endianness!
- guidoism 4y agoI seem to recall discussion of this in the RFC. They do ultimately have an order if transmitted over a network! But most of us are used to getting them a byte at a time through the memory interface so we are ignorant of this.
- wtallis 4y ago> They do ultimately have an order if transmitted over a network! They do if transmitted over a serial port. But today's networks encode whole bytes or more into the symbols actually sent over the wire/air, so there's not really any way to point out on an oscilloscope trace where particular bits out of a byte are sent in a specific order.
- guidoism 4y agoWow! I didn’t know that
- AnimalMuppet 4y agoMost CPUs have assembly instructions called something like "shift left/right" or "rotate left/right". That implies an endianness, and it is (to my knowledge) always the endianness of normal written numbers - most significant is "left". That is, big endian.
- anonymousiam 4y agoIn the old days of satellite development (40 years ago at Hughes Space & Communications), I observed that the bit order in digital subsystems was always Big Endian, but all the other subsystems were usually Little Endian. Sometimes this caused issues, which fortunately were discovered during the integration & test phase. Usually problems like this were "corrected" by altering the documentation instead of the hardware so for a while, the majority of satellites in Earth orbit had a strange mix of endianness throughout the subsystems. The digital culture eventually won at Hughes/Boeing, but the standards committee punted and allows the endianness to be arbitrary (as long as it's documented) in the CCSDS standards. https://public.ccsds.org/default.aspx https://public.ccsds.org/default.aspx
- TazeTSchnitzel 4y agoA nice property of little endian is that you can index bits with (arr[x/8] >> (x%8)) & 1
- timbit42 4y agoYou think that code looks nice? You've been too deep into the code for too long.
- TazeTSchnitzel 4y agoWhy the mean comment? It's as nice as bit-indexing code can look. If you have to reverse the ordering of one of the indices it looks worse.
- CalChris 4y ago1. Network Byte Order is Big Endian. 2. x86, ARMv8 and RISC-V are Little Endian. 3. IBM 360 through s390x are Big Endian. 4. MIPS, PowerPC and ARM can be bysexual.
- arcticbull 4y agoARM is actually bi-endian too as of ARMv3. [edit] even ARMv8, the instruction fetch is always treated as little endian however the data access is implementation defined and therefore can also be bi-endian. [1] [1] https://developer.arm.com/documentation/102376/0100/Alignment-and-endianness https://developer.arm.com/documentation/102376/0100/Alignmen...
- bitwize 4y ago
- jfim 4y agoSome of these advantages are pretty dubious. Detecting whether a number is odd or even or getting the sign bit aren't impacted by endianness, as the CPU isn't working with single bytes at a time.
- samatman 4y agoOk, this is a huge reach, but: with big endian, you can determine the sign of the int at a pointer address without knowing the width of what's pointed at, with little endian, the parity. It's hard for me to call either of those advantages in anything but the most technical sense, but it's there.
- Const-me 4y ago> Variable length encoding with a prepended length field The OP's method only works under the following assumptions: (1) length in bytes is power of 2 so the re-interpret trick works, and (2) the complete source data is in memory. In many popular formats (MKV, WebM, etc) the header is variable-length as well e.g. 1-8 bits, the complete integer is therefore [ 1 .. 8 ] bytes, and because these files can be very large they don't fit in memory, need to stream from disk. Under these conditions, the big-endian format (a) eliminates need for any bit shifts, only need to mask away the header in higher bits, slightly faster on most CPUs (b) preserves original bytes in the number except the MSB which contain the header, helps with debugging. I think that's the only use case where big endian has a substantial advantage. Fortunately, all modern CPUs have fast instructions to flip the endianness, e.g. on Intel/AMD they are bswap for 4 or 8 bytes, ror/rol for 2 bytes, available as intrinsic or standard library functions in most programming languages. For all other cases, including binary formats I design, I only use little-endian convention. Little-endian CPUs have won, and I don't like doing unnecessary work flipping these bytes. I'm lazy, also a boilerplate like that is a common source of bugs.
- ZephyrBlu 4y agoThis is completely unrelated to the contents of your comment, but did you use (a) to denominate the first item and then (2) to denominate the second in your third paragraph?!
- watersb 4y agoNot the poster, but my Hacker News comments are sometimes written on my phone. Small screen. Limited viewport. Primitive text. Wrestling mightily with the autocorrect vs. technical terminology... My brain working set often loses things like proper list enumeration. :-)
- Const-me 4y agoThanks, fixed.
- ffwszgf 4y agoWas Hindu (I assume that’s the original language for which Indo-Arabic numerals were made for) written right to left ? I know Arabic was but I don’t think Indian languages ever have been.
- gumby 4y agoThe article has some bugs. Alphabets in India (for as far back as we have written records, e.g. Pali) were all LTR until Muslim rule was established at which point some people began to write Hindi using the Persian/Arab script (giving rise to Urdu, a very close sister language to Hindi). Also network order is big endian because most of the machines back then were big endian, most notably the PDP-10 which was the common research machine in that era. Big endian offered simplicity in implementation (remember the earlier models of these machines were mostly hand made, even if made in a factory, with only a few semiconductors; the ALUs and instruction decoding was all done with wires, not traces). Bytes weren't necessarily 8 bits wide and while it was trivial (easier than in C actually) to do pointer arithmetic, a "pointer cast" is a weird way to think of it. So the article is full of the assumption that the world is basically a PDP-11. I think the obsession with such machines has held computing back as much as it has sped it up.
- aap_ 4y agoHow is the PDP-10 big endian if it has no byte addressable memory? I would agree that culturally it's in the big endian camp because bits are numbered MSB to LSB, but I see no technical reason.
- gumby 4y agoBytes are very much addressable, they are simply not limited to only being 8 bits. The 8-bit convention was uncommon when these machines were designed -- I think it was only on some IBM machines at the time. Remember the 10s, like several contemporary machines, had 36-bit words with an 18-bit address space (yes Gordon Bell specifically designed them with Lisp in mind) As for the technical reason, consider wiring up the arithmetic unit by hand (work it out on a piece of paper). If you don't have an architecture book handy, look at Ken Sheriff's walkthrough of the Z80 die and consider some of the extra complexity (less of an issue of course for an MPU that was an IC) to handle little endian arithmetic ops (routing carry and such).
- aap_ 4y agoNot sure I get your point. The PDP-10 has no notion of bytes in the arithmetic unit, it's all words (and hence there's no endianness). The byte instructions are handy ways to load/store parts of words, sure, but I feel like they're more similar to modern bitfield instructions and feel rather tacked on (they were optional in the KA10 even) rather than being a central part of the architecture. Come to think of it though, byte pointers move right in the word as they're incremented. Guess you could count that as big endian.
- deleted 4y ago[deleted]
- cryptonector 4y agoTFA's claims about advantages are nonsense. E.g., for "detecting odd/even" it gives the advantage to little-endian but without considering at all how that might be implemented in hardware, or even without explaining at all why there's an advantage to be had. Or when talking about bignums: > Although this scheme can be realized with either byte order, there is an extra advantage to little endian byte ordering: If the CPU is little endian, you wouldn’t even need to care about the element size in the array because the bytes would naturally arrange themselves smoothly in little endian order across the entire array. But the same is true for big-endian! The array indices will run the opposite way, but so what? This seems aimed at forcing the conclusion that little-endian is better. There's no real advantage to one or the other. The world just has both, and we have to deal with it.
- LargoLasskhyfv 4y agoWhat do you think about the fact that the linked rfc from [-] https://www.rfc-editor.org/ien/ien137.txt https://www.rfc-editor.org/ien/ien137.txt dates from APRIL the FIRST 42 years ago? Could/would/should it matter?
- IshKebab 4y agoThis is a decent attempt. I like that they actually at least try to find pros/cons. But it does seem odd that they give "convention" to big endian despite the fact that basically all modern processors are little endian and there's absolutely no advantage to following "network byte order" if you don't have to.