4 ms·
I thought I did well until the bit-shifting questions. Who in their right mind designs a language where shifting the bits of a u16 silently converts into an i32
by codeflo 4y ago
I thought I did well until the bit-shifting questions. Who in their right mind designs a language where shifting the bits of a u16 silently converts into an i32? Doubly so since that fact alone directly causes UB — keeping the u16 would have been perfectly fine.
(Edit, since some respondents seem to miss this, explanations about efficiency or ISAs might justify promoting to u32 (though even that's debatable), but not i32. A design that auto-promotes an unsigned type, where every operation is nicely defined and total, into a signed type, where you run into all kinds of undefined behavior on overflow, is simply crazy.)
- veltas 4y agoBecause everything smaller than an int is usually promoted to int. int is the 'word' in C that things are calculated in. Even character constants are ints, all enum constants are ints, the default type was int when default types were still a thing.
- jefftk 4y agoBut why not turn u16 into u32? Why switch it to being signed on promotion?
- rwmj 4y agoI'm sure the answer is going to be along the lines of "because PCC did that and they standardized it" :-/ Here's a fun standardization problem I came across recently (nothing to do with C): http://mywiki.wooledge.org/BashFAQ/105 http://mywiki.wooledge.org/BashFAQ/105
- veltas 4y agoBelieve it or not, this makes the behavior more like what you'd expect in many cases. For example: uint8_t x = 4; extern volatile uint64_t *reg; *reg &= ~x; In the last statement x is promoted to an int, and then when the logical NOT occurs every bit is set to 1, including the high bit. When it's converted to a uint64_t for the AND, the high bits are also set to 1. So the result is that the final statement clears only bit 2 in *reg. If it promoted to unsigned int, then it would also clear bits 32-63.
- codeflo 4y agoI don’t find that very convincing. It’s simply ambiguous, I might want this: *reg &= ~(uint64_t)x; or, and there’s no elegant way to even write this in C, I might want: *reg &=(uint64_t)(uint8_t)~x; The fact that I have to write two casts here to undo the damage of the auto-promotion is evidence of how broken this is.
- veltas 4y agoThat second line can be written: *reg &= (uint8_t)~x; Or: *reg &= ~x & 0xFF;
- tialaramex 4y agoextern volatile uint64_t *reg; *reg &= ~x; People should stop doing this. What this means is: extern volatile uint64_t *reg; uint64_t tmp = *reg; tmp &= ~x; *reg = temp; But of course when you write that chances are somebody will point out that you're running in interruptible context sometimes in this function, so that's actually introducing a race condition. Why didn't they say so when you wrote it your way? Because that looked like a single operation and so it wasn't obvious it might get interrupted.
- ynfnehf 4y agoANSI describes why in their rationale: https://www.lysator.liu.se/c/rat/c2.html#3-2-1 https://www.lysator.liu.se/c/rat/c2.html#3-2-1 . The unsigned preserving rules greatly increase the number of situations where unsigned int confronts signed int to yield a questionably signed result, whereas the value preserving rules minimize such confrontations. Thus, the value preserving rules were considered to be safer for the novice, or unwary, programmer. After much discussion, the Committee decided in favor of value preserving rules, despite the fact that the UNIX C compilers had evolved in the direction of unsigned preserving.
- edflsafoiewq 4y agoYes. The model C has is that the CPU has an ALU that that operates as (word,word)->word with int being the smallest word size. This explains many of C's conversion rules: to operate on a single integer, it first has to be promoted to a word size; to operate on two integers, they first have to be converted to a common word size, etc.
- lifthrasiir 4y agoAnd that makes writing a correct and portable C futile. Yes, defined-size types are a thing and I almost exclusively use them, but since the size of `int` itself is unknown and integer promotion depends on that size, the meaning of the code using only defined-size types can still vary across platforms. intN_t etc. are only typedefs and do not form their own type hierarchy, which in my opinion is a huge mistake.
- jcranmer 4y agoIf you have a 32-bit architecture, you don't necessarily have 8-bit and 16-bit hardware operations outside of memory operations (load/store). Now, I don't find this reasoning persuasive--it's not that hard to emulate an 8-bit or 16-bit operation--and judging from the history of post-C languages, most other language designers are equally unmoved by this reasoning, but I can see someone in their right mind designing a language that acts like this. Especially if the first architecture they're developing on is precisely such on architecture (PDP11 doesn't have byte-sized add/sub/mul/div).
- simias 4y agoI think the prevailing philosophy for C (and later C++) was that code should map closely to the hardware and not expand to complicated "microcode". A left shift in Cshould just be a left shift in the underlying ISA, within reason. That's where most of the undefined behaviours come from. Having to add masking and other operations to emulate a 16bit shift on a 32bit architecture feels un-C-like, for better or worse. IMO the real issue is not so much the fact that all shifts of any type < int is treated as if it were an int, it's that the language doesn't force you to acknowledge that in the code. If you got a compilation error when trying to shift a short and had to explicitly promote to int in order to make it through, at the very least it can't lead to an oversight from a careless programmer. C is trying to be clever but only goes half way, resulting in the worst of both worlds IMO.
- codeflo 4y agoThe ISA might force someone to extend a value to 32 bits (debatable, but let’s go with it). It never forces you to treat an unsigned int as signed. It also doesn’t require inserting UB into the process.
- simias 4y agoI agree, the whole signed vs. unsigned generally feels like an afterthought in C (probably because, to a certain extent, it was). `char`'s sign being implementation-defined is a pretty wild design choice that wasted many hours of my life while porting code between ARM and x86. UBs are not required, but you need them if you want C to behave as a macro-assembler as well as allowing for aggressive optimizations. For instance `a << b` if b is greater than a's width is genuinely UB if you write portable code, different CPUs will do different things in this situation. Defining the behaviour means that the compiler would have to insert additional opcodes to make the behaviour identical on all platforms. You may argue that it's still better than having UB but that's just not C's design philosophy, for better or worse.
- nayuki 4y ago> Who in their right mind designs a language where shifting the bits of a u16 silently converts into an i32? Yeah, hence why I asked this question years ago: https://stackoverflow.com/questions/39964651/is-masking-before-unsigned-left-shift-in-c-c-too-paranoid https://stackoverflow.com/questions/39964651/is-masking-befo...
- scatters 4y agoBecause sub-integer types are for storage, not computation. Yes, it'd be better if you had to explicitly cast to int or unsigned to perform arithmetic, but that ship has sailed.