14 ms·
Should I Use Signed or Unsigned Ints?
- jschwartzi 11y agoSeems like it boils down to using unsigned when you need to treat numbers as fields of bits, and signed when you need to do arithmetic.
- joe_the_user 11y agoNot experiencing undefined behavior is a desirable thing for arithmetic so unsigned has some appeal for arithmetic too, I'd say.
- jhdevos 11y agoThat seems like a false sense of safety: when doing arithmetic, your program is most likely not going to be handling overflows correctly anyway. Also, when using unsigned ints, any code involving subtractions can lead to subtle bugs that only occur when one value is larger than another. I'd recommend sticking with signed ints for all arithmetic code.
- ridiculous_fish 11y ago> when doing arithmetic, your program is most likely not going to be handling overflows correctly anyway Well with that attitude!
- jschwartzi 11y agoIt might be pretty safe if you're only ever adding, multiplying or dividing with other unsigned numbers. Once you start doing subtraction it could get ugly.
- olympus 11y agoWhile the signed vs unsigned int question doesn't concern me much[1], I really appreciate this post because I discovered the One Page CPU. The author nerd-sniped me three paragraphs in. Now I want to try to implement this on an FPGA. I think my next few weekends will be taken up with getting a real world implementation working and getting some examples working. Thanks. [1] If I was designing a system that could only have one type of int (but why?) I'd use unsigned ints and if I needed to represent negative numbers, I'd use two's complement which is fairly well behaved, much as the author points out. This is fairly common in the embedded world.
- ridiculous_fish 11y agoThis is good general advice. A key difference is whether the type is meant to be an index (counting) or for arithmetic. Indexing with unsigned integers certainly has its pitfalls: while (idx >= 0) { foo(arr[idx]); idx--; } But this is outweighed by the enormous and under-appreciated dangers of signed arithmetic! Try writing C functions add, subtract, multiply, and divide, that do anything on overflow except UB. It's trivial with unsigned, but wretched and horrible with signed types. And real software like PostgreSQL gets it wrong, with crashy consequences: http://kqueue.org/blog/2012/12/31/idiv-dos/#sql http://kqueue.org/blog/2012/12/31/idiv-dos/#sql
- revelation 11y agoWell, there are builtin functions that do this safely for signed integers. Of course that means you need to be aware which calculations are potentially dangerous, you need to know about signed integer overflow problems in the first place, etc. ..
- michaelhoffman 11y agoWhat functions are those?
- kevinnk 11y agoFor gcc see https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins.html https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins...
- jchomali 11y agoI totally agree.
- nly 11y ago> Try writing C functions add, subtract, multiply, and divide, that do anything on overflow except UB. For those curious... https://www.securecoding.cert.org/confluence/display/c/INT32-C.+Ensure+that+operations+on+signed+integers+do+not+result+in+overflow https://www.securecoding.cert.org/confluence/display/c/INT32...
- dietrichepp 11y agoThere are a lot of corner cases involved here. The corner cases for signed integers involve undefined behavior, the corner cases for unsigned integers involve overflow. Any time you mix them you get unsigned integers, which can give you big surprises. For example, if z is unsigned, then both x and y will get converted to unsigned as well. This can cause surprises when you expected the LHS to be negative, but it's not, because the right side is unsigned, and that contaminates the left side. if (x < y * z) { ... } This particular case gets caught by -Wall, but there are plenty of cases where unintended unsigned contagion doesn't caught by the compiler. Of course, if you make x long, then y * z will be unsigned, then widened to long, which gives you different results if you are on 32-bit or if you are on Windows. Using signed integers everywhere reduces the cognitive load here, although if you are paranoid, you need to do overflow checking which is going to be a bear and you might want to switch to a language with bigints or checked arithmetic. As another point, in the following statement, the compiler is allowed to assume that the loop terminates: for (i = 0; i <= N; i++) { ... } Yes, even if N is INT_MAX. The way I think of it, your use of "signed" means that you are communicating (better yet, promising) to the compiler that you believe overflow will not occur. In these cases, the compiler will usually "do the right thing" when optimizing loops like this, where it sometimes can't do that optimization for unsigned loop variables. So I'm going to disagree as a point of style. Signed arithmetic avoids unintended consequences for comparisons and arithmetic, and enables better loop optimizations by the compiler. In my experience, this is usually correct, and it is rare that I actually want numbers to overflow and do something with them afterwards. All said, if you don't quite get the subtleties of arithmetic in C (yes, it is subtle) then your C code is fairly likely to have errors, and no style guide is going to be a panacea.
- faragon 11y agoNo matter if you use signed or unsigned types, be sure you handle both underflow and overflow. From my experience, signed integers main usage is integer arithmetic (e.g. accounting things that can be negative or positive), pointer arithmetic, error codes, or multiplexing on same variable two different kind of elements (so if negative has one meaning and if positive, a different one). Unsigned, for handling countable elements, with 0 as minimum. The good point of using unsigned integers is that you have twice the available codes, because the extra bit. Example for counting backwards using unsigned types (underflow check, -1 is 111...111 binary in 2's complement representation, so -1 is equivalent to the biggest unsigned number): size_t i = ss - 1, j = sso; for (; i != (size_t)-1; i--) { switch (s[i]) { case '"': j -= 6; s_memcpy6(o + j, """); continue; case '&': j -= 5; s_memcpy5(o + j, "&"); continue; case '\'': j -= 6; s_memcpy6(o + j, "'"); continue; case '<': j -= 4; s_memcpy4(o + j, "<"); continue; case '>': j -= 4; s_memcpy4(o + j, ">"); continue; default: o[--j] = s[i]; continue; } } Example for compute percentage on unsigned integers (same idea can be applied for signed) without losing precision nor requiring bigger data container (overflow checks): size_t s_size_t_pct(const size_t q, const size_t pct) { return q > 10000 ? (q / 100) * pct : (q * pct) / 100; } Also, for cases where overflow could happen, you can handle it doing something like this: sbool_t s_size_t_overflow(const size_t off, const size_t inc) { return inc > (S_SIZET_MAX - off) ? S_TRUE : S_FALSE; } size_t s_size_t_add(size_t a, size_t b, size_t val_if_saturated) { return s_size_t_overflow(a, b) ? val_if_saturated : a + b; } size_t s_size_t_mul(size_t a, size_t b, size_t val_if_saturated) { RETURN_IF(b == 0, 0); return a > SIZE_MAX / b ? val_if_saturated : a * b; } P.S. examples taken from: https://github.com/faragon/libsrt/blob/master/src/senc.c https://github.com/faragon/libsrt/blob/master/src/senc.c (underflow example) https://github.com/faragon/libsrt/blob/master/src/scommon.h https://github.com/faragon/libsrt/blob/master/src/scommon.h (overflow examples)
- SamReidHughes 11y ago> size_t i = ss - 1, j = sso; for (; i != (size_t)-1; i--) { The idiomatic way is for (size_t i = ss, j = sso; i-- > 0; ) { ... }
- 0x0 11y agoFunnily enough the bash crash in the comment still crashes bash on OSX 10.10.4: ($((-2**63/-1)))
- ah- 11y agoProbably because Apple hasn't really updated bash since the license switch to GPLv3.
- mmaunder 11y agoUnsigned int's for storing IPv4's FTW. Or conversely, storing IP's as signed int's will make you sad.
- dfbrown 11y agoDespite the fact that signed overflow/underflow is undefined behavior I'm pretty sure that many more bugs have resulted from unsigned underflow than signed overflow or underflow. When working with integers you're usually working with numbers relatively close to zero so it's very easy to unintentionally cross that 0 barrier. With signed integers it is much more difficult to reach those overflow/underflow limits.
- nickff 11y agoBut you have to balance the probability of a problem with its severity. I do a lot of work in C with pointers and arrays, and there are many situations where I select unsigned ints because (, although problems are more likely,) there are less likely to cause a memory read or write error (for my application, though not necessarily for others). I also often use unsigned when I am doing math that includes logs, because I am less likely to cause a math error. I would like to be clear that I agree with your assessment for most situations, with some exceptions.
- modeless 11y agoUnsigned overflow is defined but that doesn't necessarily make it any more expected when it happens. It can still ruin your day, and in fact it happens much more often than signed overflow, because it can happen with seemingly innocent subtraction of small values. After personally fixing several unsigned overflow bugs in the past few months I'm going to have to side with the Google style guide on this one.
- Peaker 11y agoIf you underflow an unsigned you get a huge value that's going to be invalid by any range check unless you assume entire unsigned range is available. But if you use signed, you have to check for values under zero AND the top of the range. So I don't understand how the bugs were caused by unsigned overflow. Wouldn't they still be bugs with negative results in signed numbers?
- modeless 11y agoMostly the errors are expecting that a calculation is using signed arithmetic when it's not, and getting a large positive value where you expected a small negative value. It's easy to forget the signedness of a variable, but the biggest problem is probably C's unfortunate tendency to silently convert signed to unsigned when at least one unsigned value is involved in a calculation. This causes unsigned arithmetic to contaminate parts of your code that you thought were signed.
- Peaker 11y agoI think I have a handful of negative numbers in hundreds of thousands of lines of C code. It's quite rare, IME. By the way, an underflowed/large unsigned number used when a signed number is expected in most contexts (e.g: when adding as an offset) will behave correctly. It will fail when you try to compare it using < or >. If you use enable gcc warnings (which you ought to for every conceivably useful warnings :-) ), that will be caught.
- kazinator 11y agoComplete newbie programmer naivete, I'm afraid. Unsigned integers have a big "cliff" immediately to the left of zero. Its behavior is not undefined, but it is not well-defined either. For instance, subtracting one from 0U produces the value UINT_MAX. This value is implementation-defined. It's safe in that the machine won't catch on fire, but what good is that if the program isn't prepared to deal with the sudden jump to a large value? Suppose that x and y are small values, in a small range confined reasonably close to zero. (Say, their decimal representation is at most three or four digits.) And suppose you know that x < y. If x and y signed, then you know that, for instance, x - 1 < y. If you have an expression like x < y + b in the program, you can happily change it, algebraically to x - b < y if you know that overflow isn't taking place, which you often do if you have assurance that these are smallish values. If they are unsigned, you cannot do this. In the absence of overflow, which happens away from zero, signed integers behave like ordinary mathematical integers. Unsigned integers do not. Check this out: downward counting loop: for (unsigned i = n - 1; i >= 0; i--) { /* Oops! Infinite! */ } Change to signed, fixed! Hopefully as a result of a compiler warning that the loop guard expression is always true due to the type. Even worse are mixtures of signed and unsigned operands in expressions; luckily, C compilers tend to have reasonably decent warnings about that. Unsigned integers are a tool. They handle specific jobs. They are not suitable as the reach-for all purpose integer.
- deleted 11y ago[deleted]
- ridiculous_fish 11y ago> Suppose that x and y are small values, in a small range confined reasonably close to zero. (Say, their decimal representation is at most three or four digits.) There's the rub! You need to justify characterizing your values in that way, which means either explicit range checks and assertions, or otherwise deriving them from something that applies that guarantee in turn. And that justification is more work than just making your code correct for every value. I mean, if x and y are both smallish, is x*y also smallish? The "nearby cliff" is a good thing, in that it makes errors come out during testing rather than a month after you ship. Handwaving about "reasonably close to zero" is begging for trouble. In the absence of automatic bigints, unsigned integers are easier to make correct.
- jerf 11y agoAs I learned from Haskell, the correct answer really ought to depend on the domain. I ought to be able to use an unsigned integer to represent things like length, for instance. However, since apparently we've collectively decided that we're always operating in ring 256/65,536/etc. instead of the real world where we only operate there in rare exceptional circumstances [1], instead of an exception being generated when we under or overflow, which is almost always what we actually want to have happen, the numbers just happily march along, underflowing or overflowing their way to complete gibberish with nary a care in the world. Consequently the answer is "signed unless you have a good reason to need unsigned", because at least then you are more likely to be able to detect an error condition. I'd like to be able to say "always" detect an error condition, but, alas, we threw that away decades ago. "Hooray" for efficiency over correctness! [1]: Since this seems to come up whenever I fail to mention this, bear in mind that being able to name the exceptional circumstances does not make those circumstances any less exceptional. Grep for all arithmetic in your choice of program and the vast, vast majority are not deliberately using overflow behavior for some effect. Even in those programs that do use it for hashing or encryption or something you'll find the vast, vast bulk of arithmetic is not. The exceptions leap to mind precisely because, as exceptions, they are memorable.
- emiliobumachar 11y agoWe have not "decided that we're always operating in ring 256/65,536/etc". We decided that being a little more efficient can save a room in the computer's building. If correctness becomes harder, well, there's an acceptable price. The programmer is always in the building anyway. An obsolete decision, sure, but I wouldn't call it incorrect. This begs the question, though, why don't we revert it in new architectures, already incompatible with everything. Tradition, I guess?
- Peaker 11y agoOverflow checking all arithmetic is still going to be quite expensive, even with hardware support, because you have to choose between: A) exception/panic on overflow: this is preferable to overflow but also undesirable behavior. B) arbitrary precision integers: this is nicer, but even today significantly more expensive than ordinary integers.
- pandaman 11y agoFor what it's worth, all the bugs I've seen related to integer representation (not just over/under, e.g. I've seen code casting a 32bit pointer to a 64bit signed integer) could have been fixed by changing int to unsigned and never the opposite. Of course, this could be just the effect of the vast majority of programmers choosing int as the default integer type.
- jschwartzi 11y agoI've seen people 10 years my senior cast a 32-bit pointer to an int32_t. It's funny because it's not difficult to get right, but nobody ever seems to bother.
- pandaman 11y agoThis was in a game engine, written by a guy who is adored here, if I said his name, I'd probably been downvoted to negative :)
- cautious_int 11y agoAssume sizeof(int) == sizeof(unsigned int) == 4 g << h Well defined because 2147483648 can be represented in a 32 bit unsigned int Even under those assumptions, it is implementation defined if unsigned int can hold the value 2^31. It is perfectly valid to have an UINT_MAX value of 2^31-1. In that case the code will cause undefined behavior. The only guarantee for unsigned int is that it's UINT_MAX value is at least 2^16-1, regardless of its bit size, and that it has at least as much value bits as a signed int. For example C allows these ranges: int: -2^31 , 2^31-1 unsigned int: 0 , 2^31-1
- bsder 11y agoUse unsigned. All of the mathematical operations are defined on unsigned int in C. This is not true of int. Compilers are now starting to do all manner of nasty optimizations on undefined behavior. It is now only a matter of time before an signed int burns you.
- cautious_int 11y ago1u/0u
- jhallenworld 11y agoUse signed because pointer differences in C are signed: char *a, *b; size_t foo = a - b; /* generates a warning with -Wconversion */ This tells me that C's native integer type is ptrdiff_t, not size_t. (And I agree this is crazy: unsigned modular arithmetic would be fine, but they chose a signed result for pointer subtraction). Why care about this? You should try to get a clean compile with -Wconversion, but also you should avoid adding casts all over the place (they hide potential problems). It's cleaner to wrap all instances of size_t with ptrdiff_t- you can even check for signed overflow in these cases if you are worried about it. There is another reason: loop index code written to work properly for unsigned will work for signed, but the reverse is not true. This means you have to think about every loop if you intend to do some kind of global conversion to unsigned.
- SamReidHughes 11y ago> (And I agree this is crazy: unsigned modular arithmetic would be fine, but they chose a signed result for pointer subtraction). But. Signed means that i < j implies that p + i < p + j, or that p < q implies that p - r < q - r.
- pixelbeat 11y agoSome practical notes for handling/avoid signed integer overflow http://www.pixelbeat.org/programming/gcc/integer_overflow.html http://www.pixelbeat.org/programming/gcc/integer_overflow.ht...
- wuch 11y agoIn fact, prior to Java 8, you could not even declare unsigned integers. In Java 8 you still can't declare unsigned integers. It just that API have been extended to provide operations that treat an int or a long as having unsigned value. This seems to imply that Java leaves the behaviour of signed integer overflow up to the underlying hardware, but guarantees a two's complement representation even if the architecture does not. It seems reasonable to expect that it would simply wrap to 0 (I wasn't able to find a more conclusive reference for this). In fact Java Specification does specify what happens in case of overflow very precisely. For example for multiplication (section 15.17.1): If an integer multiplication overflows, then the result is the low-order bits of the mathematical product as represented in some sufficiently large two's-complement format. And for addition (section 15.18.2): If an integer addition overflows, then the result is the low-order bits of the mathematical sum as represented in some sufficiently large two's-complement format. Java approach without undefined behaviour would seem to be quite convenient for a programmer, but it does not seem to be the case in practice. Of course you gain ability to check for overflow after performing the operation, which is nice, but at the same time loose the straightforward way to detect bugs provided by undefined behaviour (UB). If an application executes UB, then it is obviously wrong, thus a simple instrumentation of arithmetic operation followed by a check for overflow gives a way to detect such problems without false positives (see for example -fsanitize=undefined and similar options available in modern C/C++ compilers). IMHO only very small fraction of overflow bugs would have been fixed by just wrapping result around.
- cautious_int 11y agoOf course you gain ability to check for overflow after performing the operation, Using exceptions? IMHO only very small fraction of overflow bugs would have been fixed by just wrapping result around. Yes, switching to wrapping integers is not the solution to overflow. Rather check the operation beforehand. I still think using unsigned integers is always preferred if you don't need negative values.
- wuch 11y agoWhat I had in mind, is the fact that with defined overflow you could unconditionally perform the operation, and then observe the result to decide if overflow have in fact happened. For example, following pattern to check if adding 100 to an integer 'a' overflows would be valid: int a = ...; if (a + 100 < a) overflow So this is one of cases where wrapping behaviour could potentially fix bugs in existing applications. Obviously, rewriting this to work with overflow is trivial if you are already aware of undefined behaviour that could occur.
- atilaneves 11y agoUse the most appropriate type. If you don't know what that is, use `int`. My general rule of thumb is "use unsigned integers when interfacing with hardware, otherwise use int.". Sometimes you need int64_t. You usually don't. The problem with unsigned integers is that most people, past me included, think that it prevents nonsense negative values. It doesn't. The compiler will quite gladly let you call a `void func(unsigned x)` with -1 and let you have "fun" later debugging why who and how this function got called with an enormous value. And, of course, that's not the only source of negative numbers that get passed in by mistake. More often it's the result of a subtraction that can "never" be negative.
- helmut_hed 11y agoBecause properly handling overflow if you're using signed values involves UB, it's super hard to get it right. Even Apple has struggled with it: http://www.slashslash.info/2014/02/undefined-behavior-and-apples-secure-coding-guide/ http://www.slashslash.info/2014/02/undefined-behavior-and-ap... If you get it wrong you can also see your program's behavior change with the compiler optimization level, which is incredibly frustrating (and causes a lot of "the compiler is buggy!" beliefs).