9 ms·
I've spent two decades writing C and C++, but the last 8-9 years in really high-level languages (Ruby, Javascript, Python). From either end of the spectrum, I'
by mieko 10y ago
I've spent two decades writing C and C++, but the last 8-9 years in really high-level languages (Ruby, Javascript, Python). From either end of the spectrum, I've never felt the need for such emphasis on fixed-sized numeric types.
I've commonly needed access to fixed size numerics, like when sending texture formats to the GPU, defining struct layout in file formats and network protocols, but I have never once thought: "You know what, I'd like to make a decision as to the width of an integer every time I declare a function."
"Just has to work" low level code was the norm during the 16-bit to 32-bit transition, so it was a fucking pain. Notice how smooth the 32-bit to 64-bit transition went? (and yes, it was smooth.) I credit that to high-level languages that don't care about this stuff, and people using better practices like generic word-sized ints and size_t's in lower level code. Keep that stuff on the borders of the application.
I've noticed a decent-sized emphasis on type size in both Crystal and Swift, two not-entirely braindead newer languages. I don't get it, it's a big step backward.
- lake99 10y agoLooks like you haven't worked on embedded systems. We need to be very careful with sizes of ints here. Not just because we run the risk of overflows, but also because when we create code that may have to be ported from one architecture to another, we want to minimize re-work. > Notice how smooth the 32-bit to 64-bit transition went? There are many reasons for that. For most PC work, a 32-bit int is more than large enough. Going to 64-bit should not have affected that at all. When you're working with embedded systems, you're often working at sizes that are the bare minimum that you can live with. You might also get to work on architectures where char, short, and int are all 32-bit wide. Assume something, and the communications protocol stops working. Moreover, the PC architecture itself supported coexistence of 32-bit and 64-bit executables.
- mieko 10y ago> Looks like you haven't worked on embedded systems. I've deleted a defensive technical response to point out that these dismissive assumptions (usually unjustified) are pretty prevalent on HN, and I don't think it promotes level-headed discussion. It reads like "Let me discredit a stranger's background that I don't know, and then argue my opposing view." It creates a defensive mindset off the bat. Edit: I mean, your points are valid. We won't agree on them, but I don't like the assumption that we won't agree because I'm ignorant to them.
- wott 10y agoYou shouldn't have worried: his points are not valid at all. Using fixed-width types across various embedded (as in "small and exotic") platforms is a nightmare for portability. There is no reason not to use standard integer types, providing you understood their definition. Use fixed-width types only when absolutely really needed; don't use standard types when you need fixed-width width, assuming they will be an exact size, and everything will be fine.
- TheNewAndy 10y agoI work with embedded systems daily, writing code that is expected to work on multiple different architectures (e.g. systems where char is 8-bit, 16-bit, 24-bit or 32-bit, systems without floating point units, systems with SIMD instructions, systems without, etc). In doing this, I have found that fixed width types make the job harder, and any code that uses fixed width types is generally more difficult to work with. You have an int16_t? How do you port it to a machine without a 16-bit type? What would happen if you replaced it with a 32-bit type? If you are relying on wraparound, then your code is already undefined, so some compiler is likely going to mess with you anyway. The stuff about communicating with the outside world is quite different. At this point you have protocols, and it depends on how a protocol is written, and how you are able to actually interact with the world to fulfil that protocol. If the protocol is an ABI, then generally you are fine, because the compiler too will match the ABI. If it is a file format, then you need to know how your file io functions work (do they just write the bottom 8-bits of a char, or do they blat the whole thing across?).
- lake99 10y ago> How do you port it to a machine without a 16-bit type? The problem is made more severe when you use "int" and let the compiler decide the size of the variable. I guess I don't know what you're advocating. > The stuff about communicating with the outside world is quite different. I meant communicating as a throwaway example. It's a valid example, but really, everything gets affected. For example, a CAN identifier is 29 bits. What generic "int" would you use when you want to fill in this identifier? And no, I work at a level that's lower than an ABI. So, no OS, no file IO, etc.
- adrianN 10y agoHow does it make portability worse if you write your code in a way that doesn't assume anything about the size of your ints? If you don't know that your ints are 16 or 32 or 27 bits long you have to be more careful with checking for overflows, but if the code does that I don't see what problems you'll have when you port it to a platform with a different int size.
- raverbashing 10y agoYou're not wrong, but it seems you're not abstracting enough (and depends on your embedded system) However the smaller they are the less complex the system is. And running higher level languages is not as taboo today
- tbirdz 10y agoC's normal integer types work the closer to the way you want. When you say you're returning an int, you're not really specifying the exact size at all. For those who don't know, C's char, short, int and long data types aren't really "fixed size". The standard defines their order: sizeof(char) <= sizeof(short) <= sizeof(int) <= sizeof(long), but the actual size of these integer types depends on the platform and compiler. If you want to use fixed size ints, you'd need to use the uint8_t, uint16_t, uint32_t types from stdint.h. Although interestingly enough C99 doesn't require the platform to provide these types, it only mandates the uint_least_n_t and uint_fast_n_t types (n=8,16,32..) which are guaranteed to be at least, but not necessarily exactly the requested width. Now you may be thinking "sure the standard says they could be different, but they aren't really, a char is 8bit, a short is 16-bit an int is 32-bit a long is 64-bit". So here's a couple examples for you: many DSP platforms have a C compiler where char is 16-bit. Also Microsoft Visual c++ on windows on x86_64 compiles longs as 32-bit while GCC on Linux on x86_64 compiles longs as 64-bits. >Notice how smooth the 32-bit to 64-bit transition went? (and yes, it was smooth.) I know you said it was smooth, and maybe it was for some applications. But for many others, the 32-bit to 64-bit transition actually caused a lot of problems! Andrey Karpov has already done a great job of categorizing many of them, so I won't waste my time repeating him, but you can read his list here: http://www.viva64.com/en/a/0065/ http://www.viva64.com/en/a/0065/
- mieko 10y agoYeah, I'm arguing that the situation you describe (accurately) is better than baked-in sizes all over the source code. Use the platform-native types (whatever size they may be) unless you have a reason not to, then use <stdint.h> or an analogue. If CHAR_BIT is 13, let char be 13 bits: the platform probably chose that for a reason. When you have to pack it into a TCP header, do your strict fixed-width stuff there.
- tbirdz 10y agoI completely agree with you when it comes to function return types. For data structure fields, especially structs used many, many times like in very large arrays, I'd say it's sometimes worth using fixed size types to get better control over memory use. Using a 64-bit int for a field a 16-bit integer can handle will use up 4x as much memory. And if you've got a ten or a hundred million structs of that type, then it really adds up. For example size_t is 64-bit on my x86_64 system. But modern x86_64 systems can only use 48-bit address spaces, so a 64-bit sized object can't even be addressed! Even worse is my cpu and motherboard have a 32gb maximum of RAM (for an effective 35-bit physical addressing limitation). And size_t is supposed to be able to store the size of any object in memory, but on this platform it stores things that won't fit in memory. So most of these 64-bits are wasted on modern systems. For files you can use off_t if you're on POSIX, but the C standard doesn't say anything about requiring size_t to be able to store any filesize in a filesystem. Just using size_t is not enough to make your code work correctly on 64-bit size quantities either. For example, you make your strlen implementation return size_t, and you use size_t everywhere you do anything with a string. But can your application really handle strings that are bigger than the system RAM , or the hardware address space? Are your algorithms even efficient enough to handle the 4,294,967,295 byte maximum string size for a 32-bit system? The effort to get your program to work efficiently on >32-bit quanities is often much harder than just using size_t instead of int. So to me, when I see 64-bit size_ts being used everywhere in code that won't actually be able to handle working with >32-bit quantities, it just feels a little useless. Of course this is really more a complaint about how big size_t is on the x86_64 platform than it is a complaint about the idea of size_t in general. If only we had 48-bit size_ts (24-bit would be handy too!)
- keeperofdakeys 10y agoSo you really have two arguments in there. For pointers, I'm sure no one will disagree that size_t is a good thing. Which is why any modern low-level language has these built-in. As for fixed size types, what are the alternatives? If you wanted a dynamically sized int, you'd either need to reserve all the possible space you'd need (in which case why not just use a int64_t), or you'd need some kind of heap-allocated integer that can resize itself. We can't hide the fact that the machine itself uses fixed-size ints, unless we are willing to live with a leaky abstraction.
- ploxiln 10y agoOverflow is often something kept in mind in careful C programming. It's something you don't have to worry about nearly as much in Python (promotes to bignum) or Javascript (double precision float). That 32 to 64 bit transition? Windows couldn't change the size of "long" because it would break too much windows software that assumed long was 4 bytes. Thus "long long". The Linux ecosystem fared a bit better because of the greater variety of Unix systems and cross-platform C that ran on them. It was a time of fragmentation and incompatibility, but also a time of posix and attempting to define minimal standards in a portable and flexible way. Windows of course didn't have to care about any of that. I actually fixed some old-ish classic open source networking software for 64-bit mips, around 2010. You see, it worked OK on x86_64, which is little endian, because in that case the low-order 4 bytes (an ipv4 address) end up in the same place relative to the start address, whether it's a 4-byte or 8-byte type. But for big-endian 64-bit mips, that doesn't work out, the low-order 4 bytes are on the unlucky end of the 8-byte type. There's also the awkwardly named "ntohl()" and "htonl()" standard library functions that were named in the distant past when it was assumed a long would always be 4 bytes. As soon as that wasn't true they changed to work with uint32_t, as they always should have. Sure, I'll use a plain int whenever I'm returning a status or error code, and whenever iterating over some trivial range. I'll use a "char[]" or similar for a string of ascii. But for many things, like "count of bytes transferred", or "seconds until timeout", or any meaningful quantity value or anything stored in a struct (meaningful and non-transient), I'll pick the appropriate bit-size, to make the code clear. It's good style because you have to know the limits of the type whenever you do anything non-trivial with it. These days, there's always a value here or there where 4 billion will sometimes not be enough: bytes, microseconds. I need 64 bits. I could write "long long", but why not write what I really mean: "uint64_t" aka give me 64 bits. (also, "unsigned" is longer to write and read :)
- user5994461 10y ago> Notice how smooth the 32-bit to 64-bit transition went? (and yes, it was smooth.) I credit most of that to Microsoft and the insane efforts they make to support backward compatibility. They made an extremely complex abstraction layer to run 32 bits applications on 64 bits OS, yet it works very well. To this day, it's possible to run lot of applications that were written 15 years ago, seamlessly, as they were built and distributed 15 years ago.
- mieko 10y agoI've heard Microsoft had a great story there, but I was drawing from experiences with the open source Unix-likes. Seems like it was across the board until you get too far into the weeds.
- cyber_dude 10y agoYou should credit AMD for introducing AMD64/x86_64. Without that backward compatibility wouldn't have been possible.
- trentnelson 10y agoA great paper if you haven't already seen it: https://github.com/tpn/pdfs/blob/master/A%20History%20of%20Modern%2064-bit%20Computing%20-%20Feb%202007%20(CSEP590A).pdf https://github.com/tpn/pdfs/blob/master/A%20History%20of%20M...
- wott 10y agoWhen browsing code posted or advertised on various forums, I noticed a dramatic increase in the use of uint8_t and such. In places where they really were not useful, and a basic standard type would suffice, and be more portable, and be as fast or faster. I have seen intN_t used as booleans! I have seen uintN_t types used and signed values put in them! (which meant the author didn't understand yet the basis of types but was somehow taught it was a good practice to use exact-width types). I also suspect the influence of languages with fanboys, as Rust. I had a 'fight' with these, they really didn't grasp the concept of having a type defined as the natural type of the platform and as having at least (and possibly at most) N bits. I gave up. As long as they have fun with their thing and do not try to spread "the good word" and influence other languages, that would be fine. Let's hope we won't see the influence of languages whose designers had a gripe against unsigned integers...
- karyon 10y agoFor a university project I programmed for an Arduino Uno (ATmega328P). Debugging performance problems, we found out that a switch (implemented with a bunch of if/else ifs) with about 12 branches used ca. 200 clock cycles just checking the 12 conditions. Turns out the ATmega328P has only 8bit registers and the ALU operates on 8bit values. We were using 32-bit datatypes in the conditions, which the ATmega328P loaded and compared byte per byte, each comparison (incl. jump etc) cost us something like 16 cycles. So yeah, the size of the datatypes definitely mattered for us :) edit: as an addition, we needed quite a while to figure out (and were quite surprised when we did) that ints are 16bit and longs 32bit on this platform/compiler. i guess this comes down to the general question of explicit vs implicit, and i usually do prefer the former.
- Cyph0n 10y agoWhenever I write ATmega32P code, I add the following to a header file: >> typedef unsigned char uint8; And then I use uint8 by default (loop counter, accumulator, constant, etc.), unless a larger size number is required.
- dahart 10y agoI was working in console games during the 32->64 transition, and watched our game engine code suddenly go crazy with data size specificity. It was a little weird. I don't know whether it's good or bad overall, but there was definitely a reason for it, the need was clear. Suddenly you didn't know how big an int was, and that matters a lot when you're sending ints over the network and stuffing ints into save games, and it matters a lot when you use an int expecting 64 bits, but you're still compiling code for a Wii and only get 32. It matters for multiply overflows and negative numbers and bit flags and for a bunch of things besides deciding how high your for loop will go or what your largest return value is. As you say, very few people want to decide what size to use when declaring functions. But everyone wants to know exactly what to expect. If you do declare something and don't actually know what size it is, you have a real problem that can and will lead to crashes. > Notice how smooth the 32-bit to 64-bit transition went? I'm not certain, but is it possible it went smooth because everyone working in C/C++ started paying attention to their data sizes? As a user of Ruby/JS/Python, it would be worth checking what happened to the source code for the interpreters of those languages. When you program in them, yes, you're buffered from a lot of data size issues, but the internals of the languages themselves may be just as fixed-size centric as anything.
- mieko 10y ago> I'm not certain, but is it possible it went smooth because everyone working in C/C++ started paying attention to their data sizes? We don't have to guess. A huge amount of the code that made this transition was strewn across the internet in CVS and SVN repos, and now in preserved git history. From compilers and libcs to kernels to linkers, loaders and interpreters. They make concessions where they had to (check out glibc), but stay with traditional K&R-style types when possible. The "i32 calculate_age()" thing is way newer than that.
- guscost 10y agoI'm not very experienced with low-level programming, but doesn't the single-responsibility principle suggest that a purpose-built module should handle packing/unpacking of memory objects even if you are writing this module by hand? I'm sure there are cases where even the overhead of calling and running these procedures is too much, but in many cases no further optimization would be needed.
- api 10y agoThere's one more gotcha: security concerns. Writing code that expects long to be 64 bit might be dangerous as it could overflow and create a security bug on a 32 bit machine. I find that I do want to know the sizes of things fairly often.