6 ms·
I think it's best to regard intmax_t as a failed experiment. E.g. GCC has long supported __int128, clang has _ExtInt, other compilers have similar extensions, b
by pdw 4y ago
I think it's best to regard intmax_t as a failed experiment. E.g. GCC has long supported __int128, clang has _ExtInt, other compilers have similar extensions, but they refuse to increase intmax_t due to ABI concerns.
- qalmakka 4y agoIf people use intmax_t and then they expect the ABI to stay stable, than they are at fault, not libc, LLVM or GCC for changing it when the actual "max int" changes.
- teddyh 4y agoI feel the fault instead lies in the operating system for wanting to run old binaries so much that the operating system pretends that the system architecture (and associated ABI) has not changed, even when there exists larger integers than what the previous architecture had set as intmax_t.
- afiori 4y agoThe problem is not old binaries, it is old code that assumes an upper bound to the size of intmax_t.
- teddyh 4y agoI’m not sure what you mean. I agree that the problem is not the binaries, but I believe the problem to be the operating system. The source code may not have assumed a specific upper size limit on intmax_t, but the binaries were compiled in good faith to a specific machine architecture, which had a specific width on intmax_t, namely the largest integer type which that architecture could support. If then new CPU models come out with larger integer types, but the operating systems still runs that same binary straight up (without any emulation layer), this exposes the old binary to larger integer types, breaking the binary. The binary is not at fault, it’s the operating system which committed the sin of running the binary on what is effectively a different architecture without any considerations on what architecture the binary was actually compiled for.
- afiori 4y agoThe OS could do this for old binaries but no for old source. A lot of FOSS code has baked in assumptions about the max size of intmax_t and the compiler and the OS cannot do much about broken logic.
- teddyh 4y ago> A lot of FOSS code has baked in assumptions about the max size of intmax_t Well, that code is not, and was not ever, standards compliant.
- shadowgovt 4y agoAnd it turns out that doesn't matter much in terms of adoption and popularity. After all, if predictable behavior in the future across all possible configurations were a high priority requirement for the tools people use to make computers do things, neither C nor C++ would have gotten off the ground. Instead, people don't memorize the entire standard. They write code that works on their machine, publish it, fix the bugs when people say it doesn't work on their machine, and maybe converge towards something that is standards compliant but, more likely, they converge towards something that works on the standard implementation on 90% of the operating systems available.
- kazinator 4y ago> A lot of FOSS code has baked in assumptions about the max size of intmax_t Would you be able to name three examples?
- afiori 4y agoI realize I used FOSS improperly, I had no reason to single it out; my comment reads like a critique of open source which was not my intention. My intention was to make a good case for old code being compiled today by third parties. But to my understanding intmax_t was introduced in C99, before widespread dominance of 64 bits platforms. So it would not be surprising that there exists code with lower assumptions about its size. But actually I made also another mistake, the major problem had little to do with faulty logic and more to do with dynamic linking [0] I will see myself out today. [0] https://thephd.dev/intmax_t-hell-c++-c https://thephd.dev/intmax_t-hell-c++-c
- pdw 4y agoBut all major C compilers promise that they won't make ABI changes. So I'd rather say that standardizing intmax_t was a mistake.
- kazinator 4y agoIf you don't standardize intmax_t, programs will invent the concept for themselves through some combination of preprocessing directives and shell scripts that probe the toolchain. If you want to know "what is the widest signed integer type available", you will get some kind of answer one way or another.
- mtlmtlmtlmtl 4y agoOr even worse, if they don't know their build tools, some kind of union based monstrosity...
- kevin_thibedeau 4y agoGCC changed long to 64-bits and x86 codebases had to accept that. The base integer types are designed to be variable sized and you code accordingly based on the minimum size guarantee. Assuming types are frozen in time is how you get things like LLP64.
- klodolph 4y agoLong is 32 bits on x86, 64 bits on x86_64. x86 code bases broke because they made non-portable assumptions.
- kevin_thibedeau 4y agoNo. It's 32-bits minimum. It can be anything larger. That's what the language says. Assuming otherwise is nonportable.
- deleted 4y ago[deleted]
- teddyh 4y agoOK, fine, but how, then, can the two problems I described be solved?
- deleted 4y ago[deleted]
- pdw 4y agoWell, what did you do before intmax_t was added to C99? - POSIX defines that in_port_t is equal to uint16_t, so for that type there's no problem at all ;) - Don't use printf/scanf for this, roll your own - Refuse to print/read uid_t that are outside intmax_t range - Use an autoconf test
- teddyh 4y agoMy program was first written when intmax_t already existed; AFAICT, intmax_t was introduced in C99. I also need to parse and/or print uid_d, gid_t, and pid_t, so I guess that in the future I’ll have to implement my own integer parser and printer, like user “orlp” suggested elsethread. Until then I guess I’ll have to do something like #if sizeof(pid_t) > sizeof(uintmax_t) #error Nope #endif …repeated for all the types I need to parse and/or print.
- ynik 4y agoYou can't use sizeof in #if, but you can use C11 static_assert instead. Also this doesn't just apply "in the future" -- integer types larger than intmax_t already exist, C23 just updates the standard to match reality.
- kazinator 4y agoYou can do this: #if SIZEOF_PID_T > SIZEOF_UINTMAX_T ... #endif Where you detect these sizes from the toolchain and deposit them as #define constants in some "config.h" header. There are ways to detect sizes by compiling a source file to an object file, and then analyzing the object file (no execution), so things work under cross-compiling. I've used a number of tricks over the years, and settled on this one: https://www.kylheku.com/cgit/txr/tree/configure?h=txr-278#n1473 https://www.kylheku.com/cgit/txr/tree/configure?h=txr-278#n1... The basic idea is that we can take a value like sizeof(long), which is a constant, and using nothing but constant arithmetic, we can convert it to the characters " 8". These characters can be placed into a character array where they are delimited by some prefix that we can look for in the compiled object file with some common utilities. This is quite durable in the sense that it can be reasonably expected to work through a wide variety of object formats. In that test program, I have a structure with some character arrays. The larger character arrays hold the informative prefixes. The two-byte arrays hold decimal digits for sizes. Two digits go up to 99 bytes so we are safe for a number of years. The DEC macro calculates the ASCII decimal value of its argument, expanding to a two-byte initializer for a character array: #define D(N, Z) ((N) ? (N) + '0' : Z) #define UD(S) D((S) / 10, ' ') #define LD(S) D((S) % 10, '0') #define DEC(S) { UD(S), LD(S) } E.g. DEC(42) -> /* the equivalent of */ { '4', '2' } But actually DEC(42) -> { UD(42), LD(42) } -> { D((42) / 10, ' '), D((42) % 10, '0') } -> { ((42) / 10) ? ((42 / 10)) + '0' : ' ', ((42) % 10) ? ((42 % 10)) + '0' : '0' } The space instead of a leading zero is so that if we pull this into shell arithmetic, it isn't confused for an octal number.
- kazinator 4y agointmax_t can rationally be regarded as being outside of the ABI concept. It should give you the maximum integer available, not the maximum integer that your compiler had 15 years ago for the sake of compatibility. Anyone who uses intmax_t in a durable API (public function argument, member of public structure) is just a goof.
- eropple 4y agoLotta goofs write a lotta code, though. I'm not sure that I disagree with you, but this feels like the sort of thing that could really hurt to break.
- teddyh 4y ago> It should give you the maximum integer available How would you compile that to machine code which would, when run, still give you the maximum integer on next year’s CPU with all-new super-duper wide integers? It is more reasonable to think of a binary to be compiled to a specific architecture triplet, and if the CPU changes, the architecture must change, and therefore the ABI, and if the OS wants to run a binary from an older architecture, some emulation layer is needed. Of course, this would be a lot of work for the operating system people if they want to run binaries from lots of what are now different architectures, and it was apparently easier to just force the C standard to abandon the entire concept of intmax_t being the largest.
- thayne 4y ago> if the CPU changes, the architecture must change, and therefore the ABI That would add a lot of friction to changing CPUs, and making distributing software more complicated.
- dianeb 4y agoIn C, you get what you pay for. When the platform changes, it's not unreasonable to recompile in order to take advantage of new hardware features. The OS is certainly different, even if it's still called "Whatever OS", it had to be recompiled in order to support the new hardware, unless, of course, it doesn't. We're all lazy when it comes to our own work, whether it's supporting new hardware in a compiler or OS or super web service on some new cloud platform or a large system vendor. intmax_t by intent is the maximum integer available on a given platform. Abandoning that idea is a disservice to compiler writers, programmers -- especially scientific programmers -- everywhere.