10 ms·
C Portability Lessons from Weird Machines
- enjoyyourlife 5y agoWhy don't C programmers always use stdint.h?
- PhantomGremlin 5y agothe MIPS R3000 processor ... raises an exception for signed integer overflow, unlike many other processors which silently wrap to negative values. Too bad programmer laziness won and most current hardware doesn't support this. As a teenager I remember getting hit by this all the time in assembly language programming for the IBM S/360. (FORTRAN turned it off). S0C8 Fixed-point overflow exception When you're a kid you just do things quickly. This was the machine's way of slapping you upside your head and saying: "are you sure about that?"
- laumars 5y ago> When you're a kid you just do things quickly. I don’t think this is a age problem. Plenty of adults are lazy and plenty of kids aren’t.
- flohofwoe 5y agoModulo wraparound is just as much a feature in some situations as it is a bug in others. And signed vs unsigned are just different views on the same bag of bits (assuming two's complement numbers), most operations on two's complement numbers are 'sign agnostic' - I guess from a hardware designer's pov, that's the whole point :) The question is rather: was it really a good idea to bake 'signedness' into the type system? ;)
- zozbot234 5y agoModulo wraparound is convenient in non-trivial expressions involving addition, subtraction and multiplication because it will always give a correct in-range result if one exists. "Checking for overflow" in such cases is necessarily more complex than a simple check per operation; it must be designed case by case.
- pornel 5y agoThat's why Rust has separate operations for wrapping and non-wrapping arithmetic. When wrapping matters (e.g. you're writing a hash function), you make it explicit you want wrapping. Otherwise arithmetic can check for overflow (and does by default in debug builds).
- masklinn 5y ago> Modulo wraparound is just as much a feature in some situations as it is a bug in others. That’s an extremely disingenuous line of reasoning, the situations where it’s a feature is a microscopic fraction of total code: most code is neither interested in nor built to handle modular arithmetics, and most of the code which is interested in modular arithmetics needs custom modulos (e.g. hashmaps), in which case register-size-modulo is useless. That leaves a few cryptographic routines built specifically to leverage hardware modular arithmetics, which could trivially be opted in, because the developers of those specific routines know very well what they want. > signed vs unsigned are just different views on the same bag of bits […] The question is rather: was it really a good idea to bake 'signedness' into the type system? ;) The entire point of a type system is to interpret bytes in different ways, so that’s no different from asking whether it’s really a good idea to have a type system. As to your final question, Java removed signedness from the type system (by making everything signed). It’s a pain in the ass.
- flohofwoe 5y ago> Java removed signedness from the type system (by making everything signed) That's not removing signedness. Removing signedness would be treating integers as sign-less "bags of bits", and just map a signed or unsigned 'view' over those bits when actually needed (for instance when converting to a human-readable string). Essentially Schroedinger's Cat integers. There'd need to be a handful 'signed operations' (for instance widening with sign extension vs filling with zero bits, or arithmetic vs logic right-shift), but most operations would be "sign-agnostic". It would be a more explicit way of doing integer math, closer to assembly, but it would also be a lot less confusing.
- zozbot234 5y agoOverflow checks are trivial, there's no need for special hardware support. It's pretty much exclusively a language-level concern.
- addaon 5y agoOverflow checks can be very expensive without hardware support. Even on platforms with lightweight support (e.g. x86 'INTO'), you're replacing one of the fastest instructions out there -- think of how many execution units can handle a basic add -- with a sequence of two dependent instructions.
- zozbot234 5y agoA vast majority of the cost is missed optimization due to having to compute partial states in connection to overflow errors. The checks themselves are trivially predicted, and that's when the compiler can't optimize them out.
- monocasa 5y agoIn practice the vast majority of MIPS code uses addu, the non trapping variant. And in x86 land there's the into instruction, interrupt if overflow bit set, so you're left with the same options.
- spc476 5y agoWhich has to be done after every instruction (http://boston.conman.org/2015/09/05.2 http://boston.conman.org/2015/09/05.2) but it quite slow. Using a conditional jump after each instruction is faster than using INTO (http://boston.conman.org/2015/09/07.1 http://boston.conman.org/2015/09/07.1).
- monocasa 5y agoIt's more complicated than shows up in micro benchmarks like that. Since when you do it, it's pretty much every add, you end up polluting your branch predictor by using jo instructions everywhere and it can lead to worse overall perf.
- colejohnson66 5y agoMy guess would be a pipelining issue where `INTO` isn't treated as a `Jcc`, but as an `INT` (mainly because it is an interrupt). Agner Fog's instruction tables[0] show (for the Pentium 4) `Jcc` takes one uOP with a throughput of 2-4. `INTO`, OTOH, when not taken uses four uOPs with a throughput of 18! Zen 3 is much better with a throughput of 2, but that's still worse than `JO raiseINTO`. [0]: https://www.agner.org/optimize/instruction_tables.pdf https://www.agner.org/optimize/instruction_tables.pdf
- masklinn 5y ago> Too bad programmer laziness won and most current hardware doesn't support this. There were discussions around this a few years back when Regher brought up the subject, and one of the issues commonly brought up is if you want to handle (or force handling of) overflow, traps are pretty shit, because it means you have to update the trap handler before each instruction which can overflow, because a global interrupt handler won't help you as it will just be a slower overflow flag (at which point you might as well just use an overflow flag). Traps are fine if you can set up a single trap handler then go through the entire program, but that's not how high-level languages deal with these issues. 32b x86 had INTO, and compilers didn't bother using it.
- Findecanor 5y agoModern programming language exception handler implementations use tables with entries describing the code at each call-site instead of having costly setjmp()/longjmp() calls. I think you could do something similar with trap-sites, but the tables would probably be larger. BTW. The Mill architecture handles exceptions from integer code like floating point NaN: setting a meta-data flag called NaR - Not a Result. It gets carried through calculations just like NaNs do: setting every result to NaR if used as an operand... Up until it gets used as operand to an actual trapping instruction, such as a store. And of course you could also test for NaR instead of trapping.
- rsecora 5y agoIt still amazes me how the PDP-11 has the NUXI [1] problem at nibble level and how the PDP-11 was bytesexual [2]. [1] http://catb.org/jargon/html/N/NUXI-problem.html http://catb.org/jargon/html/N/NUXI-problem.html [2] http://catb.org/jargon/html/B/bytesexual.html http://catb.org/jargon/html/B/bytesexual.html
- gwern 5y agoNote: "weird machines" here has nothing to do with the well-known security concept, just referring to unusual or obscure computers.
- kwertyoowiyop 5y agoGiven C’s origin on the PDP-11, it’s amazing it ended up so portable to all these crazy architectures. Even as an old-timer, the 8051 section made me say “WTF”!
- AnimalMuppet 5y agoOverall a good article. I was a bit amused and/or disgruntled to see a TRS80 in the "Motorola 68000" section, though...
- rjsw 5y agoWhy disgruntled? I never saw a Model 16 [1] but they did exist. [1] https://en.wikipedia.org/wiki/TRS-80_Model_II#model16 https://en.wikipedia.org/wiki/TRS-80_Model_II#model16
- AnimalMuppet 5y agoWell, OK, but the vast majority were the Z80-only versions. I can't read the model number on the frame of video it statically gives me, and I'm not going to play the video to find out...
- astrobe_ 5y ago> Everyone who writes about programming the Intel 286 says what a pain its segmented memory architecture was Actually this concerns more pre-80286 processors, since 80286 introduced virtual memory, and the segment registers were less prominent in "protected mode". Moreover I wouldn't say it was a pain, at least at the assembly level, once you understood the trick. C had not concept of segmented memory, so you had to tell the compiler which "memory model" it should use. > One significant quirk is that the machine is very sensitive to data alignment. I remembered from school time about a "barrel register" that allowed to remove this limitation, but it was introduced in 68020. On the topic itself, I like to say that a program is portable if it has been ported once (likewise a module is reusable if it has been reused once). I remember porting a program from a 68K descendant to ARM, the only non-obvious portability issue was that in C, the char type is that the standard doesn't mandate the char type to be signed or unsigned (it's implementation-defined).
- spc476 5y agoThe segment registers were less prominent on the 80386 in protected mode since you also have paging, and each segment can be 4G in size. On the 80286 in protected mode the segment registers are still very much there (no paging, each segment is still limited to 64k).
- bombcar 5y agoSegments are awesome if they're much larger than available RAM - handling pointers as "base + offset" is often much easier to understand than just raw pointers.
- zwieback 5y ago> > Everyone who writes about programming the Intel 286 says what a pain its segmented memory architecture was > Actually this concerns more pre-80286 processors, since 80286 introduced virtual memory, 86 had segments, 286 added protected mode, 386 added virtual. I would agree, though, 286 wasn't as bad as people make it sound. In OS/2 1.x it was quite usable.
- shadowofneptune 5y ago
- eqvinox 5y agoOn a slightly related note, chances are good anyone reading this has an 8051 within a few meters of them - they're close to omnipresent in USB chips, particularly hubs, storage bridges and keyboards / mice. The architecture is equally atrocious as the 6502. btw: a good indicator is GCC support - AVR, also an 8-bit µC - is perfectly well supported by GCC. 8051 and 6502, you need something like SDCC [http://sdcc.sourceforge.net/ http://sdcc.sourceforge.net/]
- RicoElectrico 5y agoHope RISC-V will displace 8051 over time. It's such an absurd thing to extend this architecture in myriad non-interoperable (although backwards-compatible with OG 8051) ways. And don't forget about the XRAM nonsense. Yuck.
- jazzyjackson 5y ago> The architecture is equally atrocious as the 6502. I only ever hear glowing/nostalgic reviews of 6502 programming, I guess from the retro/8bit gaming scene, curious what you find so atrocious.
- tenebrisalietum 5y ago6502 is awesome to program from assembly. What makes the 6502 atrocious for C is: - internal CPU registers are 8 bits, no more, no less and you only have 3 of them (5 if you count the stack pointer and processor status register). - fixed 8-bit stack pointer so things like automatic variables and pass-by-value can't take up a lot of space. - things like "access memory location Z + contents of register A + this offset" aren't possible without a lot of instructions. - no hardware divide or multiply. Many CPUs have instructions that map neatly to C operations, but not 6502. With enough instructions C is hostable by any CPU (e.g. ZPU) but a lot of work is needed to do that on the 6502 and the real question is - will it fit in 16K, 32K, as most 6502 CPUs only have 16 address lines - meaning they only see 64K of addresses at once. Mechanisms exist to extend that but they are platform specific. IMHO Z80 is better in this regard with it's 16-bit stack pointer and combinable index registers.
- zwieback 5y agoI wrote a fair amount of code for TI's TMS320C4x DSPs. They had 32 bit sized char, short, int, long, float and double and a long double with 40 bits. Took a bit to get used to but really the only way to get to the good stuff was by writing assembly code and hand-tuning all the pipeline stuff.
- rwmj 5y agoC23 just dropped support for any non-twos-complement architectures. No more C on Unisys for you! http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2412.pdf http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2412.pdf
- eqvinox 5y agoThat doesn't preclude C23 on Unisys, it just forces the compiler to hide that fact from the programmer ;D (SCNR)
- bombcar 5y agoWhich will be wonderful for code compiled to C23 linked against C89 libraries, if that's even possible.
- titzer 5y agoIt's amazing the abilities that 50 years can bring a programming language. Longest, most painful design debate ever.
- namibj 5y agoOne of the most widespread potential targets for non-two's-complement C would be JavaScript, or otherwise abusing floating point hardware for C's signed integers. Because their overflow is indeed weird (and breaks associativity).
- mmoskal 5y agoIt looks like signed overflow will be still undefined in C23 (but defined in C++).
- isomel 5y agoSigned overflow is not defined in C++
- rjsw 5y agoThere are C compilers for the PDP-10, it must count as fairly weird.
- nivertech 5y agoThe author forgot to mention that 8051 has a bit-addressable lower part of RAM. PDP-11 had a weird RAM overlay scheme of squeezing 256KB RAM into a 64KB 16-bit address space. IBM System/360 also had a weird addressing scheme with base register and up to 4KB offsets. https://en.wikipedia.org/wiki/IBM_System/360_architecture#Addressing https://en.wikipedia.org/wiki/IBM_System/360_architecture#Ad...
- zokier 5y agoI don't think there is anything wrong in writing platform-specific code; in certain circles there is this weird fetishitization of portability, placing it on the highest pedestal as a metric of quality. This happens in C programming and also for example in shell scripting, people advocating for relying only on POSIX-defined behavior. If a platform specific way of doing something works better in some use-case then there should be no shame in using that. What is important is that the code relies on well-defined behavior, and also that the platform assumptions/requirements are documented to a degree. Of course it is wonderful that you can make programs that are indeed portable between this huge range of computers; just that not every program needs to do so.
- bombcar 5y agoC (and to some extent shell) programmers are the ones with the most experience of the machine under them changing, perhaps drastically - few other languages have even been around long enough for that to have happened. Java sidesteps this, of course, by defining a JVM to run on and leaving the underlying details to the implementation.
- Filligree 5y ago> Of course it is wonderful that you can make programs that are indeed portable between this huge range of computers; just that not every program needs to do so. Isn't most code that would behave differently on different architectures subject to undefined behaviour, however? The signed overflow case mentioned, for example. Sure, some of it is implementation-defined, but in practice you need to write ultra-portable code anyway in order for your compiler not to pull the rug out underneath you.
- scaramanga 5y agoThe behaviour of signed integer overflow is undefined, but only if it actually overflows, and whether it overflows or not can be based on the widths of types which are implementation-defined, so there is a link between the two things. Many "correct" programs rely on undefined behaviour for optimizations while the authors knowingly assert (or assume) that the undefined behaviour doesn't actually occur. So the question is, if you are writing non-portable code for a specific environment, is a UB that actually triggers on your environment "non-portability" or is it a "bug"? I think most people would define that as being a bug. Edit: Okay, bad example cos you could maybe use int_leastN_t or something in that example to document the assumptions... But you don't always have control over types, and there may be multiple constraints, etc..
- ChuckMcM 5y agoI scored 7. (have written C code on six of the architectures mentioned (PDP 11, i86, VAX, 68K, IBM 360, AT&T 3B2, and DG Eclipse) I have also written C code on the DEC KL-10 (36 bit machine) which isn't covered. And while I have a PDP-8, I only have FOCAL and FORTRAN for it rather than C. I'm sure there is a C compiler out there somewhere :-). With the imminent C23 spec I'm really amazed at how well C has held up over the last half century. A lot of things in computers are very 'temporal' (in that there are a lot of things that are all available at a certain point in time that are required for the system to work) but C has managed to dodge much of that.
- viddi 5y agoHaven't read the article yet, but I have noticed that the tab keeps loading even after 10 minutes. Aborting the loading process leads to broken media. I am no expert in HTML video delivery and haven't tried it out, but maybe setting the preload attribute to "none" or "metadata" might help?
- WalterBright 5y agoI recall discussing C with a person using a processor where chars, shorts, and ints were all 32 bits. He stressed that writing portable code was necessary. I pointed out that any programs that needed to manipulate byte data would require very special coding to work. Special enough to make it pointless to try to write portable code between it and a regular machine. It's unreasonable to contort the standard to support such machines. It is reasonable for the compiler author on such machines to make some adjustments to the language. For example, one can say C++ is technically portable to 16 bit machines. But it is not in practice, because: 1. segmented 16 bit machines require near/far pointer extensions 2. exceptions will not work on 16 bit machines, because supporting it consumes way too much memory 3. ditto for RTTI
- jnwatson 5y agoI love these weird machines. I'll give an example of another. The Texas Instruments' C40 was a DSP made in the late 90's. It had 32-bit words, but inefficient byte manipulation. The compiler/ABI writer's solution was simple: make char 32 bits. So sizeof(char)==sizeof(short)==sizeof(int)==sizeof(long). I remember writing routines for "packing" and "unpacking" native strings to byte strings and back.
- adonovan 5y agoI remember using that machine and wondering why my program used 4x the memory I expected. The hardware manual used the term "byte" in the old-fashioned way, to mean "minimum addressable unit": 32 bits. Today "byte" only ever means "octet".
- anonymousiam 5y agoGood article. I've done C on half of those platforms. 8051 is actually more complex than described. There are two different kinds of RAM with two different addressing modes. IIRC there are 128 bytes of "zero page" RAM and, depending on the specific variant, somewhere between 256 bytes and a few kilobytes of "normal" RAM. Both RAM types can be present, and the addresses can both be the same value, but point to different memory, so the context of the RAM type is critical. The variants usually have a lot more ROM than RAM so coding styles may need to be adjusted, such as using a lot of (ROM initialized data) constants instead of computing things at run time. 6502 has a similar "zero page" addressing mode to the 8051. I never encountered any alignment exceptions on 68k (Aztec C). Either I was oblivious and lucky, or just naturally wrote good code. I do remember something about PDP-11 where there was a maximum code segment size (32k words?). C on the VAX (where I first learned it) was a superset of the (not yet) ANSI standard. I vaguely remember some cases where the compiler/environment would allow some lazy notation with regard to initialized data structures. They left out some interesting platforms (such as 6809/OS9, TI MSP430 and PPC), which have their own quirks.
- kazinator 5y agoTheoretical portability is not very useful. If you've not tested on a Unisys with 36 bit integers and 8 word function pointers, the theoretical port is garbage. It may be easier to fix the code than if you didn't think about the Unisys, but that effort has to be weighed against the incredibly vanishing odds of it ever being required.
- ziml77 5y agoWhile skimming did I miss an example of code that is portable between most of these systems? I'd love to see that because I'm having a very hard time believing that's possible. Or maybe you can make something that will technically compile and function on any of those, but if it doesn't perform reasonably then it's really hard to call the portability aspect a success. Also you're very limited as to what you can actually do in ANSI C. You're going to need to start poking directly at the hardware which is not going to be portable. Hell, even stuff like checking if a letter is within a certain range in the English alphabet might not work between machines. Letters might not be contiguous in the machine's native character encoding.
- ForOldHack 5y agoNo mention of the Interdata 8/32 or the (cough cough) IBM RT?
- jecel 5y agoIndeed, the effort to port C and Unix to the Interdata forced C to evolve quite a bit (adding unions, typedefs and so on): https://www.bell-labs.com/usr/dmr/www/portpap.html https://www.bell-labs.com/usr/dmr/www/portpap.html
- gnufx 5y agoI thought there'd be mention of CHERI as an up-to-the-minute architecture (as Arm Morello). I don't remember whether it requires modifications to standard C, but there's a C programming guide [1]. People who say everything is little endian presumably don't maintain packages for the major GNU/Linux systems which support s390x. I don't remember how many big endian targets I could count in GCC's list when I looked fairly recently. The place to look for "weird" features there is presumably the embedded systems processors. 1. https://www.cl.cam.ac.uk/techreports/UCAM-CL-TR-947.pdf https://www.cl.cam.ac.uk/techreports/UCAM-CL-TR-947.pdf