5 ms·
This was three or four jobs ago, but I remember reviewing someone's C code and they kept different collections of char* and int* pointers where they could have
by billpg 4y ago
This was three or four jobs ago, but I remember reviewing someone's C code and they kept different collections of char* and int* pointers where they could have used a single collection of void* and the handler code would have been a lot simpler.
The justification was that on this particular platform, char* pointers were differently structured to int* pointers, because char* pointers had to reference a single byte and int* pointers didn't.
EDIT - I appear to have cut this story short. See my response to "wyldfire" for the rest. Sorry for causing confusion.
- wyldfire 4y ago> because char* pointers had to reference a single byte and int* pointers didn't. I must be missing some context or you have a typo. Probably most architectures I've ever worked with had `int *` refer to a register/word-sized value, and I've not yet worked with an architecture that had single-byte registers. Decades ago I worked on a codebase that used void * everywhere and rampant casting of pointer types to and fro. It was a total nightmare - the compiler was completely out of the loop and runtime was the only place to find your bugs.
- billpg 4y agoI forget the details (long time ago) but char* and int* pointers had a different internal structure. The assembly generated by the compiler when code accessed a char* pointer was optimized for accessing single bytes and was very different to the code generated for an int* pointer. Digging deeper, this particular microcontroller was tuned for accessing 32 bits at a time. Accessing individual bytes needed extra bit-shuffling code to be added by the compiler.
- leeter 4y agoSounds like M68K or something similar, although Alpha AXP had similar byte level access issues. A compiler on either of those platforms likely would add a lot of fix up code to deal with the fact they have to load the aligned (either 16bit in M68K case or 32Bit IIRC in Alpha) and then do bitwise and shifts depending on the pointers lower bits. Raymond's blog on the Alpha https://devblogs.microsoft.com/oldnewthing/20170816-00/?p=96825 https://devblogs.microsoft.com/oldnewthing/20170816-00/?p=96...
- monocasa 4y agoM68k was byte addressable just fine. Early alpha had that issue though, as did later cray compilers. Alpha fixed it with BWX (byte word extension). Early cray compilers simply defined char as being 64bits, but later added support for the shift/mask/thick pointer scheme to pack 8 chars in a word.
- leeter 4y agoMust have depended on variant, the one we used in college would throw a GP fault for misaligned access. It literally didn't have an A0 line. That said it's been over 10 years and I could be remembering the very hard instruction alignment rules as applying to data too...
- monocasa 4y ago16bits had to be aligned. It didn't have an A0 because of the 16bit pathway, but it did have byte select lines (#UDS, #LDS) for when you'd move.b d0,ADDR so that devices external to the CPU could see an 8-bit data access if that's what you were doing.
- wyldfire 4y ago> char* and int* pointers had a different internal structure. The assembly generated by the compiler when code accessed a char* pointer was optimized for accessing single bytes and was very different to the code generated for an int* pointer. But -- they are different. Architectures where they're treated the same are probably the exception. Depending on what you mean by "very different" - most architectures will emit different code for byte access versus word access.
- billpg 4y agoAccessing a 32 bit word was a simple read op. Accessing an 8 bit byte from a pointer, the compiler would insert assembly code into the generated object code. The "normal" part of the pointer would be read, loading four characters into a 32 bit register. Two extra bits were squirreled away somewhere in the pointer and these would feed into a shift instruction so the requested byte would appear in the lowest-significant 8 bits of the register. Finally, an AND instruction would clear the top 24 bits.
- mlyle 4y agoThere are architectures where all you have is word addressing to memory. If you want to get a specific byte out, you need to retrieve it and shift/mask yourself. In turn, a pointer to a byte is a software construct rather than something there's actual direct architectural support for.
- loeg 4y agoDo C compilers for those platforms transparently implement this for your char pointers as GP suggests? I would expect that you would need to do it manually and that native C pointers would only address the same words as the machine itself.
- mlyle 4y ago> Do C compilers for those platforms transparently implement this for your char pointers as GP suggests? Yes. Lots of little microcontrollers and older big machines have this "feature" and C compilers fix it for you. There are nightmarish microcontrollers with Harvard architectures and standards-compliant C compilers that fix this up all behind the scenes for you. E.g. the 8051 is ubiquitous, and it has a Harvard architecture: there are separate buses/instructions to access program memory and normal data memory. The program memory is only word addressable, and the data memory is byte addressable. So, a "pointer" in many C environments for 8051 says what bus the data is on and stashes in other bits what the byte address is, if applicable. And dereferencing the pointer involves a whole lot of conditional operations. Then there's things like the PDP-10, where there's hardware support for doing fancy things with byte pointers, but the pointers still have a different format than word pointers (e.g. they stash the byte offset in the high bits, not the low bits). The C standards makes relatively few demands upon pointers so that you can do interesting things if necessary for an architecture.
- billpg 4y agoDepends on how helpful the compiler is. This particular compiler had an option to switch off adding in bit shifting code when reading characters and instead set CHAR_BIT to 32, meaning strings would have each character taking up 32 bits of space. (So many zero bits, but already handles emojis.)
- kjs3 4y agoan architecture that had single-byte registers Wild guess, but the OP might be talking about the Intel 8051. Single-byte registers, and depending on the C compiler (and there are a few of them) 8-bit int* pointing to the first 128/256 bytes of memory, but up to 64K of (much slower) memory is supported in different memory spaces with different instructions and a 16-bit register called DPTR (and some implementations have 2 DPTR registers). C support for these additional spaces is mostly via compiler extensions analogous but different from the old 8086 NEAR and FAR pointers. I'm obviously greatly simplifying and leaving out a ton of details. Oh, yeah...on 8051 you need to support bit addressing as well, at least for the 16 bytes from 20h to 2Fh. It's an odd chip.
- dahart 4y agoIt is true that on at least some platforms an int* that is 4 byte aligned is faster to access than a pointer that is not aligned. I don’t know if there are platforms where int* is assumed to be 4-byte aligned, or if the C standard allows or disallows that, but it seems plausible that some compiler somewhere defaulted to assuming an int* is aligned. Some compilers might generate 2 load instructions for an unaligned load, which incurs extra latency even if your data is already in the cache line. These days usually you might use some kind of alignment directive to enforce these things, which works on any pointer type, but it does seem possible that the person’s code you reviewed wasn’t incorrect to assume there’s a difference between those pointer types, even if there was a better option.
- hoseja 4y agoIn C++, char* pointers are special in that they are exempt from aliasing rules.