6 ms·
Would you get simpler (faster) ops if you stole the high bit instead of the low bit? e.g. x+y becomes (x + y | (1<<63)) Or does the loss of overflow detection
by jbert 12y ago
Would you get simpler (faster) ops if you stole the high bit instead of the low bit?
e.g. x+y becomes (x + y | (1<<63))
Or does the loss of overflow detection hurt you more?
- breadbox 12y agoSetting the high bit would mean that the integer values are no longer automatically distinguishable from pointer values, which is kind of the whole point.
- TheLoneWolfling 12y agoWhy? You could just check if the high bit is set, instead of if the low bit is set. IIRC, you can do this just by checking if it is less than zero.
- infogulch 12y agoFirst, pointers are typically unsigned integers. Second, memory addresses are virtual, they don't have to exist physically. That is, you don't have to physically have up through 2 quintillion addressable bytes for a pointer with that value to be valid. This is especially true with Address Space Layout Randomization (ASLR) security techniques employed by operating systems. tl;dr: all valid pointer values are possible, regardless of how much memory you actually have.
- TheLoneWolfling 12y agoSo then store the pointer implicitly right-shifted by 1. Means that boxed access is slower, but unboxed access is faster. And, when you get down to the assembly level, it doesn't matter - you can treat an unsigned integer as signed for a comparison if it makes things easier.
- Someone 12y agoThere could be addressable memory with the high bit set, for example ROM space for booting or video memory. That may be rare nowadays (or is it? I remember reading recently that some 64-bit CPU chose to preferably use a memory range from -2GB to +2GB. AMD64?), but whether there is is outside of the control of most language implementers. They can control their own memory allocator, though, and guarantee that it never stores 64-bit quantities at odd addresses.
- breadbox 12y agoThe issue is that pointers with the high bit set are commonplace, whereas pointers that are not aligned to the processor's word size are easily avoided (and on some architectures, outright invalid).
- TheLoneWolfling 12y agoSo then store the pointer right-shifted by 1.
- jzwinck 12y agoOn amd64, a pointer must have the top 17 bits all the same, or else it is not a valid pointer. On Windows and Linux, user space is the "lower half" which means bits 47 to 63 must all be zero for a value to be a pointer. This gives you 65535 tag values without using the LSB at all.
- robert_tweed 12y agoRegardless, you would lose the automatic guarantee that the same bit is zero for pointers, which is the whole motivation for doing this (more efficient GC). Pointers are guaranteed to have LSB of zero because of data is typically aligned to the machine word size rather than individual bytes. In this case, that is enforced (or at worst everything must be 16-bit aligned). It's a pretty clever idea really.
- glandium 12y ago> Pointers are guaranteed to have LSB of zero Except pointers to functions in ARM thumb. Or more simply pointers inside strings.
- koenigdavidmj 12y agoThe way that this is overcome is that most garbage collectors only permit direct pointers to an object rather than to a particular field inside an object. For example, there's no Java equivalent to this C: void a(int *b) { *b = 5; } struct c { int d; }; int main() { struct c foo; a(&foo.d); // no way to do this in Java! printf("%lu\n", (unsigned long)foo.d); }
- robert_tweed 12y agoAs has been pointed out, you can disallow direct pointers into the middle of things and use offsets instead. I'm curious though, because I don't know much about ARM, why function pointers can't be forced into even alignment? If it's 3rd party binary code you might have a problem, but if you control the compilation you can add NOP padding to get whatever alignment you want. There's nothing that requires any particular alignment on x86/x64 either, but it's standard for most compilers because it makes everything run faster at minimal size cost.
- belovedeagle 12y ago> why function pointers can't be forced into even alignment The actual instructions are aligned on an even byte boundary; however, on AArch32, encoding interoperability is achieved with BLX or BX, which uses the low bit as a tag value: 0=A32, 1=T16/32. However, in a lisp interpreter, e.g., you should never find a function pointer in an "object" context; i.e., something which could be an integer instead.