4 ms·
Also consider instructions for efficient atomic reference counting, with traps on both inc (overflow) and dec. In particular, they can have weaker ordering sem
by devit 11y ago
Also consider instructions for efficient atomic reference counting, with traps on both inc (overflow) and dec.
In particular, they can have weaker ordering semantics and they can be buffered and elided among themselves (obviously with some sort of inter-core snooping).
And possibly support for "tagged" numbers, e.g. add integers if high bit is not set, call function otherwise, same for floats if not NaN, with a predictor for them.
- vardump 11y ago> Also consider instructions for efficient atomic reference counting, with traps on both inc (overflow) and dec. Atomic reference counting is slow and gets even worse the more CPU cores and especially CPU sockets you have. If you can afford to make as expensive operation as an atomic add, you can definitely afford to add overflow checks. Atomic add is 50-1000+ clock cycles depending on contention, core/socket count and "moon phase" -- latency is somewhat unpredictable. > In particular, they can have weaker ordering semantics and they can be buffered and elided among themselves (obviously with some sort of inter-core snooping). I'm not sure how weak ordering semantics and fetch-and-add (atomic add) could mix. Aren't atomics about strong ordering by definition? Maybe there's something I don't understand. > And possibly support for "tagged" numbers, e.g. add integers if high bit is not set, call function otherwise, same for floats if not NaN, with a predictor for them. You'd still get branch mispredict which I guess you're trying to avoid. There'd be no performance improvement.
- Taniwha 11y agodone right you'd not predict this branch, it's the exception that would get the mispredict
- vardump 11y agoTagging is done to carry information about data type. Like to mark that float64 is actually a 32-bit integer. Traps (CPU exceptions, such as traditional FPU exceptions like division by zero) usually involve kernel mode context switch. So if you trap on tag, the performance for tagged values will probably be 3-5 orders of magnitude slower. That's a lot.
- renox 11y ago> Traps (CPU exceptions, such as traditional FPU exceptions like division by zero) usually involve kernel mode context switch. Could you explain why? I thought that trapping was more like a 'slow branch': slow due to the flush the pipeline but why should the kernel be involved(1)? 1: except if you need to swap in a page, but that's just like any other memory reference.
- vardump 11y agoWhen we're talking about x86, that's true in ring 0. Otherwise first thing CPU does is to enter privileged, ring 0 mode, save registers, jump through interrupt vector table and process the trap in kernel code. Trap handler will probably need to check usermode program counter and take a look at the instruction that caused the trap. No hard data, but I think we're talking about 1-5 microseconds. Runtime/language exceptions have different mechanisms that don't require kernel context switches (but might involve slow steps like stack walk).
- deleted 11y ago[deleted]
- KMag 11y agoIf you want to get better performance out of dynamic language implementations that use NaN-tagging, you'll likely get better performance by adding one instruction that performs an indirect 64-bit load using 52-bit or 51-bit NaN-tagged addresses. The instruction should probably contain an immediate value for a PC-relative branch if the value isn't a properly formatted NaN-tagged address. All languages would benefit from instructions to more efficiently support tracing of native code. A pair of special purpose registers (trace stack and trace limit registers) to push all indirect and conditional branch and call targets would really speed up tracing of native code a la HP's Project Dynamo. Presumably upon trace stack overflow the processor would trap to the kernel or call to userspace interrupt vector entry. A small pseudorandom number generator and another pair of special purpose registers (stack and limit register) for probabilistically sampling the PC would make profiling lighter weight, both for purposes of human analysis of code and also for runtime optimization in JITs or HP Dynamo-like native code re-optimization.