3 ms·
You can definitely round a float to the nearest integer-represented-as-a-float by accessing its bits. It will work. It will not be faster on a modern processor
by pascal_cuoq 13y ago
You can definitely round a float to the nearest integer-represented-as-a-float by accessing its bits. It will work.
It will not be faster on a modern processor, though, because floating-point numbers live in their own registers, distinct from integer registers. Floating-point registers do not have bitwise operations, moving data to and from integer registers is comparatively expensive, but they have IEEE 754 arithmetic.
So the game is to do as much as possible with what instructions there are. One such instruction is for instance (for processors with SSE2) " cvttss2si %xmm0, %eax", to truncate a float in xmm0 to an int in eax. Intel provided an instruction for that because the C language defines the cast float -> int as a truncation.
The next blog post in the series will be about using IEEE 754 addition to avoid the round-trip through eax altogether.
- pbsd 13y agoSSE4 added FRNDINT (x87) and ROUNDSS/SD (xmm) to do this task purely on hardware. Further, if one is seriously determined, there are ways to perform bitwise on floating-point registers directly. On XMM, it's trivial: just take advantage of the integer (and non-integer, e.g. XORAPS) instruction available. On x87, one can take advantage of the direct mapping between MMX and x87 registers to manipulate them using MMX's integer instructions. This has problems, but as I said -- seriously determined.
- gsg 13y agoOn modern x86 hardware, integer instructions that act on float values in xmm registers (and vice versa) require a somewhat expensive change of execution domain. That the register names are the same is irrelevant: everything will be renamed into what the CPU works with internally (which is separate integer and floating point domains). Mixing integer mmx and x87 is even worse, with the crazy emms junk. Not worth it.
- DarkShikari 13y agoThe penalty is at most ~1 cycle of latency -- in practice I find it gets completely absorbed by the OOE engine. I've never measured any significant penalty in any code for mixing float and int SSE operations on any x86 microarchitecture. Floating point bitwise operations exist too: xorps, andps, and so on.
- gsg 13y agoHmm, ok. Intel recommend avoiding it pretty strongly: I guess they overstate their case. One cycle in (and one out?) isn't exactly crushing.