5 ms·
It's odd that you were using ring buffers in 1992 for low level code but don't understand the value of avoiding a modulus instruction. Masking is far more effic
by kitsuac 10y ago
It's odd that you were using ring buffers in 1992 for low level code but don't understand the value of avoiding a modulus instruction. Masking is far more efficient and often a ring buffer will be used in code where performance is absolutely critical.
- jgrahamc 10y agoYou wouldn't use the modulus operation. You aren't adding some arbitrary number that's going to make you increase either index by more than the buffer length so you know that at worse you are going to need to subtract the length of the buffer. IIRC the way we made this really fast was the write the buffer backwards. That way you can detect wrapping around the buffer because DEC will underflow and set the sign flag. Then you can JS to whatever code needs to ADD back the buffer length to handle the wrap around. But 2^n has another problem (back in that era): buffer size. You are stuck with 1K, 2K, 4K, etc. buffers. When memory is tight you likely need something very specific, so you end up with the solution we had. But, hey, if memory is free use 2^n bytes for your buffer.
- deleted 10y ago[deleted]
- kitsuac 10y agoYou are still introducing a conditional by detecting the need to subtract, and iterating backward through memory is horrific for cache performance. If you need a specific, non power-of-2 sized buffer, then of course you make that design decision and pay the performance penalty. But I restate it's odd that you weren't even aware of the cost in 1992 as a system level programmer.
- jgrahamc 10y agoThe 286/386 didn't have a cache so that wasn't a worry at that time.
- coldtea 10y ago>But I restate it's odd that you weren't even aware of the cost in 1992 as a system level programmer. Costs were very different in the pipelines (or lack thereof) of eighties/early nineties hardware.
- to3m 10y agoHave you tested this recently? I haven't for some years now but performance was identical regardless of direction. Maybe I need to try it again. I'd expect going backwards to be no worse than "not as good" - like say perhaps the prefetching mechanism doesn't cater for this case - but maybe my standards aren't high enough and this is enough to tip things over into the horrific.
- tbirdz 10y ago>iterating backward through memory is horrific for cache performance This isn't true for Intel chips since Netburst Pentium 4. The hardware prefetchers can handle predicting iterating through an array forwards, backwards, and even strided accesses [0]. The arrays takes up the same number of cache lines in both cases, so going forwards or backwards are still going to have the same number of cache misses. 0: https://software.intel.com/en-us/articles/optimizing-application-performance-on-intel-coret-microarchitecture-using-hardware-implemented-prefetchers/ https://software.intel.com/en-us/articles/optimizing-applica...
- SamReidHughes 10y agoYou don't need a conditional. You can set up a mask using sbb. ; precondition: 0 <= x <= N ; (N is constant) mov y, 0 ; set up mask cmp x, N-1 ; set carry flag if x >= N sbb y, 0 ; subtract 1 from y if carry flag set and x, y ; set x to zero if x == N
- jnordwick 10y agoThe cmov looks better than sbb, but both have data dependencies than a predicted branch wouldn't.
- SamReidHughes 10y agoAhh! Serves me right for reading books from before the 486 :P