5 ms·
These 8086 instructions were pretty cool coming from a 6502 background. The usual way to copy a block of memory in 6502 was something like this: ldy #len
by ataylor284_ 4y ago
These 8086 instructions were pretty cool coming from a 6502 background. The usual way to copy a block of memory in 6502 was something like this:
ldy #len
loop:
lda source-1,y ; memory accesses: loads opcode, address, value
sta dest-1,y ; memory accesses: loads opcode, address, stores value
dey ; memory accesses: loads opcode
bne loop ; memory accesses: loads opcode, offset
So 4 instructions and a minimum of 9 memory accesses per byte copied, more if not in the zero page. Even unrolling the loop gets you down to 6.
Compare this to the 8086:
mov cx, len
mov si, source
mov di, dest
rep movsb ; memory accesses: read the opcodes, then 1 load and 1 store per byte
Even forgetting about word moves, you're down to 2 memory accesses per byte. No instruction opcode or branching overhead.
- kabdib 4y agoDon't forget the direction bit in the 8086's status register. [Don't forget the decimal mode bit in the 6502's status register :-)]
- chrononaut 4y agoWhat's noteworthy about the decimal mode bit in OP's example? (Not familiar with 6502 nuances, so this sounds interesting.)
- ataylor284_ 4y agoThey're both called the "D" flag, and they both globally control the behavior of certain processor instructions. Besides that, they're totally different. The 8086 D flag is the direction for string instructions, whether to increment or decrement the index registers after each repetition. One use was for moving memory between overlapping regions by starting at the end and working backwards. The 6502 D flag is for decimal mode. When set, the instructions ADC and SBC do BCD arithmetic instead of binary, e.g. 0x09 + 0x01 = 0x10 instead of 0x0A as you'd expect. The 8086 also has support for BCD, but took a different approach, using separate instructions DAA/DAS for the same purpose. Both can lead to bugs when you assume the flag is in the "normal" state but it somehow got flipped.
- ataylor284_ 4y agoLearned that one the hard way. I built a 6502 single board computer and tested it out running my customer built monitor. I spent hours verifying the hardware to isolate a problem where programs would crash and memory would become corrupted. Turns out my monitor was popping garbage into the flags register when it returned control and sometimes setting decimal mode.
- tenebrisalietum 4y agoThe enhanced 6502 derivative in the PC Engine/TurboGrafx 16 games console of the early 90's was enhanced with block move instructions (MVI, MVN?) that worked similarly I think. (Hu62C80 or similar was the CPU name...)
- rasz 4y agosadly Hu6502 sucked just as much. It had dedicated instructions, but cycle cost was ridiculous (17 + 6x) = ~160KB/s http://shu.emuunlim.com/download/pcedocs/pce_cpu.html http://shu.emuunlim.com/download/pcedocs/pce_cpu.html Transfer Alternate Increment (TAI), Transfer Increment Alternate (TIA), Transfer Decrement Decrement (TDD), Transfer Increment Increment (TII) For contrast 5 years older 80286 already did 'rep movsw' at afaik 2 cycles per byte. 6 years later Pentium did 'rep movsd' at 4 bytes per cycle. Nowadays Cannonlake can do 'rep movsb' full cachelines at a time at full cache/memory controller speed.
- rep_lodsb 4y agoWell the Z80 was worse. 21 cycles per byte! The reason is that instead of running a loop in microcode, it decremented PC by 2, then fetched the instruction again every time.
- deleted 4y ago[deleted]
- peterfirefly 4y agoIf you wanted to clear an area, LDIR/LDDR were slower than using a string of PUSH instructions. When moving data PUSH/POP were slightly slower than using LDIR/LDDR, though.
- tom_ 4y agoIt's worse than this for the 6502! DEY takes 2 cycles, and BNE (when taken) takes 3. LDA abs,Y takes 4-5 cycles (the extra one happens if a page boundary is crossed when performing the address calculation - in practice you'd try to arrange for this not to happen), and STA abs,Y takes 5 cycles. So out of the 24 cycles per iteration, there's 2 useful ones: 1 read, and 1 write. You need to unroll this sort of loop a few times to get anything near 9 cycles per byte, and ~80% of the cycles are still going to be kind of wasted. (Still annoys me to this day if I find myself writing code to copy stuff around. Why isn't it in the right place already?!)