7 ms·
Solving the Mystery of ARM7TDMI Multiply Carry Flag
- skrrtww 2y agohttps://shonumi.github.io/blog/nds_rolling.html https://shonumi.github.io/blog/nds_rolling.html More context on how this value affects (at least one) DS game- see post from December 27th, 2019.
- userbinator 2y agoI'm not really familiar with ARM Asm but do you think that's handwritten Asm that its author overlooked the effects of the carry and it just happened to work, or a clever "emulator trap" added by Nintendo's compiler?
- bonzini 2y agoIt's handwritten and hand-obfuscated assembly, the author didn't know the effect of the carry but knew it was deterministic; and the code worked. Though it's not clear to me where the corruption after the second MLA instruction comes from, because the second block of three instructions should produce the same output as the first. It is possible that it was copied/pasted incorrectly.
- comex 2y agoAre you sure it's handwritten or obfuscated? I remember from when I used to disassemble compiled ARM code (not on the NDS though) that it was common to see CMP, followed by a bunch of instructions with one condition predicate, followed by a bunch of instructions with the opposite predicate. In this case, it's subtly wrong to use that pattern, but only on older versions of ARM. That could reflect a very sneaky attempt to break emulators… but it could also just be a compiler bug. That said, I too don't understand how corruption could be produced unless there was a copy/paste mistake.
- userbinator 2y agoAnd just to get this out of the way, the carry flag’s behavior after multiplication isn’t an important detail to emulate at all. Software doesn’t rely on it. On as fixed of a hardware as a game console, and with the accompanying anti-piracy/anti-cheating/emulation efforts of that industry, I'd expect it to be. From the history of emulating previous consoles, we know that any deterministic difference can and will be exploited, either to determine whether the hardware is authentic, or incidentally as a result of unintentional bugs. This reminds me of the Z80, where two undefined flags resisted analysis for several decades; a 2-year-old set of slides on the state of that here: https://archive.fosdem.org/2022/schedule/event/z80/attachments/slides/5207/export/events/attachments/z80/slides/5207/z80_last_secrets.pdf https://archive.fosdem.org/2022/schedule/event/z80/attachmen...
- comex 2y agoWhile it’s a bit newer than the GBA, there is at least one Wii game with intentional anti-emulation measures: https://tcrf.net/Cars_2_(PlayStation_3,_Xbox_360,_Windows,_Wii)#Anti-Dolphin_Code https://tcrf.net/Cars_2_(PlayStation_3,_Xbox_360,_Windows,_W...
- deleted 2y ago[deleted]
- DevilStuff 2y agoIn the GBA scene, people didn't actually tend to exploit the carry flag at all. If there was any anti emulation, it was usually flashcart related or cpu timing related.
- Dwedit 2y agoIt was mostly the Game Pak Prefetch feature that was used to foil GBA emulators. A game can detect if the number of cycles to access ROM (with prefetch enabled) is incorrect.
- f1shy 2y ago>> the carry flag’s behavior after multiplication isn’t an important detail to emulate at all. Software doesn’t rely on it. Famous last words: https://www.hyrumslaw.com/ https://www.hyrumslaw.com/
- ujikoluk 2y ago> Seriously, they decided that the program counter should be a general purpose register. Why??? Don't really understand this reaction. Why not? Seems to make for a nice regular design that the PC is just another register.
- DevilStuff 2y agoThe big issue is that it doesn't really need to be a GPR. You never find yourself using the PC in instructions other than in, say, the occasional add instruction for switch case jumptables, or pushes / pops. So it ends up wasting instruction space, when you could've had an additional register, or encoded a zero register (which is what AARCH64 does nowadays).
- joosters 2y agoBut it is (or was originally) used in lots of places, not just jump tables, generally to do relative addressing, for example when you want to refer to data nearby, e.g. ADD r0, r15, #200 LDR r1, [r15, #-100] etc
- DevilStuff 2y agoAh I miscommunicated, I still think PC can and should be used in places like the operand of an LDR / ADD. It's using it as the output of certain instructions (and allowing it to be used as such) that I take issue with. ARMv4T allowed you to set PC as the output of basically any instruction, allowing you to create cursed instructions like this lol: eor pc, pc, pc
- immibis 2y agoIsn't writing to it except by a branch instruction undefined behaviour? If you can use it as an operand, it has a register number, so you can use it as a result, unless you special-case one or the other, which ARM didn't do because it was supposed to be simple. They could have ignored it by omitting some write decode circuitry, but why?
- flohofwoe 2y ago> it allows the program counter to be used a general purpose register From a CPU emulator writer's perspective this isn't all that strange. For instance on Z80 the immediate jump instruction `JP nnnn` is loading a 16-bit immediate value into the internal PC register, which is the same thing as loading a 16-bit value into a regular register pair (e.g. 'LD HL,nnnn') - e.g. the mnemonics for the jump instruction could just as well be `LD PC,nnnn` ;) A relative jump (which does a signed-add of an 8-bit offset value to the 16-bit address in PC) is the same math as the Z80 indexed addressing mode (IX+d) and (IY+d) (I don't know though if the same transistors are used). A RET (load 16-bit value from stack into PC) is the the same operation as a POP (load 16-bit value from stack into a regular register pair). ...so it's almost surprising that the program counter isn't exposed as a regular register in most (traditional) CPUs. I guess in modern CPUs it's not so simple because of the internal pipelining though.
- adrian_b 2y agoThe reason why the PC is not normally exposed as a regular register is that the set of operations needed for the PC is different than for the regular registers. There are operations required for the PC that are not needed for regular registers, e.g. conditional add (a.k.a. conditional relative jump), add-and-store and load-and-store (a.k.a. procedure call). On the other hand, there are many operations that are needed for regular registers and which are useless for the PC, e.g. logical operations, shift/rotate, multiplication and division and others. Because of this, encoding the PC as a regular register is pointless and wasteful of the instruction encoding space. Moreover, when the ISA has an implicit stack pointer, which is also the only register that can be used as a stack pointer, like the x86 ISA, the set of operations that are used with the SP is a very small subset of the operations available for the regular registers, so encoding the SP as a regular register is also wasteful. Especially in 32-bit x86, where the number of architectural registers was very small, it would have been better if the SP would not have been encoded as a regular register, wasting a register number.
- addaon 2y ago> The reason why the PC is not normally exposed as a regular register is that the set of operations needed for the PC is different than for the regular registers. I'd add to that, that what you give is the reason it's /okay/ to expose the PC as a special register instead of a GPR. The reason that it's /important/ to is that the PC is accessed on every instruction fetch, so if it's part of a uniform register file, it basically eats up an entire read port of that register file. Register file size scaled badly with port count (much worse than it does with register count), so this ends up adding quite a bit of area. (You can hack around this by having a single dedicated read port for just the PC register, but then you're half way to an SPR.)