4 ms·
One major problem are games which generate code into ram and then execute it. I can't remember if there were any NES games doing that, but I've seen other 6502
by gulpahum 11y ago
One major problem are games which generate code into ram and then execute it. I can't remember if there were any NES games doing that, but I've seen other 6502 based games doing that.
- Drakim 11y agoMaybe a tiny VM could be added just for runtime generated code. (but to be honest I don't think there is a single NES game that does that)
- TickleSteve 11y agodynamically generated code was quite a common technique in the era of 6502. Its frowned upon now, but it was a major optimisation technique back then.
- Drakim 11y agoDo you have any resources on such techniques and tricks? I'm very much into 6502 from a NES background but I know little about dynamically generated code. I understand how it works, I just can't imagine what I would potentially use for it. :)
- woodman 11y agoThe only potential practical use that I can think of, outside of procedurally generated decompression engines, is this: You find yourself on an uncharted desert island with two sailors, a movie star, some other lady, a millionaire and his wife... and a crate of 4k roms. For reasons that would take far too long to explain here - your only salvation is to recreate the Atari game catalog on your coconut game console.
- jerf 11y agoSpeed. 6502s are 8-bit processors running in the neighborhood of 1MHz. Anything you can do to squeeze out a cycle is worth it. Rather than writing something that will repeatedly examine the data, and makes decisions, and then does things, write code that simply does things. Much faster, if you can spare the RAM for the code. Plus if you're really careful and clever, the code to "do things" can itself be the data! Sometimes, anyhow. Also, the world has changed a lot since then. Interpreters have less penalty on a modern chip than an old stupid chip, because branch prediction, prefetching, and multiple pipelines can really help with them, so it's relatively speaking cheaper to examine data and make decisions and the CPU will spend more time "doing things" as long as the data required and the branches taken are predictable, which they often are in this sort of code. And on the flip side, modern processors really want your code to be static, precisely so that all those optimizations can work well, along with code caches, micro-op caches, etc... constantly changing the code isn't good for performance on modern chips. The 6502 doesn't care how much the code is changing, it just executes the next opcode at the same speed regardless. Very different world.
- peteri 11y agoSadly I don't have any at the moment but I recall using self modifying code to write graphics to an Apple ][ hires screen. Thinking about it I wonder why I didn't copy the core of that down to the zero page and write the value into the instructions there. STA $12 is a cycle faster than STA $1234 maybe I couldn't find the space (or I was being kind to the OS + basic)
- gp2000 11y agoI don't have a 6502-specific example, but scaling an image with point sampling is one place I've seen it done. The inner-loop to scale a line might look like this: frac = 0; while (len > 0) { for (frac += scale; frac >= 1; frac--) *dst++ = *src; src++; } Instead of repeating the work of the inner loop every time, generate the code that has the scale baked in. For instance, doubling the image would output this code fragment repeated for the width of the source image: *dst++ = *src; *dst++ = *src; src++; And scaling down by half would be: *dst++ = *src; src++; src++; Though whether this is faster or the best technique really depends on the processor. It might be just as easy to pre-compute a lookup table pointing to the offset of the source pixel for each destination one. That's something an 8086 could do fairly easily and maybe a 6502, but not so much for a Z-80.
- gulpahum 11y agoOne case would be configurable code. The code is copied to ram and then tweaked based on some parameters. The modified code runs faster than having a bunch of branches to check/calculate result based on the parameters. It might even be easier to write that kind of code. Another case is to access a large range of memory, for instance, to fetch data from a large table. You have the code in ram and increment the high-byte address (HH) of a 'lda $HH00,x' or 'sta $HH00,x' to access larger ram area (because x index can only access 256 bytes). I've seen that in Vic-20 games, I don't know if NES games used it. One case is loop unrolling, a.k.a. speedcode. This is very common practice in modern 6502 demos. I don't think that the old NES games used it, but some modern NES demos may use it. See http://codebase64.org/doku.php?id=base:speedcode http://codebase64.org/doku.php?id=base:speedcode or http://csdb.dk/forums/index.php?roomid=11&topicid=96279&showallposts=1 http://csdb.dk/forums/index.php?roomid=11&topicid=96279&show...
- jimsmart 11y agoOn the C64, when scrolling, the fastest way to update the screen colour RAM (which cannot be pointed to another address, unlike the screen character map) is to do it all with immediate LDA #$xx / STA $D800 / LDA #$xx / STA $D801 / etc (where xx are your colours, and $D800 is the base address of the colour RAM) - then during the 7 frames of pixel scrolling, you both offset-copy the screen characters to the second screen buffer, and also move all the character colour info (the values throughout the previous LDAs) through the colour RAM update code, and on the 8th frame you flip character screens and then call the colour RAM splat subroutine. [edit] It's much faster than the obvious solution, which is to do LDA $D801 / STA $D800 / LDA $D802 / STA $D801 / etc. Or, even worse, a loop incrementing the X register with LDA $D801,X / STA $D800,X / etc. Or worse still: indirection via zero-page (though no-one should really ever consider that for colour RAM updates, even though at first glance it seems clever for moving the characters between screens, it eats too much time) (This technique is pretty much only applicable if your scroll is fixed-direction and fixed-speed) [edit #2] Credit for that goes to Jon Williams (Shadow Dancer, C64 - and others) for adding that optimisation to my scroll routines that he used in SD - and for then telling me the trick :)
- jimsmart 11y agoMy anecdotal evidence ties with yours, to a greater degree. Not specifically generating code to run, but I certainly wrote plenty of self-modifying code when coding games on the C64, although it was never a trick I personally used on the NES - two reasons: by default your code runs from ROM anyhow (so there's just less of a tendency to consider self-mod code in the first place), plus the RAM on the NES is so very limited.