5 ms·
Didn’t self modifying code just die out when cache came around in the late eighties and early nineties? We used it a lot on C64 and Amiga (until A1200 with 6802
by Max-q 4y ago
Didn’t self modifying code just die out when cache came around in the late eighties and early nineties? We used it a lot on C64 and Amiga (until A1200 with 68020 CPU).
- cwzwarich 4y agoAll JITs use self-modifying code, and JITs are more popular now than even 15 years ago. Although you might make a distinction between modifying code that has already executed once, from the perspective of instruction/data consistency on a CPU it doesn’t really matter.
- Jensson 4y agoJIT doesn't modify the behavior though, it typically just changes a function pointer and then that function does the same thing but more optimized. So then instruction caches just has some effect on performance instead of changing program behavior.
- cwzwarich 4y agoJITs also modify the code at the destination of the function pointer, which was likely also code at a different time (since it usually needs to be allocated from a separate heap that is marked executable). If the CPU prefetched/cached the instructions at the function pointer target prior to the modification, then the program would observe incorrect behavior, even though it's not self-modifying code in the usual sense.
- miohtama 4y agoSelf-modifying code in a traditional sense has been changing constants and instructions in a tight loops that are assembled by hand. I don't think any of JITs take this approach, as when they emit machine code it is always a fresh memory allocation and fresh compilation. JITs don't "poke and peek" already compiled code. But I could be wrong as JITs are getting complex today.
- ridiculous_fish 4y agoJITs do indeed modify already-compiled code. For example, JavaScriptCore has a notion of a "patchable jump." These are jump instructions intended to be modified later. They may initially point to slow paths, and are later patched to point to faster paths as type information accumulates. https://webkit.org/blog/10298/inline-caching-delete/ https://webkit.org/blog/10298/inline-caching-delete/
- miohtama 4y agoThis looks super interesting, thank you.
- anonymousDan 4y agoI think the Linux kernel has some limited forms of self modifying code?
- loeg 4y agoTracepoints are often implemented as self-modifying code (nop <-> trap instruction).
- retrac 4y ago1) Yes, cache, though it was specifically the introduction of split instruction/data caches. Modern processors are basically Harvard architecture internally, handling instructions and data separately to allow them to be processed in parallel. So either explicit instructions (RISC approach) or extensive tracking of cache entries (CISC approach) must be used to ensure that data cache changes are propagated to the instruction cache. 2) It greatly complicates things in terms of software engineering. It's bad practice unless strictly necessary. Code that rewrites itself is very hard to debug, trace and just generally work with, compared to immutable blocks of read-only code. Most high-level languages don't expose anything except function pointers in terms of self-modification, so compiled output would rarely need SMC. (Though it may be faster to emit SMC to implement high-level constructs, in some edge cases.) The hardware doesn't like it and maintainers, managers, and most professors don't like it, either.
- layer8 4y agoI still used it on a 486 in the mid-nineties in drawing routines to switch the drawing mode without having to have a dynamic function call in the inner loop (or having to duplicate the drawing routine for each mode). I think it was the Pentium that introduced split instruction/data caches on x86.