4 ms·
How is that assembly has code as data just like Lisp? Is it possible to create executable code during runtime, change the language's syntax, and stuff like that
by alberich 14y ago
How is that assembly has code as data just like Lisp? Is it possible to create executable code during runtime, change the language's syntax, and stuff like that, using assembly?
- danielweber 14y agoIf your architecture allows it, sure. I was reading some old Apple II assembly last year, and one very common way to do loops is like this: 1. Is 10 equal to 0? 2. If so, go to step 6. 3. Do a bunch of stuff. 4. Decrement the first number written in step 1. 5. Go to step 1. 6. Rest of program.
- ajross 14y agoTo be clear, even on architectures that allow it (or OS environments: the NX bit on modern x86 boxes means that you can't write to code space without mmap'ing the region yourself), it's a Very Bad Idea. Code and data use different L1 caches (on Sandy/Ivy Bridge, there's even a still lower level "uop" cache for code!), and keeping them coherent after self-modifying code has run is (1) extraordinarily difficult to get right and (2) hugely slow, much slower than simply indirecting on data in the first place.
- marssaxman 14y agoAssembly language is just a high level description of machine code: each assembler statement maps to one machine instruction. Those machine instructions are nothing but sequences of bytes. It is certainly possible to create executable code during runtime; you just write the appropriate bytes into a buffer, mark the buffer as executable, and jump in. Back in the old days, before memory protection, we used to do all kinds of code-as-data-as-code tricks. I remember dynamically generating trampoline functions that effectively did partial application. I'd write a little stub of assembly code that would push a bunch of literal zeros onto a stack and then jump to zero; when I wanted to use it, I'd copy it, poke actual values into the literals, and then pass it along as a zero-parameters function pointer. You don't change the language syntax; that doesn't really make sense in assembly world. It's more that assembly code lives in a world of bytes and pointers, and is itself very directly composed of bytes and pointers, and so it's as easy to use the same bytes-and-pointers concepts on code as it is on data. This style still exists on microcontrollers, where there is generally no operating system. Instead, your compiler produces an image which you flash onto the controller chip. The controller has some startup sequence which jumps to some address in flash and runs. The chip's subsystems are controlled by registers located at specific addresses; in order to drive such a device you have to know how its address space is laid out, where your code is placed in it, and how your code interacts with the different types of memory and registers located in the address space.
- mturmon 14y agoOf course, I did not say assembly is just like Lisp in every respect. But as long as we're picking nits -- It is of course completely possible to create executable code during runtime, using assembler. Often not advisable. (Recall the punch line of: http://www.cs.utah.edu/~elb/folklore/mel.html http://www.cs.utah.edu/~elb/folklore/mel.html)
- alberich 14y agoReading the story on the link made me feel envy :) I admire those who really understand how it all works under the hoods. Although I tried to learn assembly on my own a long time ago, I never went much further than reading some imput and printing it to the screen. It seems it is a very specialized knowledge nowdays.
- rayiner 14y agoThe best way to learn assembler is by writing an assembler, which you can do in pretty much any language you're comfortable with (https://github.com/rayiner/amd64-asm https://github.com/rayiner/amd64-asm) with just a manual (http://support.amd.com/us/Processor_TechDocs/24594_APM_v3.pdf http://support.amd.com/us/Processor_TechDocs/24594_APM_v3.pd...) . I think x86 is really not a bad pedagogical tool. For all the crap it gets, it's really a fairly clean architecture. And while we think of 32-bit and 64-bit extensions as having "piled on cruft" what's really happened is that they have made the architecture conceptually cleaner and more orthogonal.
- ANTSANTS 14y agoLiterally, yes. It's arguably much more readily apparent that code is just data when programming in assembly than in any other environment. It's all bytes, and you're forced to accept that from day 1. You aren't going to have the expressiveness that a Lisp provides available to you, but self-modifying code and macros assemblers are common in some assembly programming scenes. Anything to save a few bytes or cycles. I'm a little rusty, but here's a basic example off the top of my head (sorry, wall of text incoming): Functional programmers are familiar with the concept of "map," an operation that applies a function to every element in a list or what have you. Let's say I'm programming in 6502 assembly, and I have a little function that adds an amount to every byte in a page (or every byte in some particular 256 bytes). Let's say you also want a similar function that instead multiplies each byte by two (just a single left shift), or masks some bits off with a bitwise AND, or whatever. It would look something like this: store zero in X register, add constant to the memory at the address (some constant address + value in X register), increment X, branch back to the adding part if the zero flag is not set (ie, loop until X overflows to 0), then return to wherever we called this function from. You could write a handful of these almost identical subroutines, with the only real difference being the single opcode that reads, modifies, and rewrites each byte... or you could just rewrite the opcode at runtime! Now, your add-to-page, multiply-page, mask-out-page, and any similar functions all share a little block of code that you can think of as "their map," and the actual functions you call initially could look something like this: Write the constant for the relevant opcode to the instruction in map you want to replace ($7D for ADC absolute,X on the 6502), and jump to the map subroutine. That's a bit of a contrived example, but in a routine with a more complicated access pattern, it might really make a difference in byte savings: Imagine a routine that clips offscreen entities in a game, that calculates something like "if their X coordinate is below or above some value, they disappear for now. If their Y coordinate is below some value, they fell into a pit and died." You could reuse a lot of the general logic for checking the left edge of the screen for the other two edges, just by rewriting a constant and instruction or two. A perhaps more readily useful example is rewriting "constant" addresses. How would we rewrite our above "map" routine to modify arbitrary pages, and not just one in particular? We could store the address of the page we want to modify in memory, and use an indirect addressing mode for add. Indirect addressing modes of instructions do something like this: Load two bytes from some address in memory, and treat that as the address we want to operate on. Problem is, this indirect address mode adds two cycles to every operation we perform when compared to the original constant-address map. Suddenly, we're wasting over 500 cycles per map! The solution is instead to treat the "constant" address in the instruction stream of the map routine as your address variable. Just rewrite those two bytes during your "function prologue," and voila, you have a general-purpose map routine that only uses half-a-dozen cycles or so more than one that only worked on a certain page. The part I've been leaving out is using assembler macros to automate a lot of this stuff for you. Again, I'm rusty, and I never got all that experienced in writing macros, but someone could very easily write themself a macro (if they're using a powerful macro assembler, like ca65) that takes a single argument, the instruction you want to execute in your map, and generates for that stub subroutine that replaces the opcode for them. I found simpler macros to be more useful in everyday code, though: you make a macro that fills in a gap in the 6502's instruction set, like performing arithmetic between the accumulator and index registers, or basic 16-bit arithmetic, and from then on, you can pretend that the CPU had those instructions all along. I may have only used these techniques for shaving bytes and cycles off of straightforward routines, but in some ways, going from writing 6502 assembler to writing C, Lua, and JavaScript actually feels like a step back in terms of expressiveness, even if they are certainly more productive languages in actuality, and first-class functions/function pointers cover much of the most practical (and least dangerous) use-cases for self-modifying code. I suppose I won't get that feeling back until I set some time aside to really learn a Lisp.