11 ms·
Why xor eax, eax?
- rhaps0dy 10mo agoNo RSS? I want to subscribe :'(
- sph 10mo ago“Who cares about RSS, no one uses it any more” There’s dozens of us! By the way, totally unaffiliated, but I have used fetchrss for those websites that have no feed.
- daeken 10mo agoBack in 2005 or 2006, I was working at a little startup with "DVD Jon" Johansen and we'd have Quake 3 tournaments to break up the monotony of reverse-engineering and juggling storage infrastructure. His name was always "xor eax,eax" and I always just had to laugh at the idea of getting zeroed out by someone with that name. (Which happened a lot -- I was good, but he was much better!)
- VectorLock 10mo agoI was there but never got in on the Quake 3 fun; mp3t**
- OgsyedIE 10mo agoThe page crashes after 3 seconds, 100% of the time, on the latest version of Android Chrome and works fine on Brave, fyi.
- robmccoll 10mo agoThis is not my experience on the latest version of Chrome Android (142.0.7444.171). It did not crash for me.
- pansa2 10mo ago> Unlike other partial register writes, when writing to an e register like eax, the architecture zeros the top 32 bits for free. I’m familiar with 32-bit x86 assembly from writing it 10-20 years ago. So I was aware of the benefit of xor in general, but the above quote was new to me. I don’t have any experience with 64-bit assembly - is there a guide anywhere that teaches 64-bit specifics like the above? Something like “x64 for those who know x86”?
- veltas 10mo agoChapter 3 of volume 1, ctrl+f for "64-bit mode", has a lot of the essentials including e.g. the stuff about zeroing out the top half of the register. https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html https://www.intel.com/content/www/us/en/developer/articles/t...
- huflungdung 10mo ago[dead]
- matt_d 10mo agoSee https://github.com/MattPD/cpplinks/blob/master/assembly.x86.md https://github.com/MattPD/cpplinks/blob/master/assembly.x86.... - mostly focused on x86-64 (and some of the talks/tutorials offer pretty good overview)
- sparkie 10mo agoIt's not only xor that does this, but most 32-bit operations zero-extend the result of the 64-bit register. AMD did this for backward compatibility. so existing programs would mostly continue working, unlike Intel's earlier attempt at 64-bits which was an entirely new design. The reason `xor eax,eax` is preferred to `xor rax,rax` is due to how the instructions are encoded - it saves one byte which in turn reduces instruction cache usage. When using 64-bit operations, a REX prefix is required on the instruction (byte 0x40..0x4F), which serves two purposes - the MSB of the low nybble (W) being set (ie, REX prefixes 0x48..0x4f) indicates a 64-bit operation, and the low 3 bits of low nybble allow using registers r8-r15 by providing an extra bit for the ModRM register field and the base and index fields in the SIB byte, as only 3-bits (8-registers) are provided by x86. A recent addition, APX, adds an additional 16 registers (r16-r31), which need 2 additional bits. There's a REX2 prefix for this (0xD5 ...), which is a two byte prefix to the instruction. REX2 replaces the REX prefix when accessing r16-r31, still contains the W bit, but it also includes an `M0` bit, which says which of the two main opcode maps to use, which replaces the 0x0F prefix, so it has no additional cost over the REX prefix when accessing the second opcode map.
- snvzz 10mo agoBecause, unlike RISC-V, x86 has no x0 register.
- jabl 10mo agoFrom your past posting history, I presume that you're implying this makes RISC-V better? Do we have any data showing that having a dedicated zero register is better than a short and canonical instruction for zeroing an arbitrary register?
- kevin_thibedeau 10mo agoIt's a definite liability on a machine with only 8 general purpose registers. Losing 12% of the register space for a constant would be a waste of hardware.
- menaerus 10mo ago8 registers? Ever heard of register renaming?
- Polizeiposaune 10mo agoEver heard of a loop that needed to keep more than 7 variables live? Register renaming helps with pipelining and out-of-order execution, but instructions in the program can only reference the architectural registers - go beyond that and you end up needing to spill some values to (architectural) memory. There's a reason why AMD added r8-r15 to the architecture, and why intel is adding r16-r31..
- menaerus 10mo agoI have but that was not the point? My first point was exactly that there are more ISA registers and not only 8, and therefore the question mark. My second point was about register renaming which, contrary what you say, does mitigate the artifacts of running out of registers by spilling the variables to the stack memory. It does it by eliminating the false dependencies between variables/registers and xor eax, eax is a great candidate for that.
- omnicognate 10mo agoIt happens to be the first instruction of the first snippet in the wonderful xchg rax,rax. https://www.xorpd.net/pages/xchg_rax/snip_00.html https://www.xorpd.net/pages/xchg_rax/snip_00.html
- dooglius 10mo agoNot sure what I am looking at here is this just a bunch of different ways to zero registers?
- omnicognate 10mo agoIt's a collection of interesting assembly snippets ("gems and riddles" in the author's words) presented without commentary. People have posted annotated "solutions" online, but figuring out what the snippets do and why they are interesting is the fun of it. It's also available as an inscrutable printed book on Amazon.
- _jzlw 10mo agoThat music when you click "int" is awesome. Reminds me of the good ol' days of keygens.
- Audiophilip 10mo agoIt's a chiptune-style xm module, "Funky Stars" by Quazar: https://soundcloud.com/scene_music/funky-stars https://soundcloud.com/scene_music/funky-stars
- therein 10mo agoKeygen music will always have a special place in my heart. This is a good one. I do wonder who was the first cracker that thought of including a keygen music that started the tradition. I also miss how different groups competed with each other and boasted about theirs while dissing others in readmes. Readme's would have .NFO suffix and that would try to load in some Windows tool but you had to open them in notepad. Good times.
- deleted 10mo ago
- eb0la 10mo agoI remember a lot of code zeroing registrers, dating at least back from the IBM PC XT days (before the 80286). If you decode the instruction, it makes sense to use XOR: - mov ax, 0 - needs 4 bytes (66 b8 00 00) - xor ax,ax - needs 3 bytes (66 31 c0) This extra byte in a machine with less than 1 Megabyte of memory did id matter. In 386 processors it was also - mov eax,0 - needs 5 bytes (b8 00 00 00 00) - xor eax,eax - needs 2 bytes (31 c0) Here Intel made the decision to use only 2 bytes. I bet this helps both the instruction decoder and (of course) saves more memory than the old 8086 instruction.
- vardump 10mo ago> - mov ax, 0 - needs 4 bytes (66 b8 00 00) - xor ax,ax - needs 3 bytes (66 31 c0) You don't need operand size prefix 0x66 when running 16 bit code in Real Mode. So "mov ax, 0" is 3 bytes and "xor ax, ax" is just 2 bytes.
- eb0la 10mo agoMy fault: I just compiled the instruction with an assembler instead of looking up the actual instruction from documentation. It makes much more sense: resetting ax, and bc (xor ax,ax ; xor bx,bx) will be 4 octets, DWORD aligned, and a bit faster to fetch by the x86 than the 3-octet version I wrote before.
- RHSeeger 10mo ago> the IBM PC XT days (before the 80286) Fun fact - the IBM PC XT also came in a 286 model (the XT 286).
- deadcore 10mo agoMatt Godbolt also uploads to his self titled Youtube channel: https://www.youtube.com/watch?v=eLjZ48gqbyg https://www.youtube.com/watch?v=eLjZ48gqbyg
- vanderZwan 10mo agoNot sure why you got downvoted for pointing that out - it might be linked at the end of the article but people can still miss that.
- deadcore 10mo ago*shrugs* the internet being the internet I suppose. There was "See the video that accompanies this post." but NGL was just posting encase anyone didn’t have time to read or missed it.
- brucehoult 10mo agoHe also runs a site with a bunch of different compilers and versions :p
- mattgodbolt 10mo agoThat's just some weird side hobby of his.
- brucehoult 10mo agoDude. You've become a verb.
- pclmulqdq 10mo agoIn modern CPUs, a lot of these are recognized as zeroing idioms and they end up doing the same thing (often a register renaming trick). Using the shortest one makes sense. If you use a really weird zeroing pattern, you can also see it as a backend uop while many of these zeroing idioms are elided by the frontend on some cores.
- Dwedit 10mo agoBecause "sub eax,eax" looks stupid. (and also clears the carry flag, unlike "xor eax, eax")
- tom_ 10mo agoxor clears the carry as well? In fact, looks like xor and sub affect the same set of flags! xor: > The OF and CF flags are cleared; the SF, ZF, and PF flags are set according to the result. The state of the AF flag is undefined. sub: > The OF, SF, ZF, AF, PF, and CF flags are set according to the result. (I don't have an x64 system handy, but hopefully the reference manual can be trusted. I dimly remembered this, or something like it, tripping me up after coming from programming for the 6502.)
- trollbridge 10mo agoThis is a good thing since the pipeline now doesn’t have to track the state of the flags since they all got zero’d.
- sfink 10mo agoStrangely, the only difference on the flags is that AF (auxiliary carry) is undefined for `xor eax, eax` but guaranteed to be zeroed for `sub eax, eax`. I don't know what that means in practice, though I'm guessing that at the very least the hardware would not treat it as a dependency on the previous value.
- HackerThemAll 10mo agoIf I remember correctly, sub used to be slower than xor on some ancient architectures.
- tony-john12 10mo ago[dead]
- sylware 10mo agoRemnant of RISC attempt without a zero register.
- sylware 10mo agoCome on... that was a joke... this karma system...
- fooker 10mo agoIt's funny how machine code is a high level language nowadays, for this example the CPU recognizes the zeroing pattern and does something quite a bit different.
- Reubensson 10mo agoWhat do you mean that cpu does something different? Isnt cpu doing what is being asked, that being xor with consequence of zeroing when given two same values.
- horsawlarway 10mo agoNo. It's emulating the zero result when it recognizes this pattern, usually by playing clever tricks with virtual registers.
- IsTom 10mo agoI think OP means that it has come a long way from the simple mental model of µops being a direct execution of operations and with all the register renamings and so on
- dooglius 10mo agoFTA: > And, having done that it removes the operation from the execution queue - that is the xor takes zero execution cycles!1 It’s essentially optimised out by the CPU
- fooker 10mo agoSame consequence yes. But it will not execute xor, nor will it actually zero out eax in most cases. It'll do something similar to constant propagation with the information that whenever xor eax, eax occurs; all uses of eax go through a simpler execution path until eax is overwritten.
- 12_throw_away 10mo ago> with consequence of zeroing when given two same values Right, it has the same consequence, but it doesn't actually perform the stated operation. ASM is just a now just a high level language that tells the computer to "please give me the same state that a PDP-11-like computer would give me upon executing these instructions."
- silverfrost 10mo agoBack on the Z80 'xor a' is the shortest sequence to zero A
- fortran77 10mo agoBack when I did IBM 370 BAL Assembly Language, we did the same thing to clear a register to zero. XR 15,15 XOR REGISTER 15 WITH REGISTER 15 vs L 15,=F'0' LOAD REGISTER 15 WITH 0 This was alleged to be faster on the 370 because because XR operated entirely within the CPU registers, and L (Load) fetched data from memory (i.e.., the constant came from program memory).
- vanderZwan 10mo ago> In my 6502 hacking days, the presence of an exclusive OR was a sure-fire indicator you’d either found the encryption part of the code, or some kind of sprite routine. Meanwhile, people like me who got started with a Z80 instead immediately knew why, since XOR A is the smallest and fastest way to clear the accumulator and flag register. Funny how that also shows how specific this is to a particular CPU lineage or its offshoots.
- jgrahamc 10mo agoIn my 6502 hacking days, the presence of an exclusive OR was a sure-fire indicator you’d either found the encryption part of the code, or some kind of sprite routine. Yeah, sadly the 6502 didn't allow you to do EOR A; while the Z80 did allow XOR A. If I remember correctly XOR A was AF and LD A, 0 was 3E 01[1]. So saved a whole byte! And I think the XOR was 3 clock cycles fast than the LD. So less space taken up by the instruction and faster. I have a very distinct memory in my first job (writing x86 assembly) of the CEO walking up behind my desk and pointing out that I'd done MOV AX, 0 when I could have done XOR AX, AX. [1] 3E 00
- vanderZwan 10mo agoHah, we commented on the exact same paragraph within a minute of each other! My memory agrees with your memory, although I think that should be 3E 00. Let me look that up: https://jnz.dk/z80/ld_r_n.html https://jnz.dk/z80/ld_r_n.html https://jnz.dk/z80/xor_r.html https://jnz.dk/z80/xor_r.html Yep, if I'm reading this right that's 3E 00, since the second byte is the immediate value. One difference between XOR and LD is that LD A, 0 does not affect flags, which sometimes mattered.
- jgrahamc 10mo agoYou're right. Of course, it's 3E 00. Not sure how I remembered 3E 01. My only excuse is that it was 40 years ago!
- sfink 10mo agoWhat is this "LD A, 0" syntax? Is it a z80 thing? One of the random things burned into my memory for 6502 assembly is that LDA is $A9. I never separated the instruction from the register; it's not like they were general purpose. But that might be because I learned programming from the 2 books that came with my C64, a BASIC manual and a machine code reference manual, and that's how they did it. I learned assembly programming by reading through the list of supported instructions. That, and typing in games from Compute's Gazette and manually disassembling the DATA instructions to understand how they worked. Oh, and the zero-page reference. Good times.
- bitwize 10mo agoBecause mov eax, 0 requires fetching a constant and prolongs instruction fetching/execution. XOR A was a trick I learned back in the Z80 days.
- dintech 10mo agoMy brain read this is "Why not ear wax?"
- kragen 10mo agoxor wax, wax ; clear wax xor sax, sax ; clear sax xor fax, fax ; tru tru
- jabedude 10mo agosimilarly IIRC, on (some generations of) x86 chips, NOP is sugar around `XCHG EAX, EAX` which is effectively a do-nothing operation
- bitwize 10mo agoThis is pretty much all x86 chips as far as I'm aware: opcode 0x90 which is equivalent to XCHG AX,AX. The 8080 and Z80's NOP was at opcode 0. Which was neat because you could make a "NOP slide" simply by zeroing out memory.
- kccqzy 10mo agoThere are multiple variants of nop mainly because you sometimes need the nop instruction to take up a certain number of bytes for alignment purposes. You have the 1-byte nop, but there is also the 9-byte nop.
- BiraIgnacio 10mo agoAlso cool this got at the top item on the HN front page
- sixthDot 10mo agoI've wrote a lot of `xor al,al` in my youth.
- flohofwoe 10mo agoThe actually surprising part to me is that such an important instruction uses a two byte encoding instead of one byte :)
- kccqzy 10mo agoEven supporting just 8 registers that would take up 8/256=0.03125 of the instruction encoding space.
- GuB-42 10mo agoThey could have made a version just for (E)AX. "general purpose" registers in x86 are not the same. AX is the accumulator, for arithmetic, BX is for indexing, CX is the loop counter and DX is for data and extending AX in divisions. You don't have to use them for that purpose, but you will have access more optimized instructions if you do. Out of these 4, AX is the most likely you would want to set to zero. For loops, it is generally expected that you count down, with CX. The "LOOP" instruction is designed for this, so no special need to zero CX. SI and DI, the index registers may benefit from an optimized zeroing, for use with the "string" instructions. Here I think Intel engineers didn't see the need and not having a special instruction to zero AX must simplify the decoder.
- HackerThemAll 10mo ago> Interestingly, when zeroing the “extended” numbered registers (like r8), GCC still uses the d (double width, ie 32-bit) variant. Of course. I might have some data stored in the higher dword of that register.
- rfl890 10mo agoWhich will still be zeroed.
- Tuna-Fish 10mo agoClearing e8 also clears the upper half. Partial register updates are kryptonite to OoO engines. For people used to low-level programming weak machines, it seems natural to just update part of a register, but the way every modern OoO CPU works that is literally not a possible operation. Registers are written to exactly once, and this operation also frees every subsequent instruction waiting for that register to be executed. Dirty registers don't get written to again, they are garbage collected and reset for next renaming. The only way to implement partial register updates is to add 3-operand instructions, and have the old register state to be the third input. This is also more expensive than it sounds like, and on many modern CPUs you can execute only one 3-operand integer instruction per clock, vs 4+ 2-operand ones.
- charles_f 10mo ago> By using a slightly more obscure instruction, we save three bytes every time we need to set a register to zero Meanwhile, most "apps" we get nowadays contain half of npmjs neatly bundled in electron. I miss the days when default was native and devs had constraints to how big their output could be.
- Filligree 10mo agoJS is just easier and takes less code. Which isn’t an excuse anymore. UI coding isn’t that hard; if someone can’t do it, well, Claude certainly can.
- charles_f 10mo agoI'm fine with that, but keeping some consideration to optimization should still be something, even in environments when constraints are low. The problem is when no-one cares and includes 4 versions of jquery in their app so that they don't have to do const $=document.getElementById, everything grows to weigh 1Gb, use 1Gb of ram and 10% of your CPU, and your system is as sluggish nowadays (or even more) than it was 10y ago, with 10x the ram and processing power.
- anticrymactic 10mo ago> so that they don't have to do const $=document.getElementById, ``` const window.$ = (q)=>document.querySelector(q); ``` Emulates the behavior much better. This is already set on modern version of browsers[1] [1] https://firefox-source-docs.mozilla.org/devtools-user/web_console/helpers/index.html https://firefox-source-docs.mozilla.org/devtools-user/web_co...
- dapperhn 10mo agoIt is set, but only in the developer console, not for JavaScript included with the website/app.
- saagarjha 10mo ago
- grimgrin 10mo agoI'd like to learn about the earliest pronunciations of these instructions. Only because watching a video earlier, I heard "MOV" pronounced "MAUV" not "MOVE" Not sure exactly how I could dig up pronunciations, except finding the oldest recordings
- pwg 10mo ago> Only because watching a video earlier, I heard "MOV" pronounced "MAUV" not "MOVE" Was it someone from an electronics background? Because MOV is also the acronym for Metal Oxide Varistor [1] from electronics and in the electronics world the acronym it is often pronounced "MAUV". [1] https://en.wikipedia.org/wiki/Varistor https://en.wikipedia.org/wiki/Varistor
- jmmv 10mo ago> It gets better though! Since this is a very common operation, x86 CPUs spot this “zeroing idiom” early in the pipeline and can specifically optimise around it: the out-of-order tracking systems knows that the value of “eax” (or whichever register is being zeroed) does not depend on the previous value of eax, so it can allocate a fresh, dependency-free zero register renamer slot. While this is probably true ("probably" because I haven't checked it myself, but it makes sense), the CPU could do the exact same thing for "mov eax, 0", couldn't it? (Does it?)
- electroly 10mo agoSure, lots of longer instructions have this effect. "xor eax,eax" is interesting because it's short. That zero immediate in "mov eax,0" is bigger than the entire "xor eax,eax" instruction.
- addaon 10mo agoYes, "mov r, imm" also breaks dependencies -- but the immediate needs to be encoded, so the instruction is longer.
- MobiusHorizons 10mo agoI believe it does in some newer CPUs. It takes extra silicon to recognize the pattern though, and compilers emit the xor because the instruction is smaller, so I doubt there is much speed up in real workloads.
- lucozade 10mo ago> couldn't it? (Does it?) It could of course. It can do pretty much any pattern matching it likes. But I doubt very much it would because that pattern is way less common. As the article points out, the XOR saves 3 bytes of instructions for a really, really common pattern (to zero a register, particularly the return register). So there's very good reason to perform the XOR preferentially and hence good reason to optimise that very common idiom. Other approaches eg add a new "zero <reg>" instruction are basically worse as they're not backward compatible and don't really improve anything other than making the assembly a tiny bit more human readable.
- 10mo ago
- ethin 10mo ago> In this case, even though rax is needed to hold the full 64-bit long result, by writing to eax, we get a nice effect: Unlike other partial register writes, when writing to an e register like eax, the architecture zeros the top 32 bits for free. So xor eax, eax sets all 64 bits to zero. I had no idea this happened. Talk about a fascinating bit of X86 trivia! Do other architectures do this too? I'd imagine so, but you never know.
- 201984 10mo agoAArch64 also zeroes the upper 32 bits of the destination register when you use a 32 bit instruction.
- flykespice 10mo agoI'm curious, why is that? I know x86-64 zeroes the upper part of the register for backwards compability and improve instruction cache (no need for REX prefix), but AArch64 is unclear for me.
- School-Cotton 10mo agoI don't know either, but why wouldn't backwards compatibility apply to aarch64? It too is based on a pre-existing 32-bit architecture.
- Narishma 10mo agoI don't think it's backwards compatible the same way x86-64 is.
- 201984 10mo agoIt's to break dependencies for register renaming. If you have an instruction like mov w5, w6 // move low 32 bits of register 6 into low 32 bits of register 5 This instruction only depends on the value of register 6. If instead it of zeroing the upper half it left it unchanged, then it would depend on w6 and also the previous value of register 5. That would constrain the renamer and consequently out-of-order execution.
- Quitschquat 10mo agoAt some point I could disassemble 8086 (16 bit x86/real mode) as a kid. Byte sequences like 31 C9 or 31 C0 were a sure way to know if a loop of some kind was being initialized. Even simple compilers at the time made the mov xx, 0 → xor xx, xx optimization.
- kstrauser 10mo agoWhy wasn't that a standard assembler macro, like ZEROAX or something? It seems to come up enough that it seems like there'd be a common shortcut for it. (Not suggesting it should be. Maybe that's a terrible idea, but I don't know why.)
- sfink 10mo agoI don't know, but one reason might be that with 8-bit opcodes you only have 256 instructions to play with, and many of those encode registers, so ZEROAX is burning a meaningful percentage of your total opcode space. And if you're not encoding it into a single byte, then it's pure waste: you already need XOR (and SUB), so you'd just be adding a redundant way of achieving the same thing. (Note that this argument doesn't completely hold up, since eg the 6502 had a fair number of undocumented opcodes largely because they didn't need all of them.) Though technically you said "assembler macro", not opcode. For that, I suspect the argument is more psychological: we had such limited resources of all sorts back then that being parsimonious with everything was a required mindset. The mindset didn't just mean you made everything as short as possible, it also meant you reused everything you possibly could. So reusing XOR just felt more fitting and natural than carving out a separate assembler instruction name. (Also, there would be the question of what effect ZEROAX should have on the flags, which can be somewhat inferred when reusing an existing instruction.)
- kstrauser 10mo agoI meant something defined in the assembler along the lines of .macro ZEROAX xor eax, eax .endm where it was defined with a semantically meaningful name, but emitting the exact same opcodes as when writing it out. I mean, I guess taking that to the logical extreme, you'd end up with... C. I dunno, it just seemed like the sort of thing that would have caught on my convention. I use to write lots of 6502 and 68k assembler, and 68k especially tended to look quite human-readable by the time devs ended up writing macros for everything. Perhaps that wasn't the same culture around x86 asm, which I admit I've done far, far less of.
- sfink 10mo ago
- flustercan 10mo agoAs a longtime developer currently perusing their first computer science degree, it makes me happy that I understood this article. Nearly makes all the trouble seem worth it.
- JuniperMesos 10mo ago> In this case, even though rax is needed to hold the full 64-bit long result, by writing to eax, we get a nice effect: Unlike other partial register writes, when writing to an e register like eax, the architecture zeros the top 32 bits for free. So xor eax, eax sets all 64 bits to zero. Huh, news to me. Although the amount of x86-64 assembly programming I've personally done is extremely minimal. Frankly, this is exactly the sort of architecture-specific detail I'm happy to let an ASM-generating library know for me rather than know myself.
- timonoko 10mo agoI had Atari Portfolio, sold without software, except basic command.com. I could only copy raw ASCII bytes from keyboard to file. When it is in COM-file, it is directly executable. What is the simple program I made from ASCII-bytes to read bytes in hexadecimal. I thought that for a week, but finally solved it. @grok: This is a classic and brilliant puzzle from the early days of DOS. The challenge is to write a self-displaying hex dumper (.COM file) using only the copy con command, which means every byte of the program must correspond to a standard, typeable ASCII character. This is a form of code-golfing that relies on a deep understanding of 8086 machine code. The solution is a clever piece of self-modifying code. Since many essential 8086 opcodes (like MOV, INT, RET) don't have printable ASCII equivalents, the program must construct these instructions in memory and then execute them. Here is a well-known 128-byte solution. I will first show you the ASCII characters you would type, and then explain how it works. The Program to Type You would have started by typing copy con hex.com and then carefully typed the following 128 characters (including spaces and case-sensitive characters) and finally pressed Ctrl+Z to save the file. 1 j!>1!QhH-@0X,jHj.Ph1PhD0Ph 0@h, j, h.0@h, h 0@h.H-@,jH1X,h.H.@,H1j j X,h.H.@,H1j j X.H-@H-@,jHj.Ph1PhE0Ph@0@h, j, h.0@h, h 0@h. How It Works: A High-Level Explanation This program is a marvel of 8086 trickery. Here is a simplified breakdown of what's happening: etc.etc
- timonoko 10mo agoMy program was definitively shorter. I think I did not bother with real hexadecimals. Just used last four bytes of characters to make a full byte. Used it as a bootstrap program. @grok: While your exact code is lost to time, it would have looked something like one of the ultra-small ASCII "dropper" programs that were once passed around. Here is a plausible 32-byte example of what the program you typed might have looked like. You would have run copy con nibbler.com, typed the following line, and hit Ctrl+Z: `j%1!PZYfX0f1Xf1f1AYf1E_j%1!PZ` This looks like nonsense, but to the 8088/8086 processor, it's a dense set of instructions that does the following: etc etc.
- timonoko 10mo ago97% of these millenials of HN do not understand the problem and its brilliant solution. That is why I was truly astonished @grok grokked it rightaway. BTW. It is not beyond possibility that this nibbler or dropper was made by myself and published in Usenet by me myself in 1989. Who else would have such a problem. It was a bankcrupt sale and the machine was sold as "inactivated".
- kwertyoowiyop 10mo agoIn this thread, we have found all the programmers born before 1975!
- vanderZwan 10mo agoHey, some of us are younger and happened to get into programming via making games on their TI-83 graphing calculator in Z80!
- wildlogic 10mo agoI learned this trick writing shellcode - the shellcode has to be null byte (0x00) free, or it will terminate and not progress past the null byte, since it is the string terminator. of course, when you xor something with itself, the result is zero. the byte code generated by the instruction xor eax, eax doesn't contain null bytes, whereas mov eax, 0 does.
- anhldbk 10mo agoYes, it's one of my favorite trick also.
- ternaryoperator 10mo agoThe origin AFAIK stems from the mainframe days. When using BAL (the assembly language for the IBM/360 family and its descendants), xoring was faster than moving 0 to the variable. Many of the early devs who wrote assembly for PCs came from mainframe backgrounds and so the idiom was carried over.
- jakewil 10mo agoI'm building a gameboy emulator and when I was debugging the boot ROM I noticed there was the instruction `xor A` (which xor's a with itself). I was wondering why they chose such a weird way to set A to 0. Now it makes sense -- since the boot ROM is only 256 bytes, they really needed to conserve space! Thanks for this, looking forward to the rest of the series!
- 3oil3 10mo agoWhat a great article! When the author mentionned "showing-off", that's what I thought at first, I mean, most of us have the "why not spend 2 hours trying to figure it out when you can read the manual for 2 minute" kind of mind-set, which is similar to the "why not make it really complex if we can make it simple". But no, it's actually a really smart idea!!
- struc_so 10mo agoIt’s not just about code size or cycle count anymore; modern OoO (Out-of-Order) processors treat this idiomatically. The renamer recognizes xor reg, reg as a dependency-breaking zeroing idiom immediately, which frees up the physical register allocation faster than a mov. It's fascinating how hardware optimization has effectively leaked into the instruction set definition over time.
- Suzuran 10mo agoIn some older IBM-built processors (channel controllers, the various iterations of the CSP), an xor of something against itself also had the effect of safely clearing a stored bad parity without triggering a parity check from reading the operand. You would see strategic clearing in this manner done by system software or firmware during error recovery or early initialization.