4 ms·
Assembly Hall of Shame
- vardump 2mo agoA great resource for any performance deoptimization.
- 2_foos_in_a_bar 2mo ago[dead]
- arn3n 2mo agoThere’s definitely strategies here; A lot of the floating point operations use subnormals, and a lot of the worst instructions are slowed down by really, really fucking with MMIO.
- TomatoCo 2mo agoThis author also has other things like: A compiler that emits only `mov` instructions and another compiler that deliberately messes with the control flow so that, if disassembled, common debuggers will draw symbols like skulls or threats. https://github.com/xoreaxeaxeax/repsych https://github.com/xoreaxeaxeax/repsych
- inigyou 2mo agoHe also bruteforced the entire opcode space to find undocumented instructions (sandsifter).
- vanderZwan 2mo agoAlso came up with the original ..cantor.dust.. binary visualization tool, which is a tool I never used directly but Chris' presentation of it in 2012 is still one of the coolest talks I've ever seen. [0] https://github.com/Battelle/cantordust https://github.com/Battelle/cantordust [1] https://www.youtube.com/watch?v=4bM3Gut1hIk https://www.youtube.com/watch?v=4bM3Gut1hIk
- Gibbon1 2mo agoI saw a string of emoji that supposedly when copied to a file and renamed .exe prints out hello when run.
- metadat 2mo agoIt’s crazy how computers still seem to get perceivably slow every few years, given how many instructions can be executed in 1ms. Shameful, even.. What’s that law called about programmers wasting all the compute on abstraction?
- LoganDark 2mo agoHuh? A millisecond is an eternity!
- m463 2mo agoI remember reading once somewhere: If some app responds in 10ms or less, it is INTERACTIVE. makes you think.
- Xirdus 2mo agoIt is literally impossible to respond to input in 10ms on most platforms, for various reasons. The USB input lag of 12-30ms and the 60Hz refresh rate of most monitors being just the first two.
- xboxnolifes 2mo ago60Hz monitors definitely prevent it, but I'm pretty sure USB lag is far less than 12-30ms. My USB mouse can make a round-trip to a remote server faster than that.
- Xirdus 2mo agoOut of curiosity, how did you measure mouse roundtrip with sub-12ms precision?
- m463 2mo agoI stand corrected. I looked it up and it is .1 seconds (100ms) The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]: - 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result. - 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data. - 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect. from Jakob Nielsen: https://www.nngroup.com/articles/response-times-3-important-limits/ https://www.nngroup.com/articles/response-times-3-important-... less readable but the original paper: https://www.yusufarslan.net/sites/yusufarslan.net/files/upload/content/Miller1968.pdf https://www.yusufarslan.net/sites/yusufarslan.net/files/uplo...
- codeshaunted 2mo agowhat im seeing from this chart is that we should be using the nop instruction for everything
- achierius 2mo agoIt'd be really interesting to see whether the winning (losing?) instructions/strategies would be different on other architectures. At least right now the top spot (`fxrstor64` on MMIO, starve PCIe) seems relatively architecture-independent, but maybe something about MMIO ordering rules on e.g. POWER would be different enough to change that -- or perhaps open up new avenues? I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
- inigyou 2mo agoI think the idea was to find a long instruction on an ordinary PC. Of course by adding special hardware you can stall things.
- spoocecow 2mo agoOh wow, glad to see Chris Domas active online again!
- layer8 2mo agoNop should be #1, because it is infinitely slow for what it does. ;)
- jooops1 2mo agoIt increments rip by one.
- fluoridation 2mo agoNo, that's done by the decoder. It actually does nothing.
- russdill 2mo agoI mean....there are several architectures out there which has a nop that is a jump forward. Kind of a tree forest issue imho
- fluoridation 2mo agoI don't know about other architectures in as much detail. I know x86 NOP does nothing.
- loeg 2mo agoThe decoder is an implementation detail that is a subcomponent of NOP; GP was right, and your correction isn't.
- fluoridation 2mo agoIt's not an implementation detail, because the decoder runs before the execution of every instruction. If we're going to say that NOP increments IP by one, then we should also say that ADD "stores in dst the addition of src and dst, as well as incrementing IP by the length of the instruction", and JMP imm "increments JMP by imm + the length of the instruction".
- Retr0id 2mo agoRelated, and linked in the readme: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii (using the slow instructions to break SMI)
- jonathrg 2mo agoI wish they would just explain it in normal terms instead of this nasty LLM "engaging blog post" style
- twothreeone 2mo agoChris usually takes an educational angle, I don't think this is LLM-generated content at all it's just his style. I highly encourage watching some of his DefCon or BlackHat talks, they're fun!
- Brian_K_White 2mo agoI looked at both links and don't see anything weird or annoying, and I hate overblown styles myself.
- benmmurphy 2mo ago‘The counters tell the story’ might be something they consider odd. But I’m not familiar with the author’s style so they could have had these tics pre-LLM or they picked up these tics from reading a lot of LLM content.
- jonathrg 2mo agoYes, that's the point where I was like "wait a second" and started counting em dashes. A third alternative, perhaps the LLMs were trained hard on his prior work? :-)
- trebligdivad 2mo agoYeh that's a really nice one; I don't see anything in that suggesting it's been fixed (or even reported?)
- michalsustr 2mo agoVery cool! Also, huh interesting. I’ve used rdtsc to measure cycle diffs but had no idea its execution takes that long. Is that common across architectures?
- inigyou 2mo agoAFAIK it acts as some kind of execution barrier, to give meaningful timing.
- rrampage 2mo agoIsn't that rdtscp ( https://www.felixcloutier.com/x86/rdtscp https://www.felixcloutier.com/x86/rdtscp )?
- deleted 2mo ago[deleted]
- pbsd 2mo agoThe cycle count for RDTSC is ~25 cycles on Skylake-era microarchitectures. The 49 number shown in the OP seems off.
- phire 2mo agoIt's actually benchmarking 1000 repetitions of the RDTSC instruction running in parallel. My guess... On Skylake, multiple in-flight RDTSC instructions slow each other down for some reason? Possibly because it's attempting to provide a strict monotonic guarantee, that no two RSTSC instructions will return the same timestamp. Intel's manual only claims monotonic, which theoretically allows for two RSTSC instructions to return the same timestamp.
- IshKebab 2mo agoUsing MMIO is cheating and makes the results very boring. It would be much more interesting to know the results if you're only allowed to use main memory.
- Lockal 2mo agoI'm not sure how to properly measure DRAM access, but for register/immediate-only instructions there is https://uops.info/table.html https://uops.info/table.html. It does not measure x87 transcendentals (I guess because microcode results in different performance), but CPUID is a notable mention.
- markus_zhang 2mo agoDoes that mean Chris Domas is ready for his next adventure?
- monocasa 2mo agoIt says in the rules > Trapped/emulated/virtualized instructions may only time the trap, not the handler. But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.
- simonebrunozzi 2mo agoRelated, somehow: Core War [0]. [0]: https://en.wikipedia.org/wiki/Core_War https://en.wikipedia.org/wiki/Core_War
- eek2121 2mo agoThis is neat!
- baddash 2mo agojust curious, how much do these actually discover useful practices or pitfalls, on top of just being for fun?
- phire 2mo agoThe repo page says exactly why these long running instructions were researched. A very long-running instruction can be used to break SMI: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii
- darksim905 2mo agoSeems like spam from this creator since there are two things on the front page?
- john_strinlai 2mo agosubmitted by two different people, both with year+ old accounts and decent karma. i dont think either is the author. the other submitter probably read this one, looked at the github, saw something else cool and posted it. (i almost did the same, but bookmarked it instead)
- kazinator 2mo agoBus cycles can be arbitrarily long on any processor that has memory cycles with a hand shake requiring an ack, with no timeout. E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive. Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero. I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
- Joel_Mckay 2mo agoIn general, more complex processors have latency issues, and in some ways modern chips have actually become worse with each design iteration. https://en.wikipedia.org/wiki/Metastability_(electronics) https://en.wikipedia.org/wiki/Metastability_(electronics) Ultimately... failure modes must arise because naive gate design simulation models can't determine issues in a computationally feasible time frame. The consequences of "fixing" CDC prone design flaws makes a processor many times slower (8 to 16 times slower on my dumb attempt), and develops weird alien design features very different from Von Neumann architectures. This is why we can't have nice things. =3
- kazinator 2mo agoI smell Buridan's ass. :)
- inigyou 2mo agoBus cycles can also be arbitrarily long on those microprocessors if they don't use dynamic logic - you can stop the clock.
- NooneAtAll3 2mo ago> Times are normalized based on the CPU base clock frequency
- deleted 2mo ago[deleted]
- 56767865678 2mo agoSahil
- Retr0id 2mo agoI wonder if you can do some damage with scatter/gather ops within a VM, such that each fetch is a TLB miss inside the VM, and every table walk fetch is a TLB miss outside of the VM (which gets you up to 24 "fetches per fetch").
- amluto 2mo agoI believe that many (all?) x86 CPUs will cheerfully load page tables from MMIO space. And at least some of the paging formats let you set the UC memory type for page tables. (And don’t forget MTRRs.)
- userbinator 2mo agoPCIe is more like a packet-switched network than a bus, which is incidentally why things like Thunderbolt (effectively external PCIe) and sillier demonstrations like https://www.youtube.com/watch?v=q5xvwPa3r7M https://www.youtube.com/watch?v=q5xvwPa3r7M work. ...and with things like https://en.wikipedia.org/wiki/ExpEther https://en.wikipedia.org/wiki/ExpEther , you can get even higher latencies.
- inigyou 2mo agoPretty dumb right? When latency gets this high, you need a more asynchronous design to get any reasonable performance. PCIe is clearly designed with the assumption of latencies a few hundred cycles at most (or usually) - this MMIO register is an extreme outlier. It might be unmapped, and timing out on the hardware side, or it might be converted to an access on some really slow configuration bus.
- rurban 2mo ago62s for a single instruction! Wonder if an compilers cost tables knows that. But it's data dependent, and cost functions probably don't do that.
- inigyou 2mo agoCompilers don't even emit this instruction.
- Telaneo 2mo agoDomas manages to abuse x86 is ways that make me unsure whether or not I should be impressed or disgusted. I guess impressed, then disgusted over Intel (and AMD?).
- thyristan 2mo agoDepending on his interpretation of the rules about trapped instructions, one could just build a loop in the x86 page tables. Those are usually a tree linked by pointers, and any page table lookup can create another page fault that creates another lookup that... Leads to x86 page table MMU magic being turing complete: https://github.com/jbangert/trapcc https://github.com/jbangert/trapcc And the simplest thing you can do on such a system is just to loop indefinitely, thus creating a simple instruction with a memory access (mov or anything, doesn't really matter, even the instruction fetch for a nop would work) to take infinite time.
- inigyou 2mo agoPage tables are physically addressed, so can't recurse. I assume this thing actually works by causing a page fault on the first instruction of the page fault handler, which is a new instruction.
- thyristan 2mo ago> Page tables are physically addressed Nope. Not on x86. You can use either physical or virtual addresses at your choosing. Consumer OSes use virtual ones, so you can swap out page tables (yes, really!). See https://wiki.osdev.org/X86_Paging https://wiki.osdev.org/X86_Paging "Page directory".
- inigyou 2mo agoNothing in this section mentions them containing virtual addresses. In fact the word "physical" is written in bold. Are you a hallucinating LLM?
- thyristan 2mo agoYou are right. Not an LLM problem, just an undecaffeinated meat brain and some faulty memories. I've misread 'When PS=0, the page table address field represents the physical address of the page table that manages the four megabytes at that point.' to mean that when PS=1, the address isn't physical. But PS is page size... And I somehow remembered that you could induce pagefaults when walking the page tables... Sorry.
- hncsiocp9x 2mo ago[dead]
- hn9zmdcaou 2mo agoThis matches everything I've seen
- egberts1 2mo agoMy favorite is having an XOR opcode modifying the operand of its next IMUL operand. Fools QEMU and makes for a good "am I on a VM" logic test. A function of precalculating XOR operand inadvertly twice at QeMU TLB compute time AFTER retrieval of and toward its cached IMUL operand value. In short, emulation doing preparation of registers twice (negating XOR) Now you have a logic test revealing QEMU thru minute differential of IMUL operand value and its different multiplication results. ROT13, anyone? Disclaimer: works only on RXW memory page. It is literally a self-modifying code.