6 ms·
The time the x86 emulator team found code so bad they fixed it during emulation
- canucker2016 4mo agoI was looking through the compiler docs about memory allocation and I found the section about the debug version of the CRT which could fill the allocated memory with a non-zero canary value to help detect uninitialized memory (assuming you weren't calling calloc - which zero-init's allocated memory). But there wasn't any similar programmatic debugging aid for detecting uninitialized stack memory. Going further down the rabbit hole, I discovered the _chkstk function. The MS C compiler would emit a call to _chkstk on function entry to ensure that stack memory had been paged in. But further reading noted that _chkstk was only emitted if the function allocated a lot of stack memory. And there was source code! MS included the assembly language source code for _chkstk in the CRT source code, installed with compiler. I needed _chkstk to be emitted for every function not only for functions that allocated >= 4KB of stack variables. Curses, foiled again. Then, while perusing the list of compiler command line switches, I see "/Ge". /Ge (Enable Stack Probes) Activates stack probes for every function call that requires storage for local variables. Ahhhhh! The grey, storm clouds parted and the sun rays bathed shone down on me in their warmth. I had all the pieces I needed to fill uninitialized stack memory with a non-zero canary value so I could make detection of uninitialized stack variables more reliable. _stkfil was born Modifying _chkstk was easy. I needed to write to every byte of stack in a stack page instead of reading only 4 bytes and skipping to the next page of stack. While I was mucking in the bowels of modifying _chkstk, I added a 4-byte global variable to hold my canary value. Let the app override what value to use. In debug builds, _stkfil helped find a couple of bugs, but soon all the stray uninited stack vars were gone and the code was forgotten. Then I read about InitAll in https://www.microsoft.com/en-us/msrc/blog/2020/05/solving-uninitialized-stack-memory-on-windows https://www.microsoft.com/en-us/msrc/blog/2020/05/solving-un... InitAll - Automatic Initialization In addition to the previously mentioned approaches, Microsoft is now using a feature known as InitAll which performs automatic compile-time initialization of stack variables. This section documents how Windows is using this technology and the rationale for why. Current Windows Settings The following types are automatically initialized: - Scalars (arrays, pointers, floats) - Arrays of pointers - Structures (plain-old-data structures) The following are not automatically initialized: - Volatile variables - Arrays of anything other than pointers (i.e. array of int, array of structures, etc.) - Classes that are not plain-old-data For optimized retail builds, the fill pattern is zero. For floats the fill pattern is 0.0. For CHK builds or developer builds (i.e. unoptimized retail builds), the fill pattern is 0xE2. For floats the fill pattern is 1.0.
- canucker2016 4mo agoAnd Android 11+ has been doing similar userspace stack-init thing - https://android-developers.googleblog.com/2020/06/system-hardening-in-android-11.html https://android-developers.googleblog.com/2020/06/system-har...
- hodgehog11 4mo agoI think we're starting to see more of this sort of thing happening now with Proton and Wine gaining prominence in the Linux community. Some games (Elden Ring comes to mind) have bad enough PC ports when they come out that the compatibility layer can incorporate a hotfix to improve performance, while users of the software on the original platform still had to suffer.
- Gigachad 4mo agoFairly sure GPU drivers do the same thing where they include a ton of per game tweaks to make them run faster. It does feel like a fragile way of doing things where an external component that should be agnostic to the software running ends up including a handful of junk trying to fix stuff that should have been fixed by the consumer of the driver.
- Guvante 4mo agoIt goes the other way too, sometimes you trigger some optimization silliness in the driver and the game needs to adapt to avoid it.
- rickdeckard 4mo agothen the driver gets updated and the game either continues to optimize (wrong) or branches out into code that was written before that driver came out and generally wasn't that well tested, and the circle continues... It's the life of a (game) developer...
- zoenolan 4mo agoThe big one I remember was many applications, not just games assuming the buffer swap was performed by a blit into the display buffer, not an framebuffer pointer update. They relied on the previous frames data still being in the back buffer. For those applications you were forced to blit the buffer, not swap the pointer and take a performance hit. I also remember a media player being called out by name in the code for doing invalid operations, needing a work around and code to detect it was running just to function.
- classichasclass 4mo agoBetting Alpha was the native architecture in question. It seemed to have the best support.
- projektfu 4mo agoYeah, but I thought DEC wrote the FX!32 translator for the Alpha. Perhaps Raymond was talking about those people and didn't want to mention that they weren't Microsoft people.
- m1r 4mo agoCouldn't they just turn the optimization off for this loop?
- MadnessASAP 4mo agoThey didn't have the code for the offensive program, they were creating the emulator to run it on a different architecture.
- McGlockenshire 4mo ago> offensive program Agreed.
- notorandit 4mo agoWhich optimizer replaces a 64k loop with 64k instructions? Ah, yes. Microsoft's!
- selcuka 4mo agoThere is no indication that the compiler that produced the code was Microsoft's. Actually the article hints otherwise ("[...] whatever compiler was used to compile this code").
- notorandit 4mo agoWho has been validating that approach to solve their own optimization target?
- notorandit 4mo ago> they fixed it during emulation It means the fix was applied to run during the emulation loop execution, not that the fix was found and applied while the emulation loop was running. Which would have made it an emulation code escape.
- jeffbee 4mo agoPeople from Transmeta told me stories about how their translators were full of special case optimizations to fix horrors they discovered in Microsoft Windows itself.
- yieldcrv 4mo ago> All in all, it took this program 256 kilobytes of code to initialize 64 kilobytes of data. solidity sweating profusely
- dlcarrier 4mo agoSimCity had a read-after-free bug that Microsoft patched in Windows 95. That was a lot easier for customers than having Maxis fix it, which could have required exchanging copies of the game.
- Cthulhu_ 4mo agoIt feels like graphics drivers do / did this a lot too. At the very least they make specific optimizations for specific games, probably by tweaking settings and features that the game developers didn't optimize properly themselves.
- kalleboo 4mo agoFamously if you renamed Quake 3 to "Quack" 3, it would slow down on the ATI Radeon 8500 https://web.archive.org/web/20091016055550/https://hardocp.com/article/2001/10/23/optimizing_or_cheating_radeon_8500_drivers/ https://web.archive.org/web/20091016055550/https://hardocp.c...
- account42 4mo agoThat's a case of the driver cheating but there are also lots of cases where the game is just full of bugs that the driver has to work around in order to not be blamed for them.
- Gibbon1 4mo agoI've said over the years a few times, this isn't our fault but it's our problem.
- smallstepforman 4mo agoThe driver switched to lower mipmapped texture and got caught. There is a ton of that out there for popular benchmark ready games. Run pro drivers instead of adrenaline to run generic baseline real driver.
- SyzygyRhythm 4mo ago
- electroglyph 4mo agoheh, when Raymond Chen dunks on the MSVC team =)
- mkl 4mo agoThere's no indication it was MSVC, and there are lots of compilers (and used to be more).
- selcuka 4mo agoTo be fair it is possible that the developer enabled a special "unroll all loops, no matter what" optimisation flag during compilation. I agree it would be stupid for a compiler to even support such a flag, but those were the 1980s/90s.
- cyberax 4mo agoAhh... Good old funrollloops... https://www.shlomifish.org/humour/by-others/funroll-loops/Gentoo-is-Rice.html https://www.shlomifish.org/humour/by-others/funroll-loops/Ge...
- PhilipRoman 4mo agoRight up there with fun, safe math optimizations
- account42 4mo agoAt least these actually make things faster usually.
- deleted 4mo ago[deleted]
- lozf 4mo agoHeh, "funrollloops" reminds me of recompiling FreeBSD 4 on my thinkpad back in the early aughts. The word made me imagine some sort of processed breakfast cereal with too many additives.
- ack_complete 4mo agoDoesn't require any special flags, just hitting optimizer limits can do it with MSVC. https://www.reddit.com/r/cpp/comments/1i36ahd/is_this_an_msvc_bug_or_am_i_doing_something_wrong/ https://www.reddit.com/r/cpp/comments/1i36ahd/is_this_an_msv...
- psanchez 4mo agoThis reminds me of a story from 15 years ago, where I was developing a technology to download games on demand by hooking into the OS calls. There was a particular game that was superslow when this tech was applied. Original game loading took around 15-20 seconds, whereas once the tech was applied it took easily 3-5 min, even with all data already downloaded. When I started digging into it, I realized the reason was the game was using something like fread(data, 1, 65536, fptr); instead of fread(data, 65536, 1, fptr); Which basically expanded back in the day to 65k reads of 1 byte for several MB file. Each fread translated to 65k reads of ReadFile Windows API. Since my code was hooking on ReadFile system call, and my call was heavier than ReadFile, the game loading felt really slow. Unusable. It would have not been fun for players. The easy fix was to swap arguments for certain calls. The long fix required to use an internal cache to account for these cases so that the hooked ReadFile was faster when data was already in disk. Funny thing is that as we started rolling out the tech and applying it to more and more games we realized lots of games did this. We went for the cache fix and games ended up loading faster than before. Honestly, games could have load all the data in a couple of seconds by just swapping the args. I'm guessing developers did this on purpose so that games seemed like they were loading a lot of stuff, although you never know.
- Taniwha 4mo agoI used to be a graphics card/chip architect for macs in the early/mid 90s - our chips were the fastest, but some programs were resistant because they did stupid stuff: pagemaker invalidated the font cache every time it went thru its main loop, quark with ATM did an n*2 thing every time it wrote text etc etc. We had special hardware to accelerate text drawing and it did nothing because the software pissed it away. We considered creating a plugin that fixed all these things, it would have been hard to maintain, in the end we travelled around to the people who made these apps and talked them through their problems To be fair excel would erase places white that it wanted to write up to 9 times before it drew any black pixels, we made that very fast! we didn't tell them :-) At the time 24-bit framebuffers were so slow that before we built graphics acceleration hardware people would switch back to 8-bit to get stuff done, making 24-bit/true colour your daily driver was a big step forward.
- kazinator 4mo ago> Anyway, my colleague found that there was one program that needed to allocate around 64KB of memory on the stack and initialize it. The standard way of doing this is to perform a stack probe to ensure that 64KB of memory is available, then subtracting 65536 from the stack pointer, and then initializing the memory in a small, tight loop. Actually, the standard way of allocating 64 kB of memory on the stack is to just assume you can do it, subtract 64k from the stack pointer, and hope for the best. Most stack allocations in the wild are not checked.
- i_don_t_know 4mo agoIIRC you have to probe every page of the stack on Windows. You cannot just subtract a value from ESP/RSP. If you don't probe every page in order, you get a page fault or some other exception (I don't remember which one).
- justsid 4mo agoHow else would the OS know your read/write 16 pages away from the current stack pointer is in fact an attempt to increase the stack and not just really bad pointer arithmetic and a bug? How many pages should the runtime let you skip before its just a segfault?
- NobodyNada 4mo agoThe reason for this is to ensure stack overflows are detected. The OS places a guard page above the top of the stack, which will cause a segfault if accessed. That way stack overflows are guaranteed to crash rather than stomping on valid memory that belongs to something else. However, if a stack frame is larger than a page (say, because it includes a large buffer), then it is possible for the program to "jump over" the guard page and access memory beyond. In order to protect against this, the compiler inserts some dummy reads or writes as needed to ensure every page is touched in order from bottom to top. This ensures the guard page is hit before the application has a chance to write to memory beyond it. Here's an example: https://godbolt.org/z/oTbzTczM6 https://godbolt.org/z/oTbzTczM6
- rohitsriram 4mo ago[flagged]
- ant6n 4mo agoArguably more of an optimization, rather than a fix. Looks like un-unrolling a loop, or better, rolling a loop. Or rolling straight line code?
- senfiaj 4mo agoYeah, but after a certain point the win is negligible. Huge code can also increase cache misses which will slow down things.
- deleted 4mo ago[deleted]
- ashdnazg 4mo agoI worked on a transpiler from Nand2tetris assembly to WebAssembly, and had some really annoying memory corruption bug that I just couldn't solve. That is, until I checked the program I used for testing (which I didn't write), and found the following code: dealloc(this) return this->field With the original allocator, this worked fine, since the deallocation didn't touch the memory. My allocator, however, overwrote the field during the deallocation with bookkeeping stuff, which meant the returned value was not what the programmer intended and after a short while the program crashed. Unlike TFA, I had the luxury of just fixing the test program.
- wazoox 4mo agoIIRC, one of the similar old story from Raymond Chen is about SimCity 2000, that did a similar trick (free memory, then start immediately using it) that worked just fine under DOS, but was a big no-no starting with Windows 95. The game was so common that Windows had to include a special rule to make it run...
- cranx 4mo agoLoop unrolling is a basic compiler optimization and depending on the machine language and processor instruction set should be faster taking into account all the house keeping required to execute a conditional, jump, move register values etc. This article is missing the analysis of why. If someone didn’t “like” it and was offended then that seems like an equally silly reason. On the surface 256k to init less does seem silly, but what if it was faster?
- ryukoposting 4mo agoA few things to consider. In this case we're talking about a tight initialization loop with probably a single instruction in the body. The HW optimizations necessary to make a loop like this perform equally to the unrolled form are so rudimentary that they're taken for granted on basically any CPU, even 30 years ago. Seriously, we're talking about optimizations I made in an "intro to Verilog" class as an undergrad, and I'm not even a HW engineer. It also depends how often this code is being hit. Does the code run once while the program loads? Nobody will notice a 2 microsecond improvement in loading times. Does the code run in a timing-sensitive hot path, like a game loop or a GUI rendering thread? Well now optimization matters. But again, consider the HW argument above. Also remember that, back then, storage wasn't cheap. 256K of code is 18% of a 1.44MB floppy, and 35% of a 720K floppy.
- pantulis 4mo agoI was just curious and checked The Old New Thing archive... yes I've been reading Raymond Chen's stories for as long as I remember but hey, it's been 23 years of delivering consistently solid stories about Windows.
- 0xdecrypt 4mo ago256 KB of code to zero 64 KB of memory is the kind of optimization that makes you question every life choice that led to it.
- rasz 4mo agoI blame Intel. It took them 33 years (ERMSB) to finally standardize REP MOVSB as _the_ fast path. Another 10 years passed and someone discovered https://lock.cmpxchg8b.com/reptar.html https://lock.cmpxchg8b.com/reptar.html
- zimmund 4mo agoI can't stop thinking about all the unoptimized code we have around. As processors (and memory) over the last 2-3 decades improved faster than we needed to fix the inefficiencies we created, we silently accepted that we don't need efficiency everywhere. So maybe a compiler, an emulator or some critical piece of code were created with this in mind, but the average app or website just waste resources left and right and pray for the best. With more and more code being written with AI (which has notoriously inefficient solutions to simple problems), I expect this issue to become more prevalent. I just hope we optimize at the source of the problem (AI and humans using it) and not on platforms (compiler and engine/kernel heuristics)
- smallstepforman 4mo agoHalf the compute and reduce memory by factor of x4 and in a decade we’ll have double the performance we have now. I do old school embedded, the amount of desktop bloat is insane. Any function I really need to refactor, I can reduce size and improve performance. And there are better engineers out there that are more efficient than me.
- phantasmat 4mo ago[dead]
- andikleen2 4mo agoDave Jones used to have a series of "Why user space sucks" Linux kernel conference talks with many such examples, usually with dumb and redundant system calls. However as someone who looks a lot at instruction traces I could probably write on e on why Linux kernel code sucks too. One of my current pet peeves is the way Linux walks bitmasks of CPU bits, which is a reasonably common operation. Due to a chain of unfortunate changes and decisions it currently needs 16+ instructions to find the next bit for something which the x86 instruction set has a single instruction. Of course that is so big that it is even outlined, adding even more overhead.
- zftnb666 4mo ago"It works" doesn't mean it works. It means it hasn't failed spectacularly yet.