14 ms·
A Not-Called Function Can Cause a 5X Slowdown
- wmu 8y agoBruce's write-ups are just great, sheer joy of reading.
- drudru11 8y agoTrue - and also remind me why I will stay away from windows
- NicoJuicy 8y agoLol, because of bugs? I sense more bias and it's not bug related.
- rincebrain 8y agoI regret to inform you that, while I think the current Windows development methodology could use better testing (to put it mildly), things like [1][2] still crop up in other platforms. [1] - https://bugzilla.kernel.org/show_bug.cgi?id=201685 https://bugzilla.kernel.org/show_bug.cgi?id=201685 [2] - http://lkml.iu.edu/hypermail/linux/kernel/1811.2/01328.html http://lkml.iu.edu/hypermail/linux/kernel/1811.2/01328.html
- NullPrefix 8y agoDon't see anything wrong with [2]. Security by default, if you want to disable protection to get some performance on defective hardware, that should be an active choice
- rincebrain 8y agoThe problem(s) with it were that it was underspecified _how_ negative the workload impacts were for some cases, and that it ended up in stable branches default-on before getting reverted for the impact. As Linus remarked in the thread, the people who were paranoid about the security implications already went the BSD route and disabled SMT, while the people who don't worry about it that much suddenly get a nasty perf impact by default. To quote an Intel rep in the thread: "Using these tools much more surgically is fine, if a paranoid task wants it for example, or when you know you are doing a hard core security transition. But always on? Yikes." Yes, your system should default to being secure. But there's a sliding scale when deciding on security flaw mitigations of user disruption versus level of security given. In this case, the user disruption was medium-high, and the benefits did not outweigh that, so the intersection of the two factors was refined. If this were something every skiddy toolkit were leveraging to exfiltrate astonishing amounts of data, this might fall the other way. But right now, that's not where we are.
- Valmar 8y ago[2] was resolved quickly. A much revised version has landed recently. [1] is a strange bug, because the devs have consistently been unable to reproduce it, despite constantly looking over the issue. Users of ZFS have hit problems, also, suggesting that it is not an EXT4 bug, but a very subtle problem elsewhere in the block subsystem.
- rincebrain 8y ago[2] was resolved after it landed in stable release branches, which is a bit late for how much impact it had. [1] was, in fact, root-caused to a blk-mq bug. https://patchwork.kernel.org/patch/10712695/ https://patchwork.kernel.org/patch/10712695/
- PhasmaFelis 8y agoYou're only hearing about this issue because they fixed it. And you know very well that every OS has had problems that are just as dumb from time to time.
- fit2rule 8y agoThat may well be, but DLL's on Windows have been dumb for a long, long time. And they're still dumb.
- PhasmaFelis 8y agoHonest question, how does this particular problem differ from classic dependency hell? It sounds like exactly the same sort of issue I see on Linux blogs when e.g. a broken PNG display library causes crashes in a program that doesn't display PNGs (but uses a library that uses a library that does).
- fit2rule 8y agoIts correct that there are similarities, but the difference in this case is that its a system-wide policy enforcing the use of a system-provided DLL that is causing the issue (everywhere) as opposed to your scenario being more of a userspace issue that, an admittedly skilled, user can get themselves out of .. In your case, I'd fix it with some hot LD_PRELOAD action - I'm not sure how I'd endeavour to address this on a production Windows system, however.
- KIFulgore 8y agoI was lucky enough to have a class with him at DigiPen. He's an optimization genius.
- xtrapolate 8y ago> "The first fix was to avoid calling CommandLineToArgvW by manually parsing the command-line." > "The second fix was to delay load shell32.dll." If your build pipeline is continuously spawning processes all-over, to the point "delay loading" makes a significant difference - it's time to start re-evaluating the entire pipeline and the practices employed.
- wtallis 8y agoDo you know of a build system that can handle a source tree as large as an entire web browser without spawning a lot of processes? It's hard to tell what, if anything, you are recommending here. Pass thousands of files to a single compiler invocation? Ignore the problems and stop trying to make process creation and clean-up faster?
- im3w1l 8y ago> Pass thousands of files to a single compiler invocation? Sure. Or pass it a file with all the filenames. Or have the compiler work as a server that takes compilation requests over a socket. It's not like passing thousands of filenames between two processes is a deep unsolved problem.
- gpderetta 8y agoOr just spawn thousands of processes which has been done for the last 40 years without particular issues.
- mannykannot 8y agoSo, the solution to concurrency problems is to serialize everything?
- deleted 8y ago[deleted]
- fit2rule 8y ago"Concurrency is everything serialised, properly."
- gcb0 8y agotl;dr the method presence pulls in a dependency that runs slow buggy code. so click bait.
- saagarjha 8y agoI don't think that's quite right: the method presence pulls in a dependency, but the code is not buggy and the slow bit is the loading of the dependency itself.
- masklinn 8y agoTechnically it's not even the loading of the dependency but its unloading: cleaning up GDI objects[0] on process destruction has become much slower in W10AE and a system-global lock is held during that event. [0] some of which are created automatically when a specific dll is loaded, possibly transitively
- vanderZwan 8y agoThat sounds like there might be other programs with severe regressions out there
- pjc50 8y agoThis is a stupid dismissal of a well-written piece of investigative debugging. This article is the kind of thing I'd like to see more of on HN.
- jackewiehose 8y agoI also liked the article but to be fair, the title is actually a little clickbaity. The slowdown has nothing to do with not-called functions, it's just about DLL-dependencies.
- mark-r 8y agoNot clickbaity at all, it's the essence of what makes the problem interesting. The fact that a DLL dependency can slow down your program at shutdown is not at all intuitive, particularly when it's a system DLL that should be bullet-proof.
- markpapadakis 8y agoGreat read but what’s up with the 6 banner ads interspersed in the content page ? Maybe a few too many ?
- doombolt 8y agoMaybe it's mobile view or something? No banners in text block on my laptop. Some in the right pane, some down below.
- wmu 8y agoAdBlock might help.
- c0nfused 8y agoI get the congrats Android user redirect ads on mobile. So, it seems fairly over the top
- pwg 8y agoWith NoScript blocking execution of javascript, there are zero banner ads interspersed in the content page.
- mehrdadn 8y agoAlso good to know: I seem to recall similar if not worse slow-downs happen if you try to pull in Windows Sockets (WS2_32.dll).
- chii 8y agowow, how come the mere presence of a function causes a DLL to get loaded? Is it because in order to compile, the DLL (or its export definition) needs to be present, and the compiler does some magic because of that?
- _wmd 8y agoLazy binding isn't without downsides, it requires internal synchronization of its own, which means it's possible to write multithreaded programs that will suffer latencies due to lock contention during symbol resolution. Depending on OS (not sure this applies to Windows), it can also mean what used to be fatal startup errors are delayed long into process life
- TickleSteve 8y agoBecause it was linked in, extra GDI dlls were also linked in leading to GDI cleanup. The linker has no idea that the function is unreachable, it has to link it in to resolve the external symbols.
- polskibus 8y agoIs there a way to tell whether a .net program is also affected by this behavior?
- mark-r 8y agoI just checked a C# app I had laying around with depends, and it depends on user32.dll which depends on gdi32.dll. It's hard to imagine a Windows program that wouldn't depend on one of those critical DLLs. The only thing that saves us is that we don't often create and destroy hundreds of processes at a time.
- CoolGuySteve 8y agoWindows has a real problem with huge monolithic dlls that inexplicably pull in other huge monolithic dlls. Especially when each dllmain() has its own weird side effects. It leads to bizarre behavior all the time such as this. The weirdest thing that ever happened to me was when Visual Studio 2012 hung in the installer. After debugging with an older VS, it turned out the installer rendered some progress bar with Silverlight, which was hung on audio initialization, which was hung on a file system driver, which was hung on a shitty Apple-provided HFS driver. Uninstalling HFS fixed the installer. Why does an installer even need audio when it never played sound? Because Microsoft dependencies are fucked.
- agumonkey 8y agoIt's a mystery how MS manages to exist. Or maybe it's just consumer market emergent chaos mastery.
- userbinator 8y agoGetting even more to the point: why does an installer even need to use Silverlight to render a progress bar!?!? As far as I know, these things have been there since 16-bit Windows and continue to work fine in Win10: https://docs.microsoft.com/en-us/windows/desktop/controls/progress-bar-control https://docs.microsoft.com/en-us/windows/desktop/controls/pr...
- jcelerier 8y ago> continue to work fine in Win10: putting the head in the sand like this won't make the problem go away. UX designers demand stuff that is smoother than win32 progress bars, that has spline-interpolation-like behaviour with some nice little tween at the end when going to the next page, with varying smooth color gradients all over the place. go make this UI with pure Win32 primitives : https://youtu.be/v0GG8uh80V4?t=158 https://youtu.be/v0GG8uh80V4?t=158
- rkagerer 8y agoI hate animations like that. All they do is distract me, waste my time while I'm waiting to see and interact with the next thing, and move my target out from under me while I'm trying to go click it. I always turn off all the animation settings right out of the box. I wouldn't be surprised if one day soon this fad gives way back to a renaissance of clean, simple, instant UI's and everyone will feel like their devices got a lot snappier.
- userbinator 8y agoUnless the compiler can prove you're never going to run into that case, it can't remove the call, and because the call is an imported function it still has to create the import and have an entry in the IAT for it, so it needs to be resolved at load time. Not all that surprising IMHO.
- ascar 8y agoThe author even said "we immediately knew what to do", which is kinda the contrary of surprising. The interesting bit is not that a slow loading dependency got imported anyway, but why that dependency is slow and that it can get imported very easily indirectly.
- csours 8y agoGreat post, but it made me sad because of how people limit themselves when it comes to tests. Tests are the things you do after the code works, maybe if you have time. Managers generally only care about hitting a %. Colleagues dodge and avoid them. I don't remember any awards for tests or testers. But tests are so powerful and so cheap...
- brucedawson 8y agoGoogle generally has an excellent attitude towards tests, with most code reviewers refusing to approve changes that lack them. So that's great. In this particular case the person who reported and then fixed the issue was quite pleased when the LLVM tests became 5x faster (and didn't cause hangs) because they could run them far more frequently, thus catching bugs even earlier in the cycle.
- vmchale 8y ago> But tests are so powerful and so cheap... Most evidence suggests the opposite, though tests have become better in time.
- fhood 8y agoYeah, I firmly support testing, but if the code you are working on isn't explicitly written to make writing tests easier, than odds are (in my experience) that writing the tests will take longer than writing the code.
- csours 8y agoYes I think this is the problem. The attitudes I mentioned come from people who tried testing and got fed up because of the many legitimate issues they had.
- dylan604 8y agoI have been known to write shitty code, but am always looking to make the next set of code less shitty or even redo something I'm working on if the time allows. I'm getting better, but have a long way to go. I have a feeling that my code would be bad to write tests against. What would entail making code easier to test against? Small bits of code meant to be included that does one specific thing? Wrapping that included coded into functions and/or objects and then creating a test suite to hit all methods etc?
- rkagerer 8y agoThis reminds me of a desktop heap exhaustion problem IE would regularly trigger for me back in the XP days: https://weblogs.asp.net/kdente/148145 https://weblogs.asp.net/kdente/148145 It all came down to a registry setting that MS neglected to bump up much from the original Win 98 defaults. IIRC the conservative default even persisted into Win 7. That 3MB limit would bring down my 48GB system...
- IloveHN84 8y agoImagine if you use layers and layers of abstraction (e.g. Java lasagna programming style), how slow it could be
- gok 8y agoThe title is really misleading. "Dependencies cause overhead even when unneeded" would be more accurate.
- klodolph 8y agoThat strips out the interesting part, though. A 5X slowdown for an actual, useful, real-world project is interesting. Vague “overhead” could mean nothing more than some bigger binaries.
- pitterpatter 8y agoThis bug is fixed in the latest insider builds at least. Using the author's own testing tool: With the Spring 2018 release: F:\tmp>.\ProcessCreatetests.exe Main process pid is 46940. Testing with 1000 descendant processes. Process creation took 2.309 s (2.309 ms per process). Lock blocked for 0.003 s total, maximum was 0.000 s. Average block time was 0.000 s. Process termination starts now. Process destruction took 0.656 s (0.656 ms per process). Lock blocked for 0.001 s total, maximum was 0.000 s. Average block time was 0.000 s. Elapsed uptime is 7.08 days. Awake uptime is 7.08 days. F:\tmp>.\ProcessCreatetests.exe -user32 Main process pid is 44584. Testing with 1000 descendant processes with user32.dll loaded. Process creation took 2.624 s (2.624 ms per process). Lock blocked for 0.014 s total, maximum was 0.001 s. Average block time was 0.000 s. Process termination starts now. Process destruction took 1.617 s (1.617 ms per process). Lock blocked for 1.122 s total, maximum was 0.648 s. Average block time was 0.026 s. Elapsed uptime is 7.08 days. Awake uptime is 7.08 days. With an insider build: C:\tmp>.\ProcessCreatetests.exe Main process pid is 9928. Testing with 1000 descendant processes. Process creation took 2.440 s (2.440 ms per process). Lock blocked for 0.003 s total, maximum was 0.002 s. Average block time was 0.000 s. Process termination starts now. Process destruction took 1.306 s (1.306 ms per process). Lock blocked for 0.003 s total, maximum was 0.001 s. Average block time was 0.000 s. Elapsed uptime is 4.78 days. Awake uptime is 3.93 days. C:\tmp>.\ProcessCreatetests.exe -user32 Main process pid is 14144. Testing with 1000 descendant processes with user32.dll loaded. Process creation took 4.756 s (4.756 ms per process). Lock blocked for 0.022 s total, maximum was 0.004 s. Average block time was 0.000 s. Process termination starts now. Process destruction took 1.823 s (1.823 ms per process). Lock blocked for 0.003 s total, maximum was 0.001 s. Average block time was 0.000 s. Elapsed uptime is 4.78 days. Awake uptime is 3.93 days. There's no longer a difference in lock blocked time whether or not you load user32 during process destruction. Nor does the very obvious mouse stuttering still happen.
- brucedawson 8y agoWoah! That is fascinating. I had heard nothing about this. Your uptime is a bit shorter on the insider build but the change in lock blocking is too dramatic to be explained by that. I notice that all of the elapsed times are worse on the insider build - is that perhaps a slower machine? And are there enough CPUs on that machine to trigger the bug? That is, I'd like to believe that the bug is fixed but I'm skeptical.