30 ms·
Surprisingly Slow
- h2odragon 6y agoI'll throw in "hidden network dependencies / name resolution"; it's amazing how things break nowadays when there's no net.
- verdverm 6y agoI'd add SaaS dependencies as well, whether it be slowness or downtime
- segmondy 6y agoThis is a solved problem tho, timeout / retry / circuit breakers / fallback etc. See - https://github.com/resilience4j/resilience4j https://github.com/resilience4j/resilience4j
- verdverm 6y agoI wouldn't call it solved if a downstream SaaS is down and my build still times out despite the aforementioned resiliency.
- segmondy 6y agoNumber one rule of distributed systems, "the network is not reliable"
- iudqnolq 6y agoFor years I thought sudo just had to take seconds to startup. Then one day I stumbled across the fact that this is caused by a missing entry in /etc/hosts. I still don't understand why this is necessary. https://serverfault.com/a/41820 https://serverfault.com/a/41820
- Flex247A 6y agoHere's a dumb question: doesn't slow software affect the environment significantly?
- verdverm 6y agodepends on why it is slow, is it CPU bound or waiting for I/O?
- deleted 6y ago[deleted]
- vanderZwan 6y agoTo give an oversimplified answer, that depends on how things are connected and synchronized. You know how bottlenecks are often explained using funnels in real life? Imagine two funnels, one directly connecting to another one. In that case, the effective rate of flow is determined by the funnel which has the smallest tip - the fact that the larger one lets through fluid quicker is irrelevant, as it is slowed down by the smaller one (we're obviously ignoring things like funnels potentially overflowing here, this is a simplified scenario). So then the bottleneck is easy to pin-point, but also it immediately becomes obvious that widening one funnel beyond the size of the other has no effect. Now let's change the scenario: I have a bucket of water that I tip into the first funnel, going into a second bucket. Once the second bucket is full I tip the second bucket into the second funnel, going into a third bucket. In this case both funnels have an impact on the time it takes to fill the second bucket, but the smaller one sill has the bigger impact. And then of course there is the parallel scenario, you can probably see how that plays out. In our software environment is a complicated mix of all of these, and it can be really hard to pin-point what is the most significant effect. The "buckets" and "funnels" translate to all kinds of things - CPU an I/O are the most obvious ones, but there's more to it than that. Also there's tons of side-effects that make this entire picture a lot more complicated than the simplified model I just described. The article gives quite a few examples, like thermal throttling, or branch prediction, but reality is even more depressing. Here's a great talk on why benchmarking is even harder than you think it is by Emery Berger: https://www.youtube.com/watch?v=koTf7u0v41o https://www.youtube.com/watch?v=koTf7u0v41o This, by the way, is also why micro-benchmarks can be meaningless in the larger context. For those kind of situations Coz is supposedly a better option (I've never had to optimize complicated situations like that). The talk I just linked goes into detail as to how it works and why it is better
- artursapek 6y agoGreat, insightful post
- chungy 6y ago> Historically, the Windows Command Prompt and the built-in Terminal.app on macOS were very slow at handling tons of output. A very old trick I remember on Windows, is to minimize command prompts if a lot of output was expected and would otherwise slow down the process. I don't know if it turned write operations into a no-op, or bypassed some slow GDI functions, but it had an extremely noticeable difference in performance.
- lifthrasiir 6y agoIIRC the font rendering in Windows was surprisingly slow and even somewhat unsafe. In one case lots of different webfonts displayed in the MSIE rendered the whole text rendering stack broken, with all text across the entire system disappeared. I wouldn't be surprised if this is a root cause of slow command prompts.
- patates 6y agoWell that explains the crazy slow performance on an old windows forms app I had written back when I was a junior developer. I'll try to see if I can reach anyone from my first employer and make them disable the debug output. Would be an interesting contribution, considering that I left more than 10 years ago :)
- ygra 6y agoOld old Windows Forms had its own text rendering (and still has), but most controls now have a second code path that uses the system's text renderer, which got updated with better shaping, more scripts, etc., while GDI+ basically never got any updates. You can see that when there's a call to SetCompatibleTextRendering(false) in the code somewhere; then it's using GDI instead of GDI+.
- gmueckl 6y agoThis sounds a lot like GDI resource exhaustion. It looks like the limits on handle IDs and the GDI heap size are still in place even in Windows 10.
- 6y ago
- Jakobeha 6y agoWould slow build configuration be a problem though? It isn't even slow compiling, on one machine you configure once and then you can compile n times (e.g. if you're developing) He's definitely right about writing to Terminals though, or in my experience logging.
- jstimpfle 6y agoIf you need to run the configuration only once, it's probably not a problem. It's surprisingly often though that we get into a situation where we unexpectedly need to run a task many times. Trivially, this happens when you need to make a change in the configuration and require feedback.
- yitchelle 6y agoEven on subsequent updates to the build configuration, I believe it checks if there environment has changed before running a build configuration. The assumption here is that the dependencies are setup correctly.
- sigotirandolas 6y agoOften I rerun the entire build from scratch when I'm uploading some nontrivial change and I want to really make sure what I'm uploading builds. Clean builds using containers or chroots will often need to rerun a configure step to make sure the build is really clean. You could argue for a "sufficiently smart cache", but if you manage to optimize the configure step to run in a small time without a cache, it's one moving part less to keep in one's mind.
- jonstewart 6y agoUnfortunately autotools is dumb. If you add a new source file to Makefile.am, you will have to run autoreconf to generate your Makefile anew. That’s fine, but it also regenerates configure, even if you didn’t touch configure.ac. AND THEN the Makefile sees that configure is new, and it reruns it. Autotools does not hear your screams.
- voiper1 6y agoWow, a ton of nitty gritty details I was not aware of!
- totololo 6y agoGreat content but please improve the contrast of your website <3
- amarshall 6y agoHuh? The foreground is #404040 and background is #F9F9F9. That’s 9.84:1 which is high enough for WCAG AAA. The background image is speckled, but even the lowest-brightness pixel is #E8F7FA which is 4.72:1, which is not AAA, but is AA.
- klapatsibalo 6y agoWhat are these ratios and the letter sequences?
- lifthrasiir 6y agoRatios are contrast ratios [1] while letter sequences are WCAG 2.0 conformance levels (AAA is the highest). [1] https://www.w3.org/TR/WCAG20/#contrast-ratiodef https://www.w3.org/TR/WCAG20/#contrast-ratiodef
- three14 6y agoI didn't notice the contrast problem until seeing the GP comment, but I suspect that speckled backgrounds are worse than a solid #E8F7FA because they mess with people's ability to do edge detection when trying to see the shapes of the letters in the text.
- indygreg2 6y agoI removed the background image and made the text blacker. Might take a force refresh to pick up the CSS file change. Is it good enough or are further tweaks needed? If more, my web design skills are mediocre, so actionable feedback would be appreciated.
- NavinF 6y agoLooks perfect to me after ctrl+shift+r. I'm glad you removed that background image, it was rather pointless.
- balloneij 6y agoWindow's slow thread spawn time is incredibly noticeable when you use Magit in Emacs. It runs a bunch of separate git commands to populate a detailed buffer. It's instantaneous on MacOS, but I have to sit and stare on Windows
- brabel 6y agoDo you mean *process* spawn time? From the article: > On Windows, assume a new process will take 10-30ms to spawn. On Linux, new processes (often via fork() + exec() will take single digit milliseconds to spawn, if that). > However, thread creation on Windows is very fast (~dozens of microseconds).
- murkt 6y agoYes, they clearly mean process spawn time: > It runs a bunch of separate git commands
- TeMPOraL 6y agoOne of many reasons why I prefer to run Emacs under WSL1 when on Windows. WSL1 has faster process start times. But then with git, there are other challenges. It took me a while to make Magit usable on our codebase (that for various reasons needs to be on the Windows side of the filesystem) - the main culprit were submodules, and someone's bright recommendation to configure git to query submodules when running git status. Here's the things I did to get Magit status on our large codebase to show in a reasonable time (around 1-2 seconds): - git config --global core.preloadindex true # This should be defaulted to true, but sometimes might not be; it ensures git operations parallelize looking at index. - git config --global gc.auto 256 # Reduce GC threshold; didn't do much in my case, but everyone recommends it in case of performance problems on Windows... - git config status.submoduleSummary false # This did the trick! It significantly cut down time to show status output. Unfortunately, it turned out that even with submoduleSummary=false, git status still checks if submodules are there, which impacts performance. On the command line, you can use --ignore-submodules argument to solve this, but for Magit, I didn't find an easy way to configure it (and didn't want to defadvice the function that builds the status buffer), so I ended up editing .git/config and adding "ignore = all" to every single submodule entry in that config. With this, finally, I get around ~1s for Magit status (and about 0.5s for raw git status). It only gets longer if I issue a git command against the same repo from Windows side - git detects the index isn't correct for the platform, and rebuilds it, which takes several seconds. Final note: if you want to check why Git is running slow on your end, set GIT_TRACE_PERFORMANCE to true before running your command[0], and you'll learn a lot. That's how I discovered submoduleSummary = false doesn't prevent git status from poking submodules. -- [0] - https://git-scm.com/docs/git https://git-scm.com/docs/git, ctrl+f GIT_TRACE_PERFORMANCE. Other values are 1, 2 (equivalent to true), or n, where n > 2, to output to a file descriptor instead of stderr.
- ulrikrasmussen 6y agoI'd like to see some numbers comparing "backwards compatible" x86_64 performance with "bleeding edge" x86_64. That was something I had never considered, but it seems obvious in hindsight that you cannot use any modern instruction sets if you want to retain binary compatibility with all x86_64 systems.
- saagarjha 6y agoI suspect the difference will not be very large, as compilers are fairly bad at automatic vectorization. Most of the places that could benefit reside in your libc anyways and those are certainly tuned to your processor model.
- rtpg 6y agoYour libc being tuned to your processor model seems extremely unlikely unless you’re compiling from source? I dunno, I only have one amd64 DVD of ubuntu
- saagarjha 6y agoIt has multiple implementations of functions and dynamically selects the right one based on CPU features.
- vanderZwan 6y agoI dunno, the "Trickle-Down Performance" section on Cosmopolitan describing the hand-tuned memcpy optimization gives me the impression that there might still be quite a lot of missed optimization opportunities there due to pessimistic assumptions about register use: https://justine.lol/cosmopolitan/index.html https://justine.lol/cosmopolitan/index.html
- saagarjha 6y agoI probably agree with Justine about instruction cache bloat for these functions, but I remain unconvinced that diverging from System-V is something worth its tradeoffs. The discussion for that would likely be lengthy and unrelated to this topic, as compiling with newer CPU features would likely make performance comparisons worse under her scheme as the vector registers are temporaries.
- secondcoming 6y ago> Laptops are highly susceptible to thermal throttling and aggressive power throttling to conserve battery. I hold the general opinion that laptops are just too variable to have reliable performance. Given the choice, I want CPU heavy workloads running in controlled and observed desktops or server environments. Hallelujah. Running microbenchmarks on laptops is generally pointless
- nottorp 6y agoExcept then he goes on and says you can't expect consistent performance across servers either :) Probably not as bad as laptop troubles, of course.
- the8472 6y agoMeasuring instruction counts instead of cpu cycles or wall time can help. Or you can temporarily disable CPU boost clocks to stay within the thermal envelope. You should do that anyway since boosting is another source of noise. On linux it's as easy as echo 1 > /sys/devices/system/cpu/intel_pstate/no_turbo or echo 0 > /sys/devices/system/cpu/cpufreq/boost
- raverbashing 6y agoYeah, autoconf/autotools are a mishmash of old tools and scripts put together. I still can't get my head around what it actually does when you do ./configure (probably conjure some 70's Unix daemon to make sure your machine is not some crazy variant with 25-bit addresses) and I tend to avoid it whenever possible
- citrin_ru 6y ago1. It is relatively easy to see what exactly configure is doing - it is logged to config.log and you can edit configure and add an echo or change something to troubleshoot failure. Troubleshooting cmake failures is much harder in my experience. 2. In many cases it is possible to reduce number of checks configure is doing: configure.ac files often contain unnecessary stuff, probably copied from some other project and kept just in case.
- raverbashing 6y agoOh I know what it is testing for, I just don't think a modern project needs to "checking for special C compiler options needed for large files" or check for the fsync command
- citrin_ru 6y agoMost of such check are performed because a software author put some macro into autoconf.ac, sometimes without good reason or without any thought at all. I see this such attitude in many different areas e. g. people copy-paste some config options into software configs from some outdated how-to or StackOverflow and OK with now knowing what a given option does and if it is relevant in given case. May be with autotools it happens more often than with cmake (though I've seen enough bad CMakeLists.txt too) because there is a myth that autotools is very hard to learn so developer don't even try to read documentation and just put some random stuff into their .ac/.am files until it works for them albeit slow.
- account42 6y ago
- fabian2k 6y agoThe python overhead is something I've noticed as well in a system that runs a lot of python scripts. Especially with a few more modules imported, the interpreter and module loading overhead can be quite significant for short running scripts. Numpy was particularly slow during imports, but I didn't see an easy way to fix this apart from removing it entirely. My impression was that it does a significant amount of work on module loading, without a way around it. I think the other side of "surprisingly slow" is that computers are generally very fast, and the things we tend to think of as the "real" work can often be faster than this kind of stuff that we don't think about that much.
- nullify88 6y agoI see this alot with Ansible. Its not particularly slow but running it places a bigger burden on laptop cpu and fans than I'd imagined.
- throwdbaaway 6y agoNoticed the same too. It is likely that we are both impacted by the very aggressive default of 1ms for `internal_poll_interval`: https://github.com/ceph/ceph-cm-ansible/pull/308 https://github.com/ceph/ceph-cm-ansible/pull/308
- uyoakaoma 6y agoFor those with issues reading the site https://outline.com/CyzVvN https://outline.com/CyzVvN
- andreyv 6y agoAutoconf can use a cache file to speed up tests: https://www.gnu.org/software/autoconf/manual/autoconf-2.60/html_node/Cache-Files.html https://www.gnu.org/software/autoconf/manual/autoconf-2.60/h...
- fctorial 6y agoNeat.
- peter_d_sherman 6y ago>"Closing File Handles on Windows Many years ago I was profiling Mercurial to help improve the working directory checkout speed on Windows, as users were observing that checkout times on Windows were much slower than on Linux, even on the same machine. I thought I could chalk this up to NTFS versus Linux filesystems or general kernel/OS level efficiency differences. What I actually learned was much more surprising. When I started profiling Mercurial on Windows, I observed that most I/O APIs were completing in a few dozen microseconds, maybe a single millisecond or two ever now and then. Windows/NTFS performance seemed great! Except for CloseHandle(). These calls were often taking 1-10+ milliseconds to complete. It seemed odd to me that file writes - even sustained file writes that were sufficient to blow past any write buffering capacity - were fast but closes slow. It was even more perplexing that CloseHandle() was slow even if you were using completion ports (i.e. async I/O). This behavior for completion ports was counter to what the MSDN documentation said should happen (the function should return immediately and its status can be retrieved later). While I didn't realize it at the time, the cause for this was/is Windows Defender. Windows Defender (and other anti-virus / scanning software) typically work on Windows by installing what's called a filesystem filter driver. This is a kernel driver that essentially hooks itself into the kernel and receives callbacks on I/O and filesystem events. It turns out the close file callback triggers scanning of written data. And this scanning appears to occur synchronously, blocking CloseHandle() from returning. This adds milliseconds of overhead." PDS: Observation: In an OS, if I/O (or more generally, API calls) are initially written to run and return quickly -- this doesn't mean that they won't degrade (for whatever reason), as the OS expands and/or underlying hardware changes, over time... For any OS writer, present or future, a key aspect of OS development is writing I/O (and API) performance tests, running them regularly, and immediately halting development to understand/fix the root cause -- if and when performance anomalies are detected... in large software systems, in large codebases, it's usually much harder to gain back performance several versions after performance has been lost (i.e., Browsers), than to be disciplined, constantly test performance, and halt development (and understand/fix the root cause) the instant any performance anomaly is detected...
- swiley 6y agoI've had the displeasure of using machines with Mcaffe software that installed a filesystem driver. It made the machine completely unusable for development and I'm shocked Microsoft thought making that the default configuration was reasonable.
- ziml77 6y agoIn my experience, third party antivirus software does a better job than Windows Defender when it comes to file open/close performance. I always disable Defender or replace it with something else specifically because of the performance impact when working with many tiny files.
- lelanthran 6y agoMaybe the third parties have handlers that don't block? There isn't any need to, after all - simply record the file details and return immediately. The virus scanner can always check that file later. In fact, it makes more sense to do it that way because if the same file changes multiple times, the scanner will only check it once. Just have to make a trade-off on the duration - wait too long before you check the written file and it may have already gone on to infect something else.
- thu2111 6y agoI think their assumption is that this would race with another program opening or executing the file that was just downloaded if there is any possible delay at all, and thus by the time the scanner reaches it it's already too late. I mean, Microsoft aren't dumb. Windows Defender is a competent AV product. If they're blocking on close there's probably a reason for it and they probably hate it. The trick with thread pooling file closes is one I'll stash in my brain for later: performance on Windows matters, especially as Win10 is getting more and more competitive vs macOS all the time.
- mlthoughts2018 6y ago> “ Programmers need to think long and hard about your process invocation model. Consider the use of fewer processes and/or consider alternative programming languages that don't have significant startup overhead if this could become a problem (anything that compiles down to assembly is usually fine).” This is backwards. It costs extra developer overhead and code overhead to write those invocations in an AOT compiled language. The trade off is usually that occasional minor slowness from the interpreted language pales in comparison to the develop-time slowness, fights with the compiler, and long term maintenance of more total code, so even though every run is a few milliseconds slower, adding up to hours of slowness over hundreds of thousands of runs, that speed savings would never realistically amortize the 20-40 hours of extra lost developer labor time up front, plus additional larger lost time to maintenance. People who say otherwise usually have a personal, parochial attachment to some specific “systems” language and always feel they personally could code it up just as fast (or, more laughably, even faster thanks to the compiler’s help) and they naively see it as frustration that other programmers don’t have the same level of command to render the develop-time trade off moot. Except that’s just hubris and ignores tons of factors that take “skill with particular systems language” out of the equation, ranging from “well good luck hiring only people who want to work like that” to “yeah, zero of the required domain specific libraries for this use case exist in anything besides Python.” This is a case where this speed optimization actually wastes time overall.
- Jtsummers 6y ago> This is a case where this speed optimization actually wastes time overall. That's too much of an absolute to be a good rule. If your heavyweight runtime is being launched 1000s of times to get a job done but you only do this once every few months, sure, don't do many optimizations, certainly don't worry about a rewrite in another language. The savings probably aren't worth it. If your heavyweight runtime is being launched 1000s of times to get a job done every day or multiple times a day, consider optimizing. Which may include changing the language. That's hardly controversial, this is the same thing we consider with every other programming task. Is X expensive in your language and do you have to do this frequently? Then minimize X or rewrite in a language that handles it better.
- latch 6y agoSo if I'm compiling PostgreSQL from source, should I be doing: export CFLAGS='-O3 -march=native' Before ./configure? Because if I don't, it's using -O2 without specifying an architecture.
- ok123456 6y agoOnly use -march=native if you're never-ever going to run the binary on another machine. This includes upgrading the processor or data rescue.
- ddalex 6y agoWhat does data at rest has to do with the cpu architecture ?
- ok123456 6y agomarch=native turns on all the features for that are available for your processor. So if, for example, you have a processor that supports AVX512 and another that doesn't, you'll get illegal instruction errors as soon as your other machine hits a region code it decided to optimize using those instructions. You won't get any real useful error messages when this happens since it's so low level. Unless you know to look for this you'll just be scratching your head going, "but it works on my machine?!?"
- dspillett 6y ago> CPUs have somewhat plateaued in their single core performance in the past decade In fact for many cases single core performance has dropped at a given relative price-point. Look at renting inexpensive (not bleeding edge) bare-metal servers: the brand new boxes often have a little less single-core performance than units a few years old, but have two, three, or four times the number of cores at a similar inflation/other adjusted cost. For most server workloads, at least where there is more than near-zero concurrency, adding more cores is far more effective than trying to make a single core go faster (up to a point - there are diminishing returns when shoving more and more cores into one machine, even for embarrassingly parallel workloads, due to other bottlenecks, unless using specialist kit for your task). It can be more power efficient too despite all the extra silicon - one of the reasons for the slight drop (rather than a plateau) in single core oomph is that a small drop in core speed (or a reduction in the complexity via pipeline depth and other features) can create a significant reduction in power consumption. Once you take into account modern CPUs being able to properly idle unused cores (so they aren't consuming more than a trickle of energy unless actively doing something) it becomes a bit of a no-brainer in many data-centre environments. There are exceptions to every rule of course - the power dynamic flips if you are running every core at full capacity most or all of the time (i.e. crypto mining).
- sandos 6y agoYes, yes, this is why I haven't bothered to upgrade my 2500K, although it is actually time now, since games apparently learnt how to use more than 1 core. I always went to some benchmark every year and saw single-core performance barely moving upwards.
- VHRanger 6y ago2500k is borderline of the sweet spot, but something like a 4790k, 5775c or 6700k can hold up 7 years later. That said, the very latest processors (AMD 5000 series, M1 apple silicon) are starting to make real gains in single threaded speed
- CoolGuySteve 6y agoI replaced my 2500K a couple years ago using the cheapest AMD components I could find. The main improvements were mostly in the chipset/motherboard: - The PCIe 2.0 lanes on the old CPU were throttling my NVMe drive to 1GB/sec transfer rates. - USB3 compatibility and USB power delivery were vastly more reliable. My old 2500K ASUS motherboard couldn't power a Lenovo VR headset for example, and plugging too many things into my USB hub would cause device dropouts. - Some improvement in either DDR4 memory bandwidth or latency fixed occasional loading stalls I'd see in games when transitioning to new areas. Even with the same GPU, before the upgrade games would lock up for about half a second sometimes and then go back to running in 60fps.
- OskarS 6y agoThe last section is really interesting. The author presents the following algorithm as the "obvious" fast way of doing diffing: 1. Split the input into lines. 2. Hash each line to facilitate fast line equivalence testing (comparing a u32 or u64 checksum is a ton faster than memcmp() or strcmp()). 3. Identity and exclude common prefix and suffix lines. 4. Feed remaining lines into diffing algorithm. This seems like a terrible way of finding the common prefix/suffix! Hashing each line isn't magically fast, you have to scan through each line to compute the hash. And unless you have a cryptographic hash (which would be slow as anything), you can get false positives, so you still have to compare the lines anyway. Like, a hash will tell you for sure that two lines are different, but not necessarily that they are the same: different strings can have the same hash. In a diff situation, the assumption here is that 99% of the times, the lines will be the same, only small parts of the file will change. So, in reality, the hashing solution does this: 1. Split the files into lines 2. Scan through each line of both files, generating the hashes 3. For each pair of lines, compare the hashes. For 99% of pairs of lines (where the hash matches), scan through them again to make sure that the lines actually match You're essentially replacing a strcmp() with a hash() + strcmp(). Compared to the naive way of just doing this: 1. Split the files into lines 2. For each pair line, strcmp() the lines once. Start from the beginning for the prefix, start from the end for the suffix, in each case, stop when you get to a mismatch That's so much faster! Generating hashes is not free! The hashes might be useful for the actual diffing algorithm (between the prefix/suffix) because it presumably has to do a lot more line comparing. But for finding common prefix/suffix, it seems like an awful way of doing it.
- ajuc 6y agoYou assumed the diff algorithm only compares each line against one other line. That's not true. You look at each line many times in these algorithms. Running time is O(n log n) or O(n^2) not O(n). So you generate N hashes and compare each hash against log N or N other hashes. So, for big enough data it should be faster.
- OskarS 6y agoNo, you misunderstand: he mentions a common optimization where before you run your diffing algorithm, you find the common line suffix/prefix for the file, and how that will improve performance (if you have a compact 5-line diff in a 10,000 line file, it's unnecessary to run the diffing algorithm over the whole thing). His point was that this suffix/prefix finding thing was surprisingly slow totally apart from the actual diffing. I was talking about that part, how hashing there is unnecessary. As I mentioned at the end of my comment, for the actual diffing algorithm, it's fine to hash away.
- ajuc 6y ago> Currently, many Linux distributions (including RHEL and Debian) have binary compatibility with the first x86_64 processor, the AMD K8, launched in 2003. [..] What this means is that by default, binaries provided by many Linux distributions won't contain instructions from modern Instruction Set Architectures (ISAs). No SSE4. No AVX. No AVX2. And more. (Well, technically binaries can contain newer instructions. But they likely won't be in default code paths and there will likely be run-time dispatching code to opt into using them.) I've used Gentoo (everything compiled for my exact processor) and Kubuntu (default binaries) on the same laptop a few years ago and the differences in perceived software speed was negligible.
- taeric 6y agoMy understanding is that the stdlib of the machine already figures out the faster code for the machine at run time. Such that, for most of the heavy stuff in many programs, it isn't that different. Granted, I actually do think I can notice the difference on some programs.
- CoolGuySteve 6y agoIt depends on the software. I've recompiled the R core with -march=native and -ftree-vectorize and gotten 20-30% performance improvements on large dataframe operations. If it were up to me, the R process would be a small shim that detects your CPU and then loads a .so that's compiled specifically for your architecture. The same improvements would probably be seen in video/image codecs, especially on Linux where browsers seem incredibly eager to disable hardware acceleration.
- deleted 6y ago[deleted]
- drewg123 6y agoHe singles out Windows for configure slowness, but MacOS is shamefully slow as well. I've seen configure run at least 2x as fast on the same machine booted into Linux or FreeBSD as compared to the MacOS that came on it.
- taeric 6y agoI can't but think some of these fall into premature territory. Configuring a build for the machine is relatively rarely on the critical path. And it is mostly tests before the build. As such, it needs to compare to the build with tests, which typically takes longer than just the build. Similarly, the concern on interpreter startup feels like being about one of the least noticed times on the system. :(
- vanderZwan 6y agoThe author appears to be mostly speaking from their experience of working at Mozilla, so I hope most of their claims are at least somewhat backed up by empirical (although possibly anecdotal) evidence
- taeric 6y agoFair. And I should have stated my main objection is that I don't find some of these surprising. Not that I don't agree they are slow. Many of these will remain slow because they are far from the critical path of most end user systems.
- commandlinefan 6y agoWell, ok, maybe, maybe not - but why NOT make everything as efficient as you can? He's gone to all the trouble to show you what to do, it's no additional effort on your part to apply it.
- taeric 6y agoIsn't this essentially Knuth's take? I don't disagree, but I suspect the critical path had moved dramatically. In large, the time to decide to use python takes far longer than using python. (And I dislike python...)
- sfink 6y agoI agree that usually the first bottleneck is the edit-rebuild cycle. But I think bottlenecks (plural) are a better way to view things than a single critical path. If my edit-rebuild cycle is fast enough that it doesn't bother me, then depending on what I am working on, I may quickly start noticing configure times. An accurate build system will reconfigure when any of a lot of different files are touched. (And inaccurate build systems waste a lot more developer time, just less evenly distributed and with more frustration involved.) So I reconfigure if I'm working on something relevant to the build system. I reconfigure when I pull down changes and rebase. I reconfigure when I'm switching between tasks (perhaps because my tests take a long time in CI), and in fact I won't switch tasks if reconfiguring takes too long. (And yes, I have multiple work trees, but my object directories can hit 20GB and it wastes even more time shuffling stuff around when I start running out of disk space for 4 work trees * 3-5 different configurations.) And interpreter startup bites you all over the place! If it adds half a second latency to my shell prompt, I won't use it. And Greg gave the math for things where you're restarting the interpreter a million times, which isn't uncommon in my experience. You should always profile. These things are heavily dependent on the type of stuff you work on. Et cetera. But my personal experience at several very different jobs says that these things do matter, a lot. Also, if you work at a moderately large place, the productivity loss from small inefficiencies is staggering. You have to look at the full picture to really see it properly; when things are slower, people don't just wait, they context switch and may never come back that day. Good for engagement numbers on HN, bad for productivity and flow.
- twiddlebits 6y agoHas anyone experienced slowness with the Linux page cache on a machine with tons of RAM? This seems especially problematic as NVME arrays approach the speeds of main memory. It's like using main memory to cache main memory, with the additional overhead of the linux page caching algorithm which seems to get worse in relation to the size of the RAM. (i noticed this in ubuntu version 20, kernel version 5.4.0)
- fudged71 6y agoSpeaking of thermal throttling on Macbooks, it's also worth pointing out that after 2 years the thermal paste on the CPU should be replaced, which is only a few dollars. I wish Apple made this a free maintenance along with removing internal dust.
- brundolf 6y agoThis is a fascinating set of shop-knowledge from someone who's clearly spent many years in a set of trenches that I hope I never have to. Great stuff.
- tomhallett 6y agoDoes anyone have an recommendations for interesting “lessons learned” type information like this blog article has?
- brundolf 6y ago> If you are running thousands of servers and your CPU load isn't coming from a JIT'ed language like Java (JITs can emit instructions for the machine they are running on... because they compile just in time), it might very well be worth compiling CPU heavy packages (and their dependencies of course) from source targeting a modern microarchitecture level so you don't leave the benefits of modern ISAs on the table. Interesting, I wonder how this has affected language benchmarks and/or overall perception between JITed languages and native languages