9 ms·
Distcc: A fast, free distributed C/C++ compiler
- mkarliner 3y agoI've used it for remote raspberry pi compilation. Works very nicely.
- awestroke 3y agoWe used it at a previous job. Every dev computer in the building ran distcc, compiles were distributed and finished super fast.
- jahnu 3y agoAlso had much success with it in the past. In my previous job we wanted to use it when building Xcode projects with a large C++ library. However, it's not compatible with Apple's modified Clang. I spent some time maintaining a fork where I hacked in support for passing along the Apple specific switches but eventually gave up. Seems like a cool feature would be the ability to define non-standard switches in a config rather than have them hard coded. edit: just remembered I wrote up a blog post about it when I worked there https://pspdfkit.com/blog/2017/crazy-fast-builds-using-distcc/ https://pspdfkit.com/blog/2017/crazy-fast-builds-using-distc...
- anybodyz 3y agoFastbuild https://www.fastbuild.org/docs/home.html https://www.fastbuild.org/docs/home.html is the free distributed compilation system many game companies use. The combination of automatic unity builds (simply appending many .cpp source files together into very large combined files), caching and distributed compilation together gives you extremely fast C++ builds. Also supports creating project files that are compatible with XCode and Visual Studio so you can just build from those IDE's and a pretty flexible dependency based build file format that can accomodate any kind of dependency.
- rasjani 3y agoSomeone just linked Fastbuild few days ago in one community i follow and i did take a look as i havent heard about that project before. I didnt really like the idea that it would need to have its own build rules but its understandable - where as distcc would just integrate directly to what ever build tool you would be using. Also the syntax didnt look so "pleasing" but i guess the behefits can out weight the learning curve of yet an another build language (compared to meson/bazel/cmake et al) .. Having distributed builds on msvc thought with free software sounded very promising thought.
- anthk 3y agoGame companies didn't hear of ccache then.
- deleted 3y ago[deleted]
- sagarm 3y agoI really prefer icecream to distcc: it handles toolchain distribution, scheduling, and discovery.
- imp0cat 3y agoThis. If you're thinking about distcc, try icecream first, it is much nicer. > But unlike distcc, Icecream uses a central server that dynamically schedules the compile jobs to the fastest free server. https://github.com/icecc/icecream https://github.com/icecc/icecream
- ironbound 3y agoIs there a reason to uses this over something like bazel build?
- acatton 3y agoIt works transparently with any project using autotools/cmake/meson. (which is 95% of open source projects, and most likely most enterprise software) No need to convert your project to starlark BUILD files.
- pjc50 3y agoIs bazel build distributed? distcc lets you use a lot of machines to speed up the wallclock-time of your compilation.
- ReleaseCandidat 3y ago> Is bazel build distributed? Yes, that's the main feature of Bazel. And it caches the generated files. Although you could do theoretically cache using ccache with distcc.
- jeffbee 3y agoI doubt that it is the main feature. RBE is difficult to deploy and consequently rarely used. On the other hand of you assume that 99% of bazel invocations are done by google, then weighted that way RBE is nearly universal.
- DrBazza 3y agoIt can be. By default it is local. But it has protobufs interfaces (IIRC), so a distributed build farm would generate the grpc endpoints for their implementation and then you tell bazel on the command line (or via .bazelrc) the address of the build farm it can use. There's a couple of projects that implement the distributed/grpc part, the main one is https://github.com/bazelbuild/bazel-buildfarm https://github.com/bazelbuild/bazel-buildfarm
- sluongng 3y ago
- Alifatisk 3y agoHad no idea about this tech, very cool!
- bobnamob 3y agoSee also https://github.com/nelhage/llama https://github.com/nelhage/llama and https://github.com/StanfordSNR/gg https://github.com/StanfordSNR/gg
- hiyer 3y agoWe used to use this in a previous company. It reduced the build time from ~20 minutes to ~2 minutes; until a bug in the ld linker at the time (linking is not distributed - only compilation is) pushed it back up to 20 minutes. We had a hell of a time finding the issue, but luckily there was a newer versions of binutils available where it was fixed; and upgrading to it got the build time back under 2 minutes. Fun times!
- mansilladev 3y agoGood gracious. Just looked at CHANGELOG. It was 20 years ago this month that I made modest code contribution to this project.
- choeger 3y agoIIRC, when using distcc you must make sure that the Toolchain is ABI-perfect identified by its path. So: /opt/ourcompany/dev/bin/cc absolutely *needs* to be the same on all machines involved or you risk very hard to spot issues. For that reason, either use the absolute same distribution for everyone (and then run /use/bin/cc but watch out for alternatives!) or roll-out your own toolstack but make sure to put its version in the path.
- Erlangen 3y agoIsn't it a misnomer to call it a "compiler"? Even it's github README says otherwise, > distcc is not itself a compiler, but rather a front-end to the GNU C/C++ compiler (gcc), or another compiler of your choice. All the regular gcc options and features work as normal.
- nicce 3y agoI guess it depends. You use this software to convert your source code into the binary blobs efficiently. From high level, it behaves like a compiler, but internals are just using other compilers.
- Taywee 3y agoA lot of other compilers are the same way. gcc and g++ are higher level wrappers around their low level compiling and linking executables. It depends how you define the terms, including what you mean by "program".
- 0xffff2 3y ago> You use this software to convert your source code into the binary blobs efficiently. So `make` is also a compiler?
- justinsaccount 3y ago20+ years ago I used distcc + jpegtran to rotate a few thousand jpegs. From what I remember I just had to write the Makefile as you would normally. I think the end result was that the transfer time at 10mbps speeds made it take longer than it would have just transforming locally, but it was neat to see it work!
- dspillett 3y agoBack at Uni (two-and-a-half decades ago now), I setup a hacky version of this sort of thing to distribute compilations over several workstations: * use make -j to run multiple task at once * replace¹ gcc with a script that picked a host and ran the task on that² This had many limitations that distcc lists it doesn't have (the machines had to have the same everything installed, the input and output locations for gcc had to be on a shared filesystem mounted in the same place on each host, it only parallelised the gcc part not linking or anything else, and it broke if any make steps had side-effects like making environment changes, etc.) and was a little fragile (it was just a quick hack…), but it worked surprisingly well overall. For small processes like our Uni work it didn't make much difference over make -j on its own³ but for building some larger projects like libraries we were calling that were not part of the department's standard Linux build, it could significantly reduce build times. On the machines we had back then a 25% improvement⁴ could mean quite a saving in wall-clock time. One issue we ran into using it was that some makefiles were not well constructed in terms of dependency ordering, so would break⁵ with make -j even without my hack layered on top. They worked reliably for sequential processing which is presumably the only way the original author used them. ---- [1] well, not actually replace, but arrange for my script to be ahead of the real thing in the search path [2] via rlogin IIRC, this predates SSH (or at least my knowledge of it) [3] despite the single-core CPUs we had at the time, task concurrency still helped on a single host as IO bottlenecks meant that single core was less aggressively used without -j2 or -j3 [4] we sometimes measured noticeably more than that, but it depended on how patballisable the build actually was, and the shared IO resource was also shared by many other tasks over the network so that was a key bottleneck [5] sometimes intermittently, so it could be difficult to diagnose
- deleted 3y ago[deleted]
- tonetheman 3y ago[dead]
- ComputerGuru 3y ago> some makefiles were not well constructed in terms of dependency ordering, so would break⁵ with make -j even without my hack layered on top That remains true to this day of even makefiles from some really big name projects. Huge pet peeve of mine and I’ve filed PRs with fixes many times.
- torarnv 3y agoHow do these distributed build tools work across the Internet, i.e. outside of a local LAN office setting? Does the increased latency and slower bandwidth become a bottleneck?
- plq 3y agoIf you distribute preprocessed code, you need LAN speeds/latency for it to be faster than a local compilation. Otherwise, you need to have the exact same headers in addition to the exact same compiler installed everywhere. It MAY work otherwise, until it breaks in a subtle / obscure way one day and after tracking it down you will think maybe compilation doesn't take THAT long on a single machine :) So all in all, I wouldn't say works outside of a well-controlled cluster on a LAN.
- hbogert 3y agoAssuming you don't use NFS to continuously read and write, I think it would be fine. Compiling is generally speaking CPU bound. Even the website mentions that it just copies all preprocessed code to the workers, so I expect this to be a one-time overhead per job.
- sourcefrog 3y agoThey will work over the internet. You should use SSH or a VPN like Tailscale because the native protocol is not encrypted. Bandwidth and latency effects will have some effect on how well it will scale, but on a reasonably high latency and low RTT connection it's probably still faster than purely local compilation. Sending work over the network takes time but not much CPU, so the local CPU is freed up to do other work.
- _joel 3y agoUsed to use this with Gentoo when doing emerges. Reusing old P3s and stuff, those were the day.. Insert XKCD "Compiling..." meme :)
- ilyt 3y agoSame, whole office stringed together so I could compile gentoo on my 700MHz laptop in reasonable time.
- RustyRussell 3y agoWow, it's still active! I remember when MBP wrote it while we were working at OzLabs together. One day I'll get back to that rewrite of ccontrol using modern distcc's features...
- sourcefrog 3y agoHi Rusty! For a later hack in the same vein, check out https://github.com/sourcefrog/cargo-mutants https://github.com/sourcefrog/cargo-mutants
- squarefoot 3y agoGood memories! I had fun with distcc by compiling kernels across a few local machines back when desktops for mere mortals were dog slow, and it helped a lot. I never used it for cross compiling though, which is something that could help today when starting compilations from small embedded boards. Did anyone have success in a mixed environment, such as a small ARM board with native GCC plus one or more faster x86 machines with cross compiling tools installed?
- vkoskiv 3y agoJust tried this. I have a slow ARM laptop and a fast x86 desktop. The ARM laptop seems to detect the desktop as specified in the environment variable, but it doesn't accelerate compiling at all. The x86 box sits idle while the ARM laptop compiles. I have the aarch64 cross compiler installed on the desktop as well. Maybe I'm just bad at RTFM, but I can't seem to get it to work.
- mattst88 3y agoI don't have enough details to debug, but something I discovered just the other day might be helpful if you're building the kernel. For kernel builds, you have to specify CC= after `make`, not before. E.g. `make -jX CC="distcc alpha-unknown-linux-gnu-gcc"`.
- iveqy 3y agoI'm curious about the security implications with using distcc. Doesn't this mean that if one computer gets compromised, the attacker can run code on all other computers using distcc, or secretly inject malicious code in the build result. So using distcc means that all computers using it must be trusted. And that means that using it on "all developers computers to share the load" is good for performance but bad for security.
- pjc50 3y agoEverything on the same LAN should generally be treated as "compromised/not compromised" together. There's rarely just a compromise of one machine in the same way there's never just one cockroach. I'm not sure whether distcc affects reproducible builds? You could, in any case, have tighter controls on the release builds, which would be done on a CI machine before signing. (Back when I used distcc we didn't distribute across the dev machines, we had an entire build farm of two racks of 2U servers!)
- nicce 3y ago> Everything on the same LAN should generally be treated as "compromised/not compromised" together. There's rarely just a compromise of one machine in the same way there's never just one cockroach. I would not generalise so quickly. Is every computer compromised in the internet if one is compromised? No. It highly depends on the trust between those machines and whether they share similar services with critical vulnerabilities. Only then, they might be compromised together. But the world has evolved and not everyone anymore bases their total trust and security thinking for "no outside internet connection, we are fine".
- meinheld111 3y ago> Is every computer compromised in the internet if one is compromised? Internet is no lan (local area, L2) where computers do indeed typically have more generous policy regarding access between each other. Think about the windows firewall asking whether you just connected to a work/public/home network
- 3y ago
- anonymousDan 3y agoCan anyone explain the architecture/how it works at a high level? I get that it is distributed. Does it basically copy the complete source tree to every worker and have them compile some independent subset of the object files? Does performance scale linearly with the number of worker nodes?
- donaldihunter 3y agoPer the description: > distcc sends the complete preprocessed source code across the network for each job, so all it requires of the volunteer machines is that they be running the distccd daemon, and that they have an appropriate compiler installed. So all the "environment" is on the source machine and just a bare compiler is required on the remote machines for compilation.
- pjc50 3y agoLast time I looked, it basically ran per-file "cc -E" on the source machine to get a compilation unit (optionally checking for a ccache cache hit at this point), then piped the result to "cc" running on the target machine, and copied the resulting object file back. > Does performance scale linearly with the number of worker nodes Yes, for small N. Overall scaling was limited by how much "make -j" the source machine could cope with.
- sourcefrog 3y agoThis was the original approach, although later work added an optional "pump mode" in which headers are distributed: https://manpages.ubuntu.com/manpages/bionic/man1/distcc-pump.1.html https://manpages.ubuntu.com/manpages/bionic/man1/distcc-pump...
- Someone 3y agoIt runs the preprocessor locally, then sends that out to a volunteer node. https://www.distcc.org/distcc-lca-2004.html https://www.distcc.org/distcc-lca-2004.html: “The client is invoked as a wrapper around the compiler by Make. Because distcc is invoked in place of gcc, it needs to understand every pattern of command line invocations. If the arguments are such that the compilation can be run remotely, distcc forms two new sets of arguments, one to run the preprocessor locally, and one to run the. compiler remotely. If the arguments are not understood by distcc, it takes the safe default of running the command locally. Options that read or write additional local files such assembly listings or profiler tables are run locally” Scalability: “Reports from users indicate, distcc is nearly linearly scalable for small numbers of CPUs. Compiling across three identical machines is typically 2.5 to 2.8 times faster than local compilation. Builds across sixteen machines have been reported at over ten times faster than a local builds. These numbers include the overhead of distcc and Make, and the time for non-parallel or non-distributed tasks”
- IanCal 3y agoI used this in a uni lab (good lord 15 years ago) when we needed to compile for a robot. It took maybe an hour or two to compile run directly on the robot but I was able to setup distcc to run on all the lab machines in that room and get things done fast.
- donaldihunter 3y ago25+ years ago, our company used Clearcase for version control and it's clearmake had distributed build capability. Clearcase used a multi version file system (MVFS) and had build auditing so clearmake knew exactly what versions of source files were used in each build step. It could distribute build requests to any machine that could render the same "view" of the FS. Even without distributed builds, clearmake could re-use .o files built by other people if the input dependencies were identical. On a large multi-person project this meant that you would typically only need to build a very small percentage of the code base and the rest would be "winked in". If you wanted to force a full build, you could farm it out across a dozen machines and get circa 10x speedup. Clearcase lost the edge with the arrival of cheaper disk and multi-core CPUs. I'd say set the gold standard for version control and dependency tracking and nothing today comes close to it.
- oretoz 3y agoBrings back memories. Fresh out of college, I was given the additional job of being the Clearcase and Unix admin for my team. Not that I had any special skills but others didn't know a few Unix commands (System V) that I did. But Clearcase was such a good product and was used in the Telecom companies that I worked for (Motorola, Lucent etc.) It was owned by Rational at that time and if memory serves me right, were acquired by IBM. To this day, I find Clearcase's way of doing things is the better way to do version control. Git, in comparison, kind of feels alien and I could never really get the same type of comfort on it.
- pjmlp 3y agoThat was also the standard development workflow at Nokia Networks, back when NetAct was being developed for HP-UX. Nowadays most of it has been ported into Java, if the ongoing efforts when I left, finally managed to migrate everything away from C++, Perl and CORBA.
- z29LiTp5qUC30n 3y agoClearcase was utter crap. 6 hour code checkouts and 2 weeks to setup a new developer is a freaking joke. I literally did a conversion from Clearcase to git and reduced the setup time to 15 minutes and this is for a code base older than Clearcase is. Not to mention the absolutely bad design for handling merge conflicts (punt to human if more than 1 person touched a file seriously???)
- jchw 3y agoRelated: https://github.com/icecc/icecream https://github.com/icecc/icecream - another option that does what distcc does, but aimed at a somewhat different use case. https://ccache.dev/ https://ccache.dev/ - a similar idea but provides caching of build outputs instead of distributing builds. You can use it together with distcc to achieve even better performance.
- oretoz 3y agoccache is used together with distcc at the current place I am working at. Started digging at how these two work as I thought there is still room for improvement in our build times that can vary between 10 minutes to 1 hour. It is a huge code base, easily more than a million lines and around 18k files. But had to stop as there were way too many features to develop and bugs to fix. Also, management does not see that kind of work as useful so no point fighting those battles.
- thejosh 3y agoYou are probably aware, but for. others with ccache this is called "cache sloppiness", which is my favourite term. You can set this via config, as by default ccache is paranoid about being correct. But you can tweak it with things like setting a build directory home (this is great for me, as I'm the only user but compile things in say `/home/josh/dev/foo` and `/home/josh/dev/bar` and have my build directory as my dev directory and it's shared. (see https://ccache.dev/manual/latest.html https://ccache.dev/manual/latest.html for all the wonderful knobs you can turn and tweak). Fantastic tool, the compression with zstd is fantastic as well. I played with distcc (as I have a "homelab" with a couple of higher end consumer desktops), but found it not worth my time as compiling locally was faster. I'm sure with much bigger code bases (like yours) this would be great. Reason I used it it that archlinux makes it super easy to use with makepkg (their build tool script that helps to build packages).
- bombcar 3y agoThe best use of distcc "at home" is when you have one or more "big iron" (desktop, server, whatever) and a few tiny machines that work just fine but don't have much processing power. For example, with some work, you can setup distcc to cross-compile on your amd64 massive box for your raspberry pi.
- oarfish 3y agoThe biggest problem with distcc is that it fails with architecture-specific instruction sets like avx. A problem not shared by icecc.
- jeffbee 3y agoIsn’t that just a matter of setting appropriate target flags, and not letting them by implied? That’s basic build hygiene anyway.
- egwynn 3y agoXcode used to have native support for this, long ago. And it had mdns support too!
- ihaveajob 3y agoSo many memories. Back in grad school I would run this on a handful of workstations to speed up compilations. Really neat tool, and a clever setup.
- londons_explore 3y ago> distcc sends the complete preprocessed source code across the network for each job, s I assume that for template-heavy c++, that could easily be hundreds of gigabytes for 1 gigabyte of c++ code to compile...? If you're working from a laptop, surely the wifi connection will by far be the bottleneck?
- maccard 3y agoIt is. These systems (inctedibuild, fastbuild) tend to work extremely well but are network bottlenecked. You need a fast, low latency connection.
- sourcefrog 3y agoHi, distcc's original author here. It's really nice that people are still enjoying and using it 20 years later. I have a new project that is in a somewhat similar space of wrapping compilers: https://github.com/sourcefrog/cargo-mutants https://github.com/sourcefrog/cargo-mutants, a mutation testing tool for Rust.
- saulrh 3y agoPretty sure I used cargo-mutants on a lark during Advent of Code a couple years ago. Caught a couple bugs with it. Good stuff. Thanks! :D
- dzidol 3y agoWow, it's already 20 years. Was setting up distcc in my first job, with ccache for armcc. Good times, great experience. This was a life-safer in terms of compilation speed (just needed wait 20 mins for consolidation and like the same or more for loading symbols to the debugger, which was bearable comparing to spending another 2 or 3 hours taken by compilation). Having the opportunity, would like to thank you!
- plq 3y agoHey, I just would like to let you know that DistCC always has a special place in the hearts of loyal Gentoo users like me :) Kudos!
- omerhj 3y agoNearly twenty years ago I had a little server farm of old PCs. Two or three Pentium-133s, one dual Pentium Pro 200 machine, and my pride and joy, a Pentium 3 running at 600 MHz. I was trying to get familiar with Gentoo and to make recompiling everything all the time more bearable I set up distcc so my P3 could do most of the work. It worked very well! But after a few weeks every Gentoo box in the house started crashing regularly. It took me a while to figure out what was going on: one of the slower machines had developed a single-bit memory error and was sharing corrupted .so files with all other machines.
- klodolph 3y agoYeah. Everyone I’ve talked to who has run a build cluster has recommended ECC for the build cluster, even if they’ve decided not to use ECC for other systems. Some people would run Hackintosh-like setups for macOS build clusters, just for the ECC. Reproducible builds are also a big win here.
- omerhj 3y agoAbsolutely! This was a very educational experience.
- bombcar 3y agoI still need to get around to setting up distcc; I only have two Gentoo servers, but one is so much more powerful than the other, and their CPUs are close enough I might be able to use ccache, too ...
- ComputerGuru 3y ago> their CPUs are close enough I might be able to use ccache, too It’s the compiler that needs to line up for that. But my recommendation is to install sccache which will figure it out for you.
- Fatnino 3y agoI had to break into a laptop for my dad last week. Last used 7 years ago. Running win2k on a pentium 3. Had to jump through various hoops and find creative ways around problems (laptop too old to boot from USB, all my modern machines lack optical drives. And so on) but I did eventually break in.
- Cieric 3y agoWhile it's purpose is different it can be used to do distributed compiling, so I'll leave it here. https://github.com/Overv/outrun https://github.com/Overv/outrun Since I was just going down this rabbit hole recently, I kind of wonder if it's possible to set the filesystem on something more like the BitTorrent protocol so things like the libraries/compilers/headers that are used during compilation dont all need to come from the main pc. It probably wouldn't be useful until you reached a stupid number of computers and you started reaching the limits of the Ethernet wire, but for something stupid that can run on a pi cluster it would be a fun project.
- timetraveller26 3y agoThis is such a nice tool, I use it in my arch linux machines, it even supports cross compiling!
- ggerules 3y agoSome other multi machine options that have worked well for me, well beyond just compilation of C/C++ on multiple machines, with multiple gpu(s) and cores. 1) set up passwordless, ssh. and 2) use the gnu parallel. https://www.gnu.org/software/parallel/ https://www.gnu.org/software/parallel/ gnu parallel is super flexible, very useful.
- idatum 3y agoUsed distcc back in mid-2000s to crosscompile NetBSD kernel/userland for a TS-7200 SBC ARM device. Back then those devices didn't have much processing power, would have taken days otherwise.
- jdlyga 3y agoIt's a shame that distcc and ccache aren't more well known. Ccache, in particular, saved me probably years worth of compile time working with Qt.
- Avlin67 3y agostorage becomes much faster than network today, so is it still worth it ? (naive question sorry) or maybe the cpu demand is much bigger than IO demand ?
- linhns 3y agoTitle may be a bit confusing since skimming the project's homepage reveals that it's not a compiler, just frontend for other C++ compilers.
- lxe 3y agoOoh this brings back memories. This project made my embedded cross-compilation iteration speed go from like 2-30 minutes to 5. This was 20 years ago. I don't understand why this isn't more of a thing these days. I think Bazel can do this, but I've felt nothing but pain from using Bazel.