7 ms·
Author here. If you have any questions about mold/macOS, feel free to ask me as a reply.
by rui314 4y ago
Author here. If you have any questions about mold/macOS, feel free to ask me as a reply.
- comex 4y agoCan mold also serve as a replacement for dsymutil (i.e. linking DWARF info)? If not, any plans to support this in the future? :)
- rui314 4y agoI haven't thought about that that much, but it looks like dsymutil is a companion command of a linker, so the answer is probably yes, but I need to investigate it further to give you a concrete answer.
- glandium 4y agoSomeone was supposed to make llvm-dsymutil use the lld code, but that didn't happen.
- azinman2 4y agoWhy is it so much faster?
- rui314 4y agoThe biggest reason is because it is multi-threaded. When building a program, the compilation step is parallelized (the build system invokes a compiler for each source file), but the final link step is not. So it is important to make the linker itself multi-threaded. But even without multi-threading, mold is still faster than other linkers. I can think of various reasons why, but I don't know which attributes how much. I believe the biggest contributor is its efficient data structure -- it is hard to make program faster by writing fast code, but it can naturally be achieved by designing efficient data structures. That said, it is hard to compare two or more programs to find out why one program is faster than the others unless their designs are similar.
- glandium 4y agoNote that recent lld is also multithreaded, and with tweaks, can be faster than mold at low core counts: https://bugzilla.mozilla.org/show_bug.cgi?id=1746462#c2 https://bugzilla.mozilla.org/show_bug.cgi?id=1746462#c2
- rui314 4y agoI wouldn't be surprised. In mold, if we have more than one choices to implement a feature, I always take the one that scales well for more cores even if it doesn't perform the best on low-core count machines. My assumption is that future machines will have more cores than we have today on average, so I'm optimizing mold for such computers.
- petr_tik 4y ago> My assumption is that future machines will have more cores than we have today on average, so I'm optimizing mold for such computers. From the manpage of mold-1.3.0, I get the impression mold is designed to scale up to 32 cores, but not more. man mold | rg -C2 32 --threads --no-threads Use multiple threads. By default, mold uses as many threads as the number of cores or 32, whichever is the smallest. The reason why it is capped to 32 is because mold doesn't scale well beyond that point. To use only one thread, pass --no-threads or --thread-count=1. Is this correct or does the manpage need to be updated?
- bertr4nd 4y agoCould you comment on which data structures are most critical to mold’s performance, and what makes them so fast?
- bertr4nd 4y agoAh, I overlooked this explanation at first: https://github.com/rui314/mold/blob/main/docs/design.md https://github.com/rui314/mold/blob/main/docs/design.md Thanks for the great write up as well as mold itself!
- aidenn0 4y agohttps://github.com/rui314/mold#why-is-mold-so-fast https://github.com/rui314/mold#why-is-mold-so-fast
- wyldfire 4y agoI see mold does not yet support LTO [1]. But if it were to, could we still expect performance gain over lld when doing LTO? [1] https://github.com/rui314/mold/issues/181 https://github.com/rui314/mold/issues/181
- rui314 4y agomold does support LTO. It's not faster or slower than lld both in terms of link speed and the output binary's speed.
- Randor 4y agoDo you have any plans for supporting other binary executable file formats?
- rui314 4y agoWe do support ELF (Unix) already, and we are working on Mach-O (macOS/iOS/watchOS/etc) now. Once Mach-O is finished, we'll be working on PE/COFF (Windows).
- dholm 4y agoDoes mold support jobserver or some other mechanism to throttle threads? We use lld currently but had to disable threading as sometimes in CI several of our test binaries would get linked at the same time. Whenever this happened the lld instances appear to have spawned enough threads to overload our Jenkins slave to the point that the master wasn't able to reach it and failed the build.
- rui314 4y agoIt's being discussed (https://github.com/rui314/mold/issues/117 https://github.com/rui314/mold/issues/117) but haven't reached any conclusion. The problem is that the jobserver protocol assumes that one process is one job, and its model doesn't fit very well to programs such as mold.
- dholm 4y agoThank you rui314. There is a lot of good information collected in that issue.
- witcher_rat 4y agoIf you use Ninja, you can create a job pool for linking, separate from compiling, and restrict how many simultaneous linker jobs are run. You can even create and set job pools for Ninja through CMake, if you use that in your tooling. Unfortunately GNU Make offers no such mechanism. And for Ninja build generation I don't think Meson does it either.
- dholm 4y agoI was not aware of this feature and we do use CMake. Thank you for letting me know about it!
- rizzaxc 4y agodo you have a plan to supersede official linkers in gcc/ llvm?
- rui314 4y agoI don't have a plan, and that's not what I can plan. They can plan in theory, but I believe that's very unlikely to happen.
- rizzaxc 4y agoif mold has multiple advantages over the official ones and no drawback, can't you send them an RFC once mold reaches 1.0?
- rui314 4y agoNone of gold, lld or mold can replace GNU ld entirely because they don't cover all features that GNU ld has.
- 5e92cb50239222b 4y agoRead the readme. Full compatibility with gold/lld is a non-goal and will significantly slow down the linker IIUC.
- rurban 4y agothe whole embedded world uses linker scripts. for sure not mold goal. also, they don't have linker performance problems, but the big C++ apps have.
- ismaildonmez 4y agoIt might be asking for too much, but this project would be a nice addition to http://aosabook.org http://aosabook.org Thanks for your good work! I still remember the O(n^2) complexity of ld.bfd when linking C++ code.
- rui314 4y agoI want to write a book about linkers so that the knowledge I earned during the development of the lld and mold linkers wouldn't lost, but I don't have enough time to do that!
- kiru_io 4y ago> I want to write a book about linkers so that the knowledge I earned during the development of the lld and mold linkers wouldn't lost, but I don't have enough time to do that! I know this is a type, but it is interesting to see knowledge as a score or commodity you can "earn".
- epilys 4y agoI'm one data point but I'd buy it in an instant.
- pbiggar 4y agoFYI: https://www.amazon.com/Linkers-Loaders-John-R-Levine/dp/1558604960 https://www.amazon.com/Linkers-Loaders-John-R-Levine/dp/1558...
- pdimitar 4y agoYour tweet shows quicker linking on a Mac yet you say down-thread that Mac executable linking is still not fully supported. So can we or can we not use it today on a Mac? (My main use-case is Rust.)
- rui314 4y agoIt can create Mac executables, but mold/macOS is still in pre-alpha and no one should expect it to work for their programs. Once it becomes out of beta, I'll release it as mold 2.0, so please wait for it.
- pdimitar 4y agoThanks. I will wait for an official announcement then. Just subscribed to release notifications on GitHub, too.
- Jyaif 4y agoWhen are you going to start working on a 10x faster clang++? :-)
- rui314 4y agoThat's an even crazier goal which is probably 100x harder than writing a 10x faster linker. But I believe it's technically doable. At least, the world needs more crazy people who believe it is technically doable and take it as a challenge. If I get $$$ by selling the mold project to a big tech, I might be able to create a team with that money to tackle that crazy goal...
- blinkingled 4y agoMore power (and moneys) to you sir - I just used mold 1.3 to link qtwebkit right before this article appeared.
- petr_tik 4y agoLooking further into the future and the advent of io_uring in Linux, would you consider special-casing Linux IO ops to use io_uring or do you not expect any speedup there?
- rui314 4y agoI'm not sure if io_uring can improve mold's performance, as it has to access random locations while copying file contents to apply relocations. Currently, we mmap all input files and an output file and use memcpy to copy file contents.