6 ms·
> The 3990x runs a bit faster on the initial compile stage but the linking is single threaded and the M1 Max catches up at that point. Isn't linking IO-bound?
by fivea 5y ago
> The 3990x runs a bit faster on the initial compile stage but the linking is single threaded and the M1 Max catches up at that point.
Isn't linking IO-bound?
- kiratp 5y agoExposing my limited understanding of that level of the computing stack - it is but Apple seems to have very very good caching strategies - filesystem and L1/2/3. https://llvm.org/devmtg/2017-10/slides/Ueyama-lld.pdf https://llvm.org/devmtg/2017-10/slides/Ueyama-lld.pdf There is a breakdown in those slides discussing what parts of lld are single threaded and hard to parallelize so I suspect single thread performance plays a big role too. I generally observe one core pegged during linking.
- fivea 5y ago> Exposing my limited understanding of that level of the computing stack - it is but Apple seems to have very very good caching strategies - filesystem and L1/2/3. That would mean that these comparisons between Threadripper and the M1 Ultra do not reflect CPU performance but instead showcase whatever choice of SSD they've been using.
- LordDragonfang 5y agoL1/2/3 are CPU caches, not SSD. Though there is a good chance these are mostly firmware optimizations, not hardware. So still not an apples-to-apples comparison of cpu design.
- astrange 5y agoBy firmware you mean microcode, but I don't think either of those actually use microcode to control this.
- fivea 5y ago> L1/2/3 are CPU caches, not SSD. Why did you omit the reference to "file system"? Are we supposed to ignore the fact that a linker's main job is reading object files and write the output to a file? I find this sort of argument particularly comical given a very old school technique to speed up compilation is to use a RAM drive to store the build's output.
- nicoburns 5y agohttps://github.com/rui314/mold https://github.com/rui314/mold would suggest otherwise. Massive speedups by multithreading the linker. I think traditional linkers just aren't highly optimised.
- fivea 5y ago> https://github.com/rui314/mold https://github.com/rui314/mold would suggest otherwise. Does it, though? I mean, if you read that link you'll notice it boasts the linker's performance by comparing it with cp and how it's "so fast that it is only 2x slower than cp on the same machine." Is cp supposed to be CPU-bound?
- rsynnott 5y agoThe data the linker is running on will generally largely be in memory (in their posted example, which is a warm compile, completely in memory).
- nicoburns 5y agoYes, but that's for mold, which is multithreaded. The original context of this thread being the question of whether a linker would see speedups from multithreading. Most people are using traditional single-threaded linkers which are an order of magnitude slower than mold. The fact that mold is so much faster suggests that a linker does indeed see big speedups from multithreading.
- codeflo 5y agoFor a clean build and a reasonably specced machine, all the intermediate artifacts will still be in the cache during linking.