4 ms·
Not sure why you think clang having multiple not-identical executables is relevant to Backblaze having multiple identical executables.
by jfrunyon 5y ago
Not sure why you think clang having multiple not-identical executables is relevant to Backblaze having multiple identical executables.
- dataflow 5y ago> Not sure why you think clang having multiple not-identical executables is relevant to Backblaze having multiple identical executables. Really? You genuinely don't see why I would think a case with >99.99% identical executables might be relevant to a case with 100% identical executables?
- 300bps 5y agoI genuinely don’t see why either. With 100% identical executables, there does not appear to be a good reason to have multiple copies under any circumstance. With anything < 100% identical, well, maybe there’s a good reason to have multiples. Who knows? Id probably give someone the benefit of the doubt and figure there was some engineering challenge that made it faster/easier to do it that way. So yes, 100% identical is completely different than almost 100% identical.
- dataflow 5y ago> So yes, 100% identical is completely different than almost 100% identical. They're completely (!) different? And you're saying this despite the fact that the comment I replied to was discussing cases where one could "just execute one binary n times (with different argv[0] if they'd like)"... which is something you can do with different executables just as well as with identical ones? It's not just a little different but completely different? So different that not only you don't see any similarity, but you also cannot fathom why I might think there's some similarity?!
- notthathardbro 5y agoNo, actually. And your incredulous doubling-down is, well, making it more obvious that you seem to be missing the point. If they're 99% the same, it's generously easy to assume that there's a material difference. That assumption is completely nonsensical if they're 100% identical. So no, for the sake of every bit of context in this conversation, it does not make sense that you'd bring up an unrelated scenario of "similar" binaries.
- dang 5y agoWould you please stop posting in the flamewar style to HN? It's not what this site is for, and it destroys what it is for. We've had to ask you about this more than once in the past already. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- dataflow 5y agoApologies, but I still don't know how to respond in a situation like this. Can I ask how you would respond to this case yourself? I hope you can at least see why I found it difficult to swallow when people shut down my comment telling me there's zero similarity between a 99.99% case and a 100% case. It seems needlessly hyperbolic to me at best, and quite obviously frustrating. Looking back, I almost feel like it baited me into this. How do you see this from your standpoint? Does it feel like a good-faith comment to you that they can't tell why I would see 99.99% and 100% as similar?
- 300bps 5y agoI hope I can answer that for you. The core issue is that you are taking the original point out of context and kind of robotically parsing snippets. Context is important. For example, chimpanzee DNA is 99% identical to human DNA. Without context you could make the argument, “They are almost identical!” and be correct. But in the context of discussing the ability to fly to the moon and return safely, the two genomes are completely different.
- dataflow 5y agoI appreciate the reply. I don't understand what context you believe I was missing? The analogy is incredibly confusing to me too, because humans are more than their DNA: you obviously can't hardlink humans, nor could you replace them with stubs that call some "common" humans, it's not like the obstacle to doing these goes away even with 100% equal DNA vs. 99%. But with computer code, you very much can (for example) easily replace all the executables with a combined one that gets hardlinked with different names, such that they behave differently depending on argv[0]. They don't even need to be 100% identical for us to be able to do this; it works just as well with 1% identical code—you just end up with bigger combined binaries. This is the Busybox approach (notice it combines executables that have hardly anything in common), and it's in fact exactly what Clang already sometimes does (like on MSYS2); one would think they could take the same approach here. This is also precisely the context of the comment I replied to was saying, right? This is the context I was replying to. It's so confusing to me that you claim I'm missing context when I was in fact addressing the context I saw directly—and it was the other comments that were not.
- comex 5y agoThere is a reason: the majority of the file size comes from statically linked LLVM libraries. You can instead configure LLVM’s build system to build a single dynamic library and have the tools link to it, and this eliminates all of the duplication. However, it apparently comes with a “substantial performance penalty” [1] due to the nature of dynamic linking. (This actually surprises me, and I wonder whether it’s only referring to the inability to do LTO, or whether even LTO-less static linking is faster. Aside from the startup time issue.) A theoretical alternative would be to build all of the tools into a single executable, à la Busybox, where the combined executable would inspect argv[0] to figure out which tool’s code should be run. That way you could statically link the LLVM libraries without duplicating them in multiple executables. LLVM’s build system does not support this. I think it would be nice if it did, but it would be nontrivial to implement. [1] https://llvm.org/docs/BuildingADistribution.html https://llvm.org/docs/BuildingADistribution.html
- dataflow 5y ago> (This actually surprises me, and I wonder whether it’s only referring to the inability to do LTO, or whether even LTO-less static linking is faster.) I think the latter is also the case (though to a much lesser extent) on x64. One of the unfortunate features of x64 is it lacks direct 64-bit jumps, so every jump to an external library ends up being an indirect call. (In fact, with a potential memory load on top of that.) This was kind of surprising for me when I learned it too; it doesn't apply to x86.