4 ms·
This is something that always bothered me while I was working at Google too: we had an amazing compute and storage infrastructure that kept getting crazier and
by jmmv 10mo ago
This is something that always bothered me while I was working at Google too: we had an amazing compute and storage infrastructure that kept getting crazier and crazier over the years (in terms of performance, scalability and redundancy) but everything in operations felt slow because of the massive size of binaries. Running a command line binary? Slow. Building a binary for deployment? Slow. Deploying a binary? Slow.
The answer to an ever-increasing size of binaries was always "let's make the infrastructure scale up!" instead of "let's... not do this crazy thing maybe?". By the time I left, there were some new initiatives towards the latter and the feeling that "maybe we should have put limits much earlier" but retrofitting limits into the existing bloat was going to be exceedingly difficult.
- joatmon-snoo 10mo agoThere's a lot of tooling built on static binaries: - google-wide profiling: the core C++ team can collect data on how much of fleet CPU % is spent in absl::flat_hash_map re-bucketing (you can find papers on this publicly) - crashdump telemetry - dapper stack trace -> codesearch Borg literally had to pin the bash version because letting the bash version float caused bugs. I can't imagine how much harder debugging L7 proxy issues would be if I had to follow a .so rabbit hole. I can believe shrinking binary size would solve a lot of problems, and I can imagine ways to solve the .so versioning problem, but for every problem you mention I can name multiple other probable causes (eg was startup time really execvp time, or was it networked deps like FFs).
- Filligree 10mo agoThere’s no way my proxy binary actually requires 25GB of code, or even the 3GB it is. Sounds to me like the answer is a tree shaker.
- deleted 10mo ago[deleted]
- MaskRay 9mo agoWe are missing tooling to partition a huge binary into a few larger shared objects. As my https://maskray.me/blog/2023-05-14-relocation-overflow-and-code-models#:~:text=static https://maskray.me/blog/2023-05-14-relocation-overflow-and-c... (linked by author, thanks! But I maintain lld/ELF instead of "wrote" it - it's engineer work of many folks) Quoting the relevant paragraphs below: ## Static linking In this section, we will deviate slightly from the main topic to discuss static linking. By including all dependencies within the executable itself, it can run without relying on external shared objects. This eliminates the potential risks associated with updating dependencies separately. Certain users prefer static linking or mostly static linking for the sake of deployment convenience and performance aspects: * Link-time optimization is more effective when all dependencies are known. Providing shared object information during executable optimization is possible, but it may not be a worthwhile engineering effort. * Profiling techniques are more efficient dealing with one single executable. * The traditional ELF dynamic linking approach incurs overhead to support [symbol interposition](https://maskray.me/blog/2021-05-16-elf-interposition-and-bsymbolic https://maskray.me/blog/2021-05-16-elf-interposition-and-bsy...). * Dynamic linking involves PLT and GOT, which can introduce additional overhead. Static linking eliminates the overhead. * Loading libraries in the dynamic loader has a time complexity `O(|libs|^2*|libname|)`. The existing implementations are designed to handle tens of shared objects, rather than a thousand or more. Furthermore, the current lack of techniques to partition an executable into a few larger shared objects, as opposed to numerous smaller shared objects, exacerbates the overhead issue. In scenarios where the distributed program contains a significant amount of code (related: software bloat), employing full or mostly static linking can result in very large executable files. Consequently, certain relocations may be close to the distance limit, and even a minor disruption (e.g. add a function or introduce a dependency) can trigger relocation overflow linker errors.
- jcalvinowens 9mo ago> We are missing tooling to partition a huge binary into a few larger shared objects Those who do not understand dynamic linking are doomed to reinvent it.
- lenkite 9mo agoMaybe I am missing something, but why didn't they just leverage dynamic libraries ?
- tmoertel 9mo agoOne reason is that using static binaries greatly simplifies the problem of establishing Binary Provenance, upon which security claims and many other important things rely. In environments like Google’s it's important to know that what you have deployed to production is exactly what you think it is. See for more: https://google.github.io/building-secure-and-reliable-systems/raw/ch14.html https://google.github.io/building-secure-and-reliable-system...
- inkyoto 9mo ago> One reason is that using static binaries greatly simplifies the problem of establishing Binary Provenance, upon which security claims and many other important things rely. It depends. If it is a vulnerability stemming from libc, then every single binary has to be re-linked and redeployed, which can lead to a situation where something has been accidentally left out due to a unaccounted for artefact. One solution could be bundling the binary or related multiple binaries with the operating system image but that would incur a multidimensional overhead that would be unacceptable for most people and then we would be talking about «an application binary statically linked into the operating system» so to speak.
- tmoertel 9mo ago> If it is a vulnerability stemming from libc, then every single binary has to be re-linked and redeployed, which can lead to a situation where something has been accidentally left out due to a unaccounted for artefact. The whole point of Binary Provenance is that there are no unaccounted-for artifacts: Every build should produce binary provenance describing exactly how a given binary artifact was built: the inputs, the transformation, and the entity that performed the build. So, to use your example, you'll always know which artefacts were linked against that bad version of libc. See https://google.github.io/building-secure-and-reliable-systems/raw/ch14.html#:~:text=to%20insider%20risk.-,Binary%20Provenance,-Every%20build%20should https://google.github.io/building-secure-and-reliable-system...
- darubedarob 9mo agoI think google of all companies could build a good autostripper reducing binaries by adding partial load assembly on misses. It cant be much slower then shovelling a full monorepo assembly plus symbols into ram.
- loeg 9mo agoThe low-hanging fruit is just not shipping the debuginfo, of course.
- usefulcat 9mo agoIs compressed debug info a thing? It seems likely to compress well, and if it's rarely used then it might be a worthwhile thing to do?
- loeg 9mo agoIt is: https://maskray.me/blog/2022-01-23-compressed-debug-sections https://maskray.me/blog/2022-01-23-compressed-debug-sections But the compression ratio isn't magical (approx. 1:0.25, for both zlib and zstd in the examples given). You'd probably still want to set aside debuginfo in separate files.
- Gibbon1 9mo agoSmall brained primate comment. With embedded firmware you only flash the .text and and flash to the device. But you still can debug using the .elf file. In my case if I get a bus fault I'll pull the offending address off the stack and use bintools and the .elf to show me who was naughty. I think if you have a crash dump you should be able to make sense of things as long as you keep the unstripped .elf file around.
- bfrog 9mo agoSounds like Google could really use Nix