8 ms·
Could the XZ backdoor been detected with better Git/Deb packaging practices?
- ottoke 1y agoHow did the changes in the binary test files tests/files/bad-3-corrupt_lzma2.xz and tests/files/good-large_compressed.lzma, and the makefile change in m4/build-to-host.m4) manifest to the Debian maintainer? Was there a chance of noticing something odd?
- Groxx 1y agomostly no, from my reading - it was a multi-stage chain of relatively normal looking things that added up to an exploit. helped by the tests involved using compressed data that wasn't human-readable. you can of course come up with ways it could have been caught, but the code doesn't stand out as abnormal in context. that's all that really matters, unless your build system is already rigid enough to prevent it, and has no exploitable flaws you don't know about. finding a technical overview is annoyingly tricky, given all the non-technical blogspam after it, but e.g. https://securelist.com/xz-backdoor-story-part-1/112354/ https://securelist.com/xz-backdoor-story-part-1/112354/ looks pretty good from a skim.
- XorNot 1y agoCompression algorithms are deterministic over fixed data though (possibly with some effort). There's no good reason to have opaque, non generated data in the repository and it should certainly be a red flag going forwards.
- Groxx 1y agocommitted files with carefully crafted bad data is extremely common for testing how your code handles invalid data, especially with regression tests. and lzma absolutely needs to test itself against bad, possibly-malicious data.
- sanjams 1y agoI agree, but perhaps OP is suggesting that the hand-crafted data can be generated in a more transparent way. For example, via a script/tool that itself can be reviewed.
- Groxx 1y agocould have, absolutely. should not have been in any commit, which is basically necessary to prevent this case, almost definitely not. it's normal, and requiring all data to be generated just means extremely complicated generators for precise trigger conditions... where you can still hide malicious data. you just have to obfuscate it further. which does raise the difficulty, which is a good thing, but does not make it impossible. I completely agree that it's a good/best practice, but hard-requiring everywhere it has significant costs for all the (overwhelmingly more common) legitimate cases.
- XorNot 1y agoIt would be reasonable for error case data though to be thoroughly explained, and it must be explainable since otherwise what are you testing and why does the test exist? The xz exploit depended on the absence of that explanation but accepting that it was necessary for unstated reasons. Whereas it's entirely reasonable to have a test that says something like: "simulate an error where the header is corrupted with early nulls for the decoding logic" or something - i.e. an explanation, and then a generator which flips the targeted bits to their values. Sure: you _could_ try inserting an exploit, but now changes to the code have to also surface plausible data changes inline with the thing they claim is being tested. I wouldn't even regard that as a lot of work: why would a test like that exist, if not because someone has an explanation for the thing they want to test?
- crote 1y agoYes, but the carefully crafted bad data should be explainable. Instead of committing blobs, why not commit documented code which generates those blobs? For example, have a script compress a bunch of bytes of well-known data, then have it manually corrupt the bytes belonging to file size in the archive header.
- secondcoming 1y agoThere are tons of reasons to have hand-crafted data in a repository.
- sanjams 1y agoThe article references a technical write-up: https://research.swtch.com/xz-script https://research.swtch.com/xz-script
- Groxx 1y agoah, yes, this is one I remember seeing early on! thank you! I couldn't find much past the blogspam this time :/
- dhx 1y agoThere's a few obvious gaps, seemingly still unsolved today: 1. Build environments may not be adequately sandboxed. Some distributions are better than others (Gentoo being an example of a better approach). The idea is that the package specification specifies the full list of files to be downloaded initially into a sandboxed build environment, and scripts in that build environment when executed are not able to then access any network interfaces, filesystem locations outside the build environment, etc. Even within a build of a particular software package, more advanced sandboxing may segregate test suite resources from code that is built so that a compromise of the test suite can't impact built executables, or compromised documentation resources can't be accessed during build or eventual execution of the software. 2. The open source community as a whole (but ultimately in the hands of distribution package maintainers) are not being alerted to and apply caution for unverified high entropy in source repositories. Similar in concept to nothing-up-my-sleeve numbers.[1] Typical examples of unverified high entropy where a supply chain attack can hide payload: images, videos, archives, PDF documents etc in test suites or bundled with software as documentation and/or general resources (such as splash screens in software). It may also include IVs/example keys in code or code comments, s-boxes or similar matrices or arrays of high entropy data which may not be obvious to human reviewers how the entropy is low (such as a well known AES s-box) rather than high and potentially undifferentiated from attacker shellcode. Ideally when a package maintainer goes to commit a new package or package update, are they alerted to unexplained high entropy information that ends up in the build environment sandbox and required to justify why this is OK? [1] https://en.wikipedia.org/wiki/Nothing-up-my-sleeve_number https://en.wikipedia.org/wiki/Nothing-up-my-sleeve_number
- normie3000 1y ago[flagged]
- TacticalCoder 1y ago[dead]
- flerchin 1y ago> As of today only 93% of all Debian source packages are tracked in git on Debian’s GitLab instance at salsa.debian.org. Some key packages such as Coreutils and Bash are not using version control at all This bends my brain a little. I get that they were written before git, but not before the advent of version control.
- simoncion 1y ago> I get that they were written before git, but not before the advent of version control. git clone https://git.savannah.gnu.org/git/bash.git git clone https://git.savannah.gnu.org/git/coreutils.git Plug the repo name into https://savannah.gnu.org/git/?group=<REPO_NAME> to get a link to browse the repo.
- kryptiskt 1y agoLook at the commit log in the bash repo. What good does it do if it notionally is version controlled if the commits look like this: 2025-07-03 Bash-5.3 distribution sources and documentation bash-5.3 Chet Ramey 896 -103357/+174007
- simoncion 1y agoThat looks to be the headline for the public release commit. If you'd bothered to look around for a full sixty seconds, you'd have found that the commits tagged with bash-5.3 and bash-5.2 follow that format. Here are the headlines for a couple of fix commits: Bash-5.2 patch 12: fixes for compat mode leaving extglob enabled after command substitution Bash-5.2 patch 1: fix crash with unset arrays in arithmetic contexts It looks like discussion of the patches happens on the mailing list, which is easy to access from the page that brought you to the repo browser.
- oivey 1y agoAhh yes, if only the commit message was better. That would have stopped the xz attack.
- charcircuit 1y agoIt shouldn't have happened in the first place. OpenSSH should control their exact dependencies and Debian shouldn't be meddling with them and swapping them out, loading random code into OpenSSH's process. >we can only trust open source software. There is no way to audit closed source software The ability to audit software is not sufficient, nor neccessary for it to be trustworthy. >systems of a closed source vendor was compromised, like Crowdstrike some weeks ago, we can’t audit anything You can't audit open source vendors either.
- IshKebab 1y agoThat's really incidental. There are a gazillion vectors for exploitation once you control a package like xz. You can't fix this issue by plugging them one by one.
- 1718627440 1y ago> Debian shouldn't be meddling with them Debian is the OS, and the OS vendor should decide and modify the components it uses as a foundation to create the OS as he desires. That's what I am choosing Debian for and not some other OS. > You can't audit open source vendors either. What defines open source, is that you can request the sources for audit and modification, so I think this statement is just untrue.
- charcircuit 1y agoIf Debian wants to improve or modify OpenSSH and put their own code is, they should rename it and stop using the name of the project. Debian's actions created reputational damage by introducing a backdoor into someone else's product without clearly informing the consumer that they did so. >you can request the sources Organizarions that open source software can have closed source infrastructure that you can't request.
- 1718627440 1y agoDebian is famous for modifying all programs it ships, it is more the rule than the exception. That's the deal I get when choosing Debian. SSH is more of a protocol, than a trademarked program. > Organizarions that open source software can have closed source infrastructure that you can't request. Which can't be a source for the program binaries, so you can still audit them, you just can't rely on e.g. their proprietary test suite.
- ape4 1y agoWouldn't the next malware use a different way to embed itself
- xmodem 1y agoWhy would they bother if we don't act on any of the learnings from this one?
- bluGill 1y agoMaybe - but original ideas are hard, ideas without flaws are rare: there are reasonable odds someone will try this again.
- citbl 1y agoThe next one probably won't be caught by running noticeably slower than usual. It was a pure fluke that it got discovered _this early_.
- ApolloFortyNine 1y agoIf I remember correctly it's days were numbered as soon as that redhat bug report on the valgrind errors piling up was made. They weren't the ones to find the cause first (that's the person who took a deeper look due to the slowness), but the red flags had been raised.
- paulf38 11mo agoYes indeed. The backdoor author did try to claim that it was a false positive (and I’m sure that a very depressingly large number of people would happily go along with such a claim even without a scrap of evidence). The error was related to the use of the frame pointer. Optimised code does not use RBP as the frame pointer, only using RSP for stack addresses. The XZ backdoor code assumed that the stack used this layout. The RedHat regression tests use debug builds that do use the frame pointer. The result was the backdoor code writing below the bottom of the stack. I suspect also that Valgrind is unique in finding issues like this. Other tools do not check all memory accesses before main. Valgrind loads and runs the test binary from the very beginning and thus it detected errors in the ifunc code used by XZ that executed very early on during ld.so loading and symbol resolution.
- jart 1y agoFolks have been ringing the alarm bell for a decade. https://www.nongnu.org/lzip/xz_inadequate.html https://www.nongnu.org/lzip/xz_inadequate.html xz is insane because it appears to be one of the most legitimately dangerous compression formats with the potential to gigafry your data but is exclusively used by literal turbonormies who unironically want to like "shave off a few kilobytes" and basically get oneshotted by it.
- Delk 1y agoThe question of whether the xz format is a good choice for long-term archival is entirely unrelated to backdoors or open source supply chain security.
- jart 1y agoNo they're the same. Why do you think xz was targeted? It's a giant slippery hairball.
- Delk 1y ago> Why do you think xz was targeted? Possibly for any number of reasons. A sole maintainer with a bit too little capacity to keep up the development. A central role as a dependency for crucial packages in a couple of key distros. What would be the connection between the backdoor (or indeed any supply chain security) and any design details of the xz file format? How would the backdoor have been avoided if the archive format were different?
- deleted 11mo ago[deleted]
- tredre3 1y agoTurbonormies, as you say, tend to use gzip not xz. Which is sad because gzip is just as bad for archiving. A few bytes changed and your entire file is lost (in a .tar.gz it means everything is lost). Frankly, tarballs are an embarrassing relic, and it's not the turbonormies that insist they're still fit for purpose. They don't know any better, they'll do what people like you tell them to do.
- 1970-01-01 1y ago>Can we trust open source software? Yes — and I would argue that we can only trust open source software. But should we trust it? No!! That's why we're here! I'm not satisfied with the author's double-standard-conclusion. Trust, but verify does not have some kind of hall pass for OSS "because open-source is clearly better." Trust, but verify is independent of the license the coders choose.
- rcxdude 1y agoYes, I would say that being able to view the source code and build it yourself is a necessary but not sufficient condition of properly trusting the software. (which is not quite the same thing as it being open source, but it's relatively rare outside of being a very big customer that you can do this for non-open-source code).
- 1718627440 1y agoWhen you get the source code as a big costumer, that is open source. It might even be free software.
- johnny22 1y agomany folks make a distinction between source available and open source.
- 1718627440 1y agoThe latter meaning: accepts patches, or what?
- NekkoDroid 1y ago"Open Source" meaning the license is OSI approved (or at least meets the definition for "Open Source" by the OSI[1]) and source available is anything to which you can get the source to, but the license doesn't meet the above criteria. [1]: https://opensource.org/osd https://opensource.org/osd
- sega_sai 1y agoFrom reading this, it seems that one thing one can do is to be force separation of the build from testing, so the build never has access to binary code that can be injected.
- acka 1y agoI believe the XZ compromise partly stemmed from including binary files in what should have remained a source-only project. From what I remember, well-run projects such as those of the GNU project have always required that all binaries—whether executables or embedded data such as test files—be built directly from source, compiling a purpose-built DSL if necessary. This ensures transparency and reproducibility, both of which might have helped catch the issue earlier.
- dijit 1y agothats not the issue, there will always be prebuilt binaries (hell, deb/rpm are prebuilt binaries). The issue for xz was that the build system was not hermetic (and sufficiently audited). Hermitic build environments that can’t fetch random assets are a pain to maintain in this era, but are pretty crucial in stopping an attack of this kind. The other way is reproducible binaries, which is also very difficult. EDIT: Well either I responded to the wrong comment or this comment was entirely changed. I was replying to a comment that said. “The issue was that people used pre-built binaries” which is materially different to what the parent now says, though they rhyme.
- mananaysiempre 1y agoThe XZ project’s build system is and was hermetic. The exploit was right there in the source tarball. It was just hidden away inside a checked-in binary file that masqueraded as a test for handling of invalid compressed files. (The ostensibly autotools-built files in the tarball did not correspond to the source repository, admittedly, but that’s another question, and I’m of two minds about that one. I know that’s not a popular take, but I believe Autotools has a point with its approach to source distributions.)
- dijit 1y agoI thought that the exploit was not injected into the Git repository on GitHub at all, but only in the release tarballs. And that due to how Autoconf & co. work, it is common for tarballs of Autoconf projects to include extra files not in the Git repository (like the configure script). I thought the attacker exploited the fact that differences between the release tarball and the repository were not considered particularly suspicious by downstream redistributors in order to make the attack less discoverable.
- octoberfranklin 1y agoYes of course, and nixpkgs (nixos) already does, although unfortunately not for this particular package. The XZ backdoor was possible because people stick generated code (autoconf's output), which is totally impractical to audit, into the source tarballs. In nixpkgs, all you have to do is add `autoreconfHook` to the `nativeBuildInputs` and all that stuff gets regenerated at build time. Sadly this is not the default behavior yet.
- jiggawatts 1y agoSomething that the XZ back door made me realise is that the fundamental difference between proprietary and open source software is not the price or source availability for most of its users — no not developers! — it is the reputation and protected brand of the former and the anonymity of the latter. We have no clue who “Jia Tan” is, a name certain to be a pseudonym. Nobody has seen his face. He never provided ID to a HR department. He pays no taxes to a government that links these transactions to him. There is no way to hold his feet to the fire for misdeeds. The open source ecosystem of tools and libraries is built by hundreds of thousands of contributors, most of whom are identified by nothing more than an email. Just a string of characters. For all we know, they’re hyper-intelligent aliens subtly corrupting our computer systems, preparing the planet for invasion! I mean… that’s facetious, but seriously… how would we know if it was or wasn’t the case!? We can’t! We have a scenario where the only protection is peer review: but we’ve seen that fail over and over systematically. Glaring errors get published in science journals all of the time. Not just the XZ attack but also Heartbleed - an innocent error - occurred because of a lack of adequate peer review. I could waffle on about the psychology of “ownership” and how it mixes badly with anonymity and outside input, but I don’t want this to turn into war and peace. The point is that the fundamental issue from the “outside” looking in as a potential user is that things go wrong and then the perpetrators can’t be punished so there is virtually no disincentive to try again and again. Jia Tan is almost certainly a state-sponsored attacker. A paid professional, whose job appears to be to infect open source with back doors. The XZ attack was very much a slow burn, a part time effort. If he’s a full time employee, how may more irons did he have on the fire! Dozens? Hundreds!? What about his colleagues? Certainly he’s not the one and only such hacker! What about other countries doing the same with their own staff of hackers? The popular thinking has been that “Microsoft bad, open source good”, but imagine Jia Tan trying to pull something like this off with the source of Windows Server! He’d have to get employed, work in a cubicle farm, and then if caught in the act, evade arrest! That’s a scary difference.
- TheDong 1y ago> Something that the XZ back door made me realise is that the fundamental difference between proprietary and open source software is not the price or source availability for most of its users — no not developers! - it is the reputation and protected brand of the former and the anonymity of the latter. You're making a distinction not between open source and proprietary software but rather between hobbyist and corporate software. There are open source projects made by companies with no external contributions allowed (sqlite sorta, most of google and amazon's oss projects in practice etc) There are proprietary software downloads with no name attached, like practically every keygen, game crack, many indie games posted for free download on forums or 4chan, etc etc.
- aborsy 1y agoCouldn’t the submission to the Debian be possible only under real identities so that people take responsibility for what they submit? A random person or group nobody has ever seen or knows submitted a backdoor.
- 0rdinal 1y ago1. How could Debian effectively verify an identity? 2. Some people may want to remain pseudonymous for legitimate reasons.
- aborsy 1y agoIt’s not straightforward. The developers (at least important ones) could register with Debian project, just like they would with a company: submit identity and government documents, proof of physical address, bank account, credit card information, IdP account, .. It would operate like an organization. The lead developers could meet and know each other through regular meetings. Kind of web of trust with in person verification. There are already online meetings in some projects.