21 ms·
Wow, this hits close to home. Doing a page fault where you can't in the kernel is exactly what I did with my very first patch I submitted after I joined the Mic
by steelframe 2y ago
Wow, this hits close to home. Doing a page fault where you can't in the kernel is exactly what I did with my very first patch I submitted after I joined the Microsoft BitLocker team in 2009. I added a check on the driver initialization path and didn't annotate the code as non-paged because frankly I didn't know at the time that the Windows kernel was paged. All my kernel development experience up to that point was with Linux, which isn't paged.
BitLocker is a storage driver, so that code turned into a circular dependency. The attempt to page in the code resulted a call to that not-yet-paged-in code.
The reason I didn't catch it with local testing was because I never tried rebooting with BitLocker enabled on my dev box when I was working on that code. For everyone on the team that did have BitLocker enabled they got the BSOD when they rebooted. Even then the "blast radius" was only the BitLocker team with about 8 devs, since local changes were qualified at the team level before they were merged up the chain.
The controls in place not only protected Windows more generally, but they even protected the majority of the Windows development group. It blows my mind that a kernel driver with the level of proliferation in industry could make it out the door apparently without even the most basic level of qualification.
- bonestamp2 2y ago> without even the most basic level of qualification That was my first thought too. Our company does firmware updates to hundreds of thousands of devices every month and those updates always go through 3 rounds of internal testing, then to a couple dozen real world users who we have a close relationship with (and we supply them with spare hardware that is not on the early update path in case there is a problem with an early rollout). Then the update goes to a small subset of users who opt in to those updates, then they get rolled out in batches to the regular users in case we still somehow missed something along the way. Nothing has ever gotten past our two dozen real world users.
- rvnx 2y agoOr it could be made that Windows stops loading drivers that are crashing. Third-party driver/module crashed more than 3 times in a row -> Third-party driver/module is punished and has to be manually re-enabled.
- Lx1oG-AWb6h_ZG0 2y agoWouldn't this be an attack vector? Use some low-hanging bug to bring down an entire security module, allowing you to escalate?
- SAI_Peregrinus 2y agoIt's currently a DOS by the crashing component, so it's already broken the Availability part of Confidentiality/Integrity/Availability that defines the goals of security.
- hunter2_ 2y agoBut a loss of availability is so much more palatable than the others, plus the others often result in manually restricting availability anyway when discovered.
- jamie0 2y agoI think the wider societal impact from the loss of availability today - particularly for those in healthcare settings - might suggest this isn't always the case
- yibg 2y agoThis was my first thought too. I'm not that familiar with the space, but I would think for something this sensitive the rollout would be staggered at least instead of what looks like globally all at the same time.
- dudeism_est_03 2y agoThis is the bit I am still trying to understand. On CrowdStrike you can define how many updates a host is behind. I.e. n (latest), n-1 (one behind) or n-2 etc. This update was applied to a 'latest' policy hosts and the n-2 hosts. To me it appears that there was more to this than just a corrupt update, otherwise how was this policy ignored? Unless it doesn't separate the update as deeply and maybe just a small policy aspect, which would also be very concerning. I guess we won't really know until they release the post mortem...
- Xamayon 2y agoYeah, my guess is that they roll out the updates to every client at the same time, and then have the client implement the n-1/2/whatever part locally. That worked great-ish until they pushed a corrupt (empty) update file which crashed the client when it tried to interpret the contents... Not ideal, and obviously there isn't enough internal testing before sending stuff out to actual clients.
- mihaaly 2y agoExactly this what I was missing in the story. Like why not to have a limited set of users have it before going live for the whole user base at a mission critical product like this is beyond comprehension of everyone ever came across software bugs (so billions of people). And then we already overcame the part of not testing internally well, or at all? Something clusteruck must have happened there which is still better than imagining that this is the normal way the organization operates. Which is a very scary vision. Serious rethinking of trusting this organization is due everywhere!
- nikau 2y agoBut that would require hiring staff to manage the process, and that is money taken away from sponsoring an F1 racing team.
- Rinzler89 2y agoThe funniest part was seeing Mercedes F1 team pit crew staring at BSODs at their workstations[1] while wearing CrowdStrike t-shirts. Some jokes just write themselves. Imagine if they loose the race because of their sponsor. But hey, at least they actually dogfood the products of their sponsors instead of just taking money to shill random stuff. [1] https://www.thedrive.com/news/crowdstrike-sponsored-mercedes-f1-has-recovered-from-the-blue-screens-of-death https://www.thedrive.com/news/crowdstrike-sponsored-mercedes...
- cush 2y agoIt's baffling how fast and wide the blast radius was for this Crowdstrike update. Quite impressive actually, if you think about it - updating billions of systems that quickly.
- rkagerer 2y agoThat is the right way to do it.
- throwaway7356 2y agoBut do you ever get free world-wide advertisement that everyone uses your product? Crowdstrike sure did and I'm sure they'll use that to sell it to more people.
- brightlancer 2y ago> Even then the "blast radius" was only the BitLocker team with about 8 devs, since local changes were qualified at the team level before they were merged up the chain. Up the chain to automated test machines, right?
- nick7376182 2y agoYou would think automated test would come before your teammates work stations / commit to head.
- steelframe 2y agoDid I mention this was 15 years ago? Software development back then looked very different than it does now, especially in Wincore. There was none of this "Cloud-native development" stuff that we all know and love today. GitHub was just about 1 year old. Jenkins wouldn't be a thing for another 2 years. In this case the "automated test" flipped all kinds of configuration options with repeated reboots of a physical workstation. It took hours to run the tests, and your workstation would be constantly rebooting, so you wouldn't be accomplishing anything else for the rest of the day. It was faster and cheaper to require 8 devs to rollback to yesterday's build maybe once every couple of quarters than to snarl the whole development process with that. The tests still ran, but they were owned and run by a dedicated test engineer prior to merging the branch up.
- nick7376182 2y agoSorry, the comment wasn't meant to be a personal judgement on you.
- Symbiote 2y agoJenkins was called Hudson from 2005 until 2011, and version control is much, much older. I'm surprised you didn't have two or more workstations.
- dralley 2y agohttps://www.usenix.org/system/files/1311_05-08_mickens.pdf https://www.usenix.org/system/files/1311_05-08_mickens.pdf "Perhaps the worst thing about being a systems person is that other, non-systems people think that they understand the daily tragedies that compose your life. For example, a few weeks ago, I was debugging a new network file system that my research group created. The bug was inside a kernel-mode component, so my machines were crashing in spectacular and vindic- tive ways. After a few days of manually rebooting servers, I had transformed into a shambling, broken man, kind of like a computer scientist version of Saddam Hussein when he was pulled from his bunker, all scraggly beard and dead eyes and florid, nonsensical ramblings about semi-imagined enemies. As I paced the hallways, muttering Nixonian rants about my code, one of my colleagues from the HCI group asked me what my problem was. I described the bug, which involved concur- rent threads and corrupted state and asynchronous message delivery across multiple machines, and my coworker said, “Yeah, that sounds bad. Have you checked the log files for errors?” I said, “Indeed, I would do that if I hadn’t broken every component that a logging system needs to log data. I have a network file system, and I have broken the network, and I have broken the file system, and my machines crash when I make eye contact with them. I HAVE NO TOOLS BECAUSE I’VE DESTROYED MY TOOLS WITH MY TOOLS. My only logging option is to hire monks to transcribe the subjective experience of watching my machines die as I weep tears of blood.”
- Arrath 2y agoThat's beautiful.
- qingcharles 2y agoAh, the joys of trying to come up with creative ways to get feedback from your code when literally nothing is available. Can I make the beeper beep in morse code? Can I just put a variable delay in the code and time it with a stopwatch to know which value was returned from that function? Ughh.
- suzzer99 2y agoI call this "throwing dye in the water".
- 2y ago
- deleted 2y ago[deleted]
- brcmthrowaway 2y agoWhat does this mean? Windows kernel paged, linux non paged?
- ec109685 2y agoLinux kernel memory isn’t paged out to disk, while Windows kernel memory can be: https://knowledge.broadcom.com/external/article/32146/third-party-applications-windows-kernel.html#:~:text=Windows%20operating%20systems%20use%20two%20finite%20memory,and%20the%20much%20smaller%20non%2Dpaged%20pool%20(NPP)%2C https://knowledge.broadcom.com/external/article/32146/third-...
- nw05678 2y agoHas that changed? I remember always creating a swap partition that was meant to be at least the size of RAM
- ww520 2y agoThe memory used by the Windows kernel is either Paged or Non-Paged. Non-Paged means pinning the memory in physical RAM. Paged means it might be swapped out to disk and paged back in when needed. OP was working on BitLocker a file system driver, which handles disk IO. It must be pinned in physical RAM to be available all the times; otherwise, if it's paged out, an IO request coming would find the driver code missing in memory and try to page in the driver code, which triggers another IO request, creating an infinite loop. The Windows kernel usually would crash at that point to prevent a runway system and stops at the point of failure to let you fix the problem.
- deleted 2y ago[deleted]
- quaintdev 2y agoThank you!
- p_l 2y agoLinux is a bit unusual in that kernel memory is generally physically mapped and unless you use vmalloc any memory you allocate has to correspond to pages backed by RAM. This also ties into how file IO happens, swapping, and how Linux approach to IO is actually closer to Multics and OS/400 than OG Unix. Many other systems instead default to using full power of virtual memory including swapping kernel space to disk, with only things explicitly need to be kept in ram being allocated from "non-paged" or "wired" memory. EDIT: fixed spelling thanks to writing on phone.
- temac 2y ago"It blows my mind that a kernel driver with the level of proliferation in industry could make it out the door apparently without even the most basic level of qualification." It was my understanding that MS now sign 3rd party kernel mode code, with quality requirements. In which case why did they fail to prevent this?
- einpoklum 2y ago> In which case why did they fail to prevent this? "Oh, crowdstrike? Yeah, yeah, here's that Winodws kernel code signing key you paid for."
- whydoyoucare 2y agoYou can pay for it and sign a file full of null characters. Signing has nothing to do with quality from what I understand.
- einpoklum 2y ago"Yours sincerely, Crowdstrike --- PS - If you get hit by some massive crash, we refer you to our company's name. What were you expecting?"
- Drygord 2y ago[flagged]
- 2y ago
- password4321 2y agohttps://news.ycombinator.com/item?id=41006104#41006555 https://news.ycombinator.com/item?id=41006104#41006555 the flawed data was added in a post-processing step of the configuration update, which is after it's been tested internally but before it's copied to their update servers per a new/green account
- cududa 2y agoSo have we decided to stop using checksums or something?
- password4321 2y agoPerhaps it was the checksum/signature process!
- function_seven 2y agoYa gotta keep checksumming until you find a fixed point.
- spaceywilly 2y ago“And so that’s why we recommend using phased rollouts” -Every DevOps engineer from now on
- prox 2y ago“But that costs us money and time” - some suit.
- Woodi 2y ago"And they promise fast threat mitigation... Let allow them to take over EVERYTHING! With remote access, of course. Some form of overwatch of what they in/out by our staff ? Meh... And it even allow us to do cuts in headcount and infra by $<digits_here> a year."
- 2y ago
- jboy55 2y agoI was thinking, this doesn't seem like its a case of all these machines still on an old version of windows, or some specific version, that is having issues. Therefore QA just missed one particular variant in their smoke testing. It seems like its every windows instance with that software, so either they don't have basic automated testing, or someone pushed this outside of a normal process.
- EasyMark 2y agoThis is what I don’t get, it’s extremely hard for me to believe this didn’t get caught in CI when things started blue screening. Every place I ever did test rebooting/powercycling was part of CI, with various hardware configs. This was before even our lighthouse customers even saw it.
- deleted 2y ago[deleted]
- tomrod 2y agoDisgruntled employee trying to use Crowd Strike to start a General Strike?
- Fire-Dragon-DoL 2y agoWhat makes you think they have CI after what happened?
- simonh 2y agoApparently the flaw was added to the config file in post-processing after it had completed testing. So they thought they had testing, but actually didn't.
- mtlynch 2y ago>Doing a page fault where you can't in the kernel is exactly what I did with my very first patch I submitted after I joined the Microsoft BitLocker team in 2009. Hello from a fellow BitLocker dev from this time! I think I know who this is, but I'm not sure and don't want to say your name if you want it private. Was one of your Win10 features implementing passphrase support for the OS drive? In any case, feel free to reach out and catch up. My contact info is in my profile.
- steelframe 2y agoWin8. I've been seeing your blog posts show up here and there on HN over the years, so I was half expecting you to pick up on my self-doxx. I'll ping you offline.
- Fire-Dragon-DoL 2y agoI'm completely ignorant on the topic but isn't rebooting a default test for kernel code, given how sensitive it is?
- steelframe 2y agoOh I rebooted, I just didn't happen to have the right configuration options to invoke the failure when I rebooted. Not every dev workstation was bluescreening, just the ones with the particular feature enabled.
- martin-adams 2y agoThat sounds like it was caught by luck, unless there was some test explicitly with that configuration in the QA process?
- dagmx 2y agoA lot of QA, especially at the system level, is just luck. That’s why it’s so important to dogfood internally imho. And by internally I don’t just mean the development team, but anyone and everyone at the company who is allowed to have access to early builds.
- account42 2y agoThere's "something that requires highly specific conditions managed to slip past QA" and then there's "our update brought down literally everyone using the software". This isn't a matter of bad luck.
- jokab 2y agoMaybe thru luck, they're gonna uncover another xz utils backdoor MS version, but its probably gonna get covered up because, Microsoft
- sateesh 2y agoBut as someone already pointed out, the issue was seen on all kinds of windows hosts. Not just the ones running a specific version, specific update etc.
- isatty 2y agoI do not mean this to be blamey in any way shape or form and am asking only about the process: Shouldn’t that have been caught in code review?
- steelframe 2y agoMy manager actually blamed the more senior developer who reviewed my code for that one.
- usr1106 2y ago> I didn't know at the time that the Windows kernel was paged. At uni I had a professor in database systems, who did not like written exams, but mostly did oral exams. Obviously for DBMSes the page buffer is very relevant, so we chatted about virtual memory and paging. So in my explanation I made the difference for kernel space and user space. I am pretty sure I had read that in a book describing VAX/VMS internals. However, the professor claimed that a kernel never does paging for its own memory. I did not argue on that and passed the exam with the best grade. Did not check that book again to verify my claim. I have never done any kernel space development even vaguely close to memory management, so still today I don't know the exact details. However, what strikes me here: When that exam happened in 1985ish the NT kernel did not exist yet, I'd believe. However, IIRC a significant part of the DEC VMS kernel team went to Microsoft to work on the NT kernel. So the concept of paging (a part of) kernel memory went with them? Whether VMS --> WNT, every letter increased by one is just a coincidence or intentionally the next baby of those developers I have never understood. As Linux has shown us today much bigger systems can be successfully handled without the extra complications for paging kernel memory. Whether it's a good idea I don't know, at least not a necessary one.
- nullindividual 2y agoIf you want to hear the history of [DEC/VMS] NT from the horses mouth: https://www.youtube.com/watch?v=xi1Lq79mLeE https://www.youtube.com/watch?v=xi1Lq79mLeE
- usr1106 2y agoOh oh, 3 hours 10. I watched around half of it. The VMS --> WNT acronym relationship was not mentioned, maybe it was just made up later. One thing I did not know (or maybe not remember) is that NT was originally developed exclusively for the Intel i860, one of Intel's attempts to do RISC. Of course in the late 1980s CISC seemed deemed and everyone was moving to RISC. The code name of the i860 was N10. So that might well be the inside origin of NT, the marketing name New Technology retrofitted only later.
- 2y ago
- usr1106 2y ago> It blows my mind that a kernel driver with the level of proliferation in industry could make it out the door apparently without even the most basic level of qualification. Discussed elsewhere it is claimed that the file causing the crash was a data file that has been corrupted in the delivery process. So the development team and their CI have probably tested a good version, but the customer received a bad one. If that is true to problem is that the driver first uses an unsigned file at all, so all customer machines are continuously at risk for local attacks. And then it does not do any integrity check on the data it contains, which is a big no no for all untrusted data, whether user space or kernel.
- vlod 2y agoIf the file was signed, wouldn't that have prevented the corrupted transmission file from being loaded. I assume if the signed file was hacked (or parts missing), then it wouldn't pass verification.
- wedesoft 2y agoStill a staggered roll-out would have reduced the impact.
- mbreese 2y ago> And then it does not do any integrity check on the data it contains, which is a big no no for all untrusted data, whether user space or kernel. To me, this is the inexcusable sin. These updates should be signed and signatures validated before the file is read. Ideally the signing/validating would be handled before distribution so that when this file was corrupted, the validation would have failed here. But even with a good signature, when a file is read and the values don’t make sense, it should be treated as a bad input. From what I’ve seen, even a magic bytes header here would have helped.
- sandworm101 2y agoSo the key test, the test that was not run, was to turn the machine off and on again? Classic windows.
- hevisko 2y agoMust have been DNS... when they did the deployment run and the necessary code was pulled and the DNS failed and then the wrong code got compiled...</sarcasm> that they don't even do staged/A-B pushes was also <mind-blown-away> But the most.... ironical was: https://www.theregister.com/2024/07/18/security_review_failure/ https://www.theregister.com/2024/07/18/security_review_failu...