6 ms·
Can someone who actually understands what CrowdStrike does explain to me why on earth they don't have some kind of gradual rollout for changes? It seems like th
by tail_exchange 2y ago
Can someone who actually understands what CrowdStrike does explain to me why on earth they don't have some kind of gradual rollout for changes? It seems like their updates go out everywhere all at once, and this sounds absolutely insane for a company at this scale.
- deleted 2y ago[deleted]
- Zamiel_Snawley 2y agoTruly, how the extent the damage was so widespread is my main question at this point. Everyone has a buggy release at some point, but impacting global customers at this level is damn near unforgivable. Heads need to roll for this oversight.
- hatsunearu 2y agoIt sounds like Channel files are just basically definition updates in normal antivirus software; it's not actually code, just some stuff on what the software should "look out for". And it sounds like they shipped some malformed channel file and the software that interprets it can't handle malformed inputs and ate shit. That software happened to be kernel mode, and also marked as boot-critical, so it if falls over, it causes a BSOD and inability to boot. and it's kind of understandable that channel files might seem safe to update constantly without oversight, but that's just assuming that the file that interprets the channel file isn't a bunch of dogshit code.
- Murky3515 2y agoIt's not understandable imo. At the very least they should have tests for the loader component that shows it can handle corrupted input. Amateur hour.
- throwaway346434 2y agoHalting problem is undecidable. On the scale of "no one bothered to put error handling or validation in" to "a subtle problem exists for this given input"; you and I lack the information to make a judgement.
- vanillax 2y agoI mean, the whole world was impacted. All they had to do was test this change in a lab with several pcs. Clearly this wasn't a edge case nor a subtle problem. This was clearly a lack of testing.
- andrewinardeer 2y agoIt was a Friday. Devs just wanted to go home for the weekend.
- acdha 2y agoLeave the spin to the PR people. Their customers pay a great deal of money for 24x7 service, and this wasn’t even a code change but a definition update – a process which should be as well defined and tested as McDonald’s making a hamburger. You wouldn’t excuse getting E. coli from your lunch with “the cook just wanted to go home for the weekend”, and this is a much more expensive service.
- acdha 2y ago> you and I lack the information to make a judgement. Think about this a little harder: what do you know about the number of customers affected? We do actually have enough information to make a judgement - bricking millions of critical systems, a very high percentage of their total Windows customer population, tells us that they don’t have progressive rollouts, don’t fail into a safe mode, and that if they do have tests those tests are catastrophically unlike anything their customers run – all they had to do was launch an EC2 instance and see if it kept running.
- liuliu 2y agoNot doing fuzzing on user-input supported feature, especially for AV, is damning.
- SoftTalker 2y agoAnd though I don't know, I'm guessing it's not a certainty to say they don't contain "code." It would seem to me that they would have to, otherwise novel attacks that weren't caught by one of their existing algorithms could never be detected. I'm guessing they contain some combination of pattern/regexp type stuff, and interpreted code/scripting with trigger criteria, etc. that all gets loaded into the "engine" that actually runs the threat detection.
- numbsafari 2y agoAgreed. We all know about a really interesting vector for infecting the kernel now. One that is poorly tested, poorly implemented, and poorly secured.
- hatsunearu 2y agoYeah, I re-read my comment and it sounds like I am understanding of them. But no, saying "channel files aren't kernel code" is just hilarious, considering the channel files define how the actual kernel code is supposed to behave, so it might as well be kernel code. Especially when the bad behavior in question is triggered by bad channel files!!
- Hizonner 2y ago> it's not actually code, just some stuff on what the software should "look out for" If it controls the behavior of a computer, then it's code. > and it's kind of understandable that channel files might seem safe to update constantly without oversight Yeah, no, it's not. They pushed an update that crashed the majority of their Windows installed base in a way that couldn't be fixed remotely. It doesn't matter what the update was to. It needed to be tested. There is no way that any deployment pipeline that could fail to catch something that blatant could possibly be "understandable". ... and that kernel mode code shouldn't have been parsing anything with any complexity to begin with. And should have been tested into oblivion, and possibly formally verified. This is amateur-hour nonsense. Which is what you expect from most of these "Enterprise Cyber Security(TM)" vendors. ... AND the users shouldn't have just gone and shoved that kind of thing into every critical path they could think of.
- omoikane 2y agoConfiguration files should be treated like code and follow the same gradual rollout practices. See also: https://sre.google/workbook/canarying-releases/ https://sre.google/workbook/canarying-releases/ Which starts with "a majority of incidents are triggered by binary or configuration pushes". The stats for config related failures is one link away at https://sre.google/workbook/postmortem-analysis/ https://sre.google/workbook/postmortem-analysis/ Where it says 31% of outages in 2010-2017 are caused by "configuration push".
- timbelina 2y agoI was reading these two threads: https://x.com/perpetualmaniac/status/1814376668095754753?s=46&t=HI4y6iDN2EUVjNCTjrCQZQ https://x.com/perpetualmaniac/status/1814376668095754753?s=4... https://x.com/ananayarora/status/1814269058088304760 https://x.com/ananayarora/status/1814269058088304760 The authors explain the coding error and coredump well, but I'm lost: Is the buggy code that they're describing the channel file, or some kernel code that consumes the channel file? Is there a way to tell?
- timbelina 2y agoOK, and another question:-) Can tools like Valgrind and ASan pick up the kinds of errors that are described in those two links from my previous post?
- ananayarora 2y agoAuthor of the second post here. The first author's stack trace seems to show a fault on csagent.sys which is a bad read on 0x9c. There are some other .sys files loaded up by csagent.sys, and that's where the crash seems to happen, apparently. As for detection, Zach mentions that modern tooling could've been used to find this, so I'm assuming Valgrind can find this: https://x.com/Perpetualmaniac/status/1814376690958868979 https://x.com/Perpetualmaniac/status/1814376690958868979 Hope this helps!
- timbelina 2y agoCheers Ananay! So if I put this all together: a) The driver (sensor) csagent.sys includes code that hasn't checked with a tool like Valgrind or ASan or something and so includes some kind of memory management bug. b) Since n, n-1 and n-2 versions of the sensor all died equally spectacularly, that bug as been around for at least three versions of csagent.sys. c) The bug can be triggered by getting the csagent.sys to swallow a shitty channel file and since csagent runs in kernel mode, when it crashes it BSOD's the system. d) Someone at Crowdstrike uploaded a shitty channel file as part of an update process that apparently happens many times a day. Am I on the right track so far? If so, there's no/inadequate memory management checks in the csagent driver, and either: 1)There were also no checks before the borked channel file was uploaded because of a failure to follow process, or because there was no process, but whatever the case it was an accident. or 2) Someone uploaded on purpose, not by accident, the borked channel file intending for a nasty outcome (probably not BSOD) I can't believe that there are not a million checks and balances in place to let (1) happen, but as my grandma used to say, "Don't assume malice where stupidity will do" :-)
- Murky3515 2y agoBecause if release immediately, velocity go up
- userbinator 2y agoI'm more surprised at the fact that they didn't appear to have tested it on themselves first. FWIW, at least Microsoft still "dogfoods" (and it's what coined that term), and even if the results of that aren't great, I'm sure they would've caught something of this severity... but then again, maybe not[1]. [1] https://news.ycombinator.com/item?id=18189139 https://news.ycombinator.com/item?id=18189139
- Ekaros 2y agoThis is what really would concern me too. With this wide spread issue any reasonable testing should have detected it. Having a few dozen machines with different configurations for an few hours should have detected this. This should have been in a smoke test. Push update to machines, observe, power cycle them, observe... I could understand error in some rarer setup, but this was so common that it should have been obvious error.
- SkyPuncher 2y agoMy understanding is they basically deployed a configuration file. It seems like these files might be akin to virus signatures or other frequently updated run-time configuration. I actually don't think it's outrageous that these files are rolled out globally, simultaneously. I'm guessing they're updated frequently and _should_ be largely benign. What stands out to me is the fact that a bad config file can crash the system. No rollback mechanism. No safety checks. No safe, failure mode. Just BSOD. Given the fix is simply deleting the broken file, it's astounding to me that the system's behavior is BSOD. To me, that's more damning that a bad "software update". These files seem to change often and frequently. Given they're critical path, they shouldn't have the ability to completely crash the system.
- Analemma_ 2y agoThat’s the danger of running in kernel mode. I’ve seen some people claim this is because the bad file starts a chain of events which concludes in trying to page an unpageable file, which is an application crash in user space but brings down the whole system if it happens in the kernel.
- SkyPuncher 2y agoThat seems like programming 101 for these systems. In the past, I've worked around this by validating the configuration of a file before attempting to run it. You bail out in a safe way during validation, but still allow a hard error during run time. Doesn't prevent all misconfigured files, but prevents the stuff like.
- YZF 2y agoPerfect example of where instrumentation guided fuzzing like AFL would almost certainly have found a problem. I agree with the amateur hour observation. But then most things seem to be.
- userbinator 2y agoI think all programmers should have the experience of using and developing on a single-address-space OS with absolutely no protections like DOS, just to encourage them to improve their skills at writing better, actually correct code. When the smallest bugs will crash your system and cause you to lose work, you tend to be a lot more careful with thinking about what your code does instead of just running it to see what happens.
- notepad0x90 2y agoThis "channel file" is equivalent to an AV signature file. Crowdstrike is the company, the product here is "Falcon" which does behavioral monitoring of processes both on the device and using logs collected from the device in the cloud. I can see your perspective, but you should consider this: They protect these many companies, industries and even countries at such a global scale and you haven't even heard of them in the last 15 years of their operation until this one outage. You can't take days testing gradual roll outs for this type of content, because that's how long customers are left unprotected by that content. Although the root cause is on the channel files, I feel like the driver that processes them should have been able to handle the "logic bug" in question so we'll find out more over time I guess. For example, with windows defender which runs on virtually all windows systems, the signature updates on billions of devices are pushed immediately (with exception to enterprise systems, but even then there is usually not much testing on signature files themselves, if at all). As far as the devops process Crowdstrike uses to test the channel files, I think it's best to leave commentary on that to actual insiders but these updates happen several times a day sometimes and get pushed to every Crowdstrike customer.
- zug_zug 2y ago>> You can't take days testing gradual roll outs for this type of content, because that's how long customers are left unprotected by that content. If you can't take days to do it then do a gradual rollout in hours. It's not a high bar.
- notepad0x90 2y agothey reverted it after about one hour. but sure, they didn't need to target all customers all at once, that's a good point.
- lotr5 2y agothey are dumb enough to process their "channel files" in kernel, this should be only done in usermode
- 2y ago
- jefurii 2y agoI have a friend who is a security guard at a bank in Hollywood, CA, who told me the computers at his location started going down between 12:00 and 13:00PDT (19:00-20:00UTC). I don't understand CrowdStrike's rollout system, but given that people started seeing trouble earlier in the day, surely by that time they could have shut down the servers that were serving the updates, or something?? He also told me that soon after that the street outside the bank (another bank across the street, a hospital several blocks down) was lined with police who started barring entry to the buildings unless people had bank cards. By the time I woke up this morning technical people already knew basically what was going on, but I really underestimated how freaked out the average person must have been today.