11 ms·
Meta quickly detects silent data corruptions at scale
- londons_explore 5y agoIn a fleet of 100,000 machines, there will always be some clear failures... When the machine has 2x the number of segfaults of any other machine in the fleet, you send it for repairs and someone replaces the motherboard, ram and CPU... easy! But the painful ones are the 'subtle' failures. Why does machine PABL12 sometimes give NaN as a result while all 99,999 machines return sensible numbers? But all the burn in hardware tests pass... The solution was to simply exclude any machines that were outliers. Anything in the top or bottom 0.01% for any metric simply exclude that machine from future workloads. Sure, in most cases there was nothing wrong with the hardware, but when you're spending hours debugging some fault caused by a sometimes-bad floating point unit on one core of one machine out of 100,000, you're just wasting your time. By auto-banning outliers, the machine will end up doing some other task where data consistency matters less.
- notacoward 5y ago> When the machine has 2x the number of segfaults of any other machine in the fleet, you send it for repairs At that scale, it's quite likely sent to repair automatically and whoever's on call just gets a notification.
- jeffbee 5y agoWas pabl12 an actual bad machine? Sounds somehow plausible, as if I heard of it before. It was an annoying struggle trying to raise the visibility of broken CPUs during my years at Google SRE. The SRE org and the rest of the software side of Tech Infra resisted the whole concept, even though it was well-known among platforms hardware eng. The process for taking a known-bad machine out of service involved 1) the machine being reported independently by three different teams; 2) the machine continuing to be in service for days or weeks, at the leisure of some very asynchronous automation; and 3) the machine being returned immediately to service because it passed all of the cursory checks during reinstall. Really irritating. Consequently every major service had to maintain their own private blacklist. It's nice to see that some influential people on the software side are starting to come around, with papers like "Cores That Don't Count" etc, but man they could have been on this boat a decade ago.
- mjevans 5y agoReminds me of the typical story of someone with a complete damage protection plan and a flaky device. Take it in for repairs, passes all the tests, but they know it's funky, so snap it in half or otherwise completely wreck it right in front of the tech and demand that repair.
- bryan_w 5y agoUsually teams would consider a machine "bad" if that node in the cluster had elevated errors compared to the rest of the cluster they were running. Unfortunately this doesn't tell hardware teams what actually went wrong. If one could show that the CPU said 2+2=9, I'm sure they would yank it out right away, but "it returns 500 errors a lot" isn't very debugable. The only thing they can do is run the diag and return it to service if nothing comes up.
- jeffbee 5y agoWell that's one of the reasons this is difficult to handle as an organization. The novice says "the machine is broken" and is mistaken. But the expert says the same thing, and is correct. Same with compiler bugs: novices believe the compiler is full of bugs, journeymen believe the compiler is infallible, but the wise return to the knowledge that the compiler is full of bugs. Maybe that company just needs "bad machine readability" or something. And your last statement is definitely not true. I can recall multiple instances of demonstrable logic errors in which the machine repeatedly returned to service. This includes all of the machines of a certain generation of a certain vendor's CPUs that were found to have latent ALU bugs, 8 years after going into service.
- jgrahamc 5y agoSome might enjoy this old Cloudflare debugging story about random crashes in production. https://blog.cloudflare.com/however-improbable-the-story-of-a-processor-bug/ https://blog.cloudflare.com/however-improbable-the-story-of-...
- dfdz 5y agoThanks for sharing.
- ignoramous 5y agoAdd to that a bunch of "rare" / "unlikely" / "silent" CPU bugs (compute errors) that Google and Facebook see with regularity: https://muratbuffalo.blogspot.com/2021/06/cores-that-dont-count.html https://muratbuffalo.blogspot.com/2021/06/cores-that-dont-co... > So Google found fail-silent Corruption Execution Errors (CEEs) at CPU/cores. This is interesting because we thought tested CPUs do not have logic errors, and if they had an error it would be a fail-stop or at least fail-noisy hardware errors triggering machine checks. Previously we had known about fail-silent storage and network errors due to bit flips, but the CEEs are new because they are computation errors. While it is easy to detect data corruption due to bit flips, it is hard to detect CEEs because they are rare and require expensive methods to detect/correct in real-time. https://muratbuffalo.blogspot.com/2021/06/silent-data-corruptions-at-scale.html https://muratbuffalo.blogspot.com/2021/06/silent-data-corrup... > The paper claims that silent data corruptions can occur due to device characteristics and are repeatable at scale. They observed that these failures are reproducible and not transient. Then, how come did these CPUs pass the quality control tests by the chip producers? In soft-error based fault injection studies by chip producers, CPU CEEs are evaluated to be a one in a million occurrence, not 1 in 1000 observed at deployment at Facebook and Google... The paper also says that increased density, technology scaling, and wider datapaths increase the probability of silent errors.
- notacoward 5y agoTo be clear, this is about corruption in the CPU/GPU/memory complex. There's a whole separate set of techniques (some of which I worked on) to detect and correct data corruption on disk.
- huhtenberg 5y agoI'm in the same boat and my takeaway is that the vast majority of a "silent" on-disk corruption actually happens on the way to the storage, i.e. the data gets corrupted in some RAM it passes through and then just ends up being written out in corrupted state. This is because, virtually all modern drives implement per-sector FEC coding, so if a bit does flip on the disk, you will either get back original data (now FEC-corrected) or you will get a read error. That is, the so-called "bitrot" phenomenon is largely mis-attributed. Bitrot doesn't happen at rest. It happens in transit.
- notacoward 5y agoI can state categorically that bitrot on disk does exist, because that's one of the parts I worked on. It's pretty rare - unfortunately I don't think I can give you the numbers - but across enough exabytes it does happen enough to justify slow scans to detect it.
- huhtenberg 5y agoHow did you know it was a change at rest? The only correct way to test for bitrot is to read the data back immediately after it was written and the cache flushed. If it's the same as the original, we know it made it to the disk undamaged. Then re-read it again after some time. If it doesn't match, re-read immediately again, ideally using a different physical memory block. Compare again. If it doesn't match, take the disk to another machine and re-read again. If it doesn't match, only then it's an actual at-rest bitrot... OR it's a drive's firmware bug, because corrupted data must be corrected or it must not be returned at all.
- notacoward 5y ago
- HL33tibCe7 5y agoCompletely off-topic digression: I still think the name change to “Meta” is a big mistake. Subjectively, for some reason I just really dislike the name. More objectively, the branding is very muddled, e.g: serving an “Engineering at Meta” blog post on fb.com. Often with these things it’s just about time; it feels wrong because you’re just not used to the change yet. Maybe that will happen, but it’s been months now. Usually with these changes I change my mind quicker than that.
- ATsch 5y agoI feel like, given the negative connotations of "Facebook", that's by design.
- nwsm 5y ago> the name change to "Meta" is a big mistake I think it's too soon to tell. Facebook has really negative brand recognition (from my POV), and who knows, maybe "metaverse" style online interaction is the future. (For the record I'm anti-web3 and indifferent on metaverse communities)
- CiPHPerCoder 5y agoI will always say VR, I will never say "metaverse". Their branding move was bold, yet unconvincing.
- tmn 5y agoVr is a subset of the ‘metaverse’. The metaverse isn’t really something new. It’s just a rebranding of the portion of our lives that are contained within the digital realm. On top of that there are obviously ideas for how to adapt and grow that space, which is all to be seen
- tonguez 5y ago“It’s just a rebranding of the portion of our lives that are contained within the digital realm.” none of my life is “contained” within shitbook
- kache_ 5y agoThe scale at which Meta operates at really boggles my mind. I work with an ex facebook guy who was on the infra side of things and the numbers he told me.. I couldn't even imagine. And I'm working on the order of magnitude of 100m/h, but still, completely different set of challenges.
- silisili 5y agoSame. I remember asking one guy at FB the process to ask for a new server. He said he can't even open a request for anything less than a thousand boxes. The largest fleet I'd worked on at that point was 12... different worlds.
- lclarkmichalek 5y agoI mean, that's not true in the general case. That'd be incredibly wasteful. (Work at Meta, mostly on capacity)
- mescaline 5y ago
- ryeguy 5y agoThis is a technical post, stay on topic and stop posting flamewar bait.
- mescaline 5y ago
- ethanwillis 5y agoIt's not flamewar bait unless you turn it into that. It can be a fruitful discussion.
- 5y ago
- mad44 5y agohttps://muratbuffalo.blogspot.com/2021/06/silent-data-corruptions-at-scale.html https://muratbuffalo.blogspot.com/2021/06/silent-data-corrup...
- tupac_speedrap 5y agoContent seems interesting but the generic corporate image at the top, crap font and off-black low contrast text colour is getting on my nerves.
- throw03172019 5y agoReader mode works great on mobile Safari.
- Hnrobert42 5y agoInterestingly, this site fails ungracefully (HTTP error code 500) when I try to visit from NordVPN, even after cycling through a few IP addresses. I’m noticing more and more sites block all VPN track. I get why, but it’s not good.
- doix 5y agoDepending on why you're using a VPN, you can just pay for a tiny VPS from ovh/hetzner, setup wireguard and use that as your VPN. Obviously don't do anything illegal, since everything is going through a server that is directly tied your credit card. But for privacy/security, it's good enough (for me anyway). I'm guessing it's luck of the draw if your IP has been blacklisted or not, but I've not had any issues the last 6 months whilst I travelled around.
- wiredfool 5y agoHetzner commonly triggers cloudflare's bot detection, and there are some things that just refuse to talk to it’s ip space.
- chaorace 5y agoI've noticed this is quite popular among the kings of cargo-cult security: banking websites. I can only hope the proliferation of VPN-gating is more contained compared to the recent (banking-led) upswing in Android root-checks. This type of security theater can be easily bypassed by any determined attacker and thus only serves to deter honest users.
- someotherperson 5y ago> This type of security theater can be easily bypassed by any determined attacker and thus only serves to deter honest users. To play devil's advocate, the large amount of attackers aren't really determined. They're just fishing for easy targets. If you check the logs on a VPS you'll see an endless stream of people trying to exploit things like Wordpress 24/7 on your brand new VPS that has nothing but a html landing page. With banks, I imagine they have a compliance check list they have to tick off to make sure that -- if and when a successful attack happens -- their insurance would pay out. If they haven't taken simple steps like blocking VPNs it could lead to the insurance company claiming negligence.
- raphaelj 5y agoIt would be better if Meta would focus on detecting spam at scale. I put a desk chair on Marketplace last Friday, and got 8 messages that were actually scams. These were trying to "schedule" a Fedex/DHL pickup, and would redirect me to fake branded websites that were requesting my personal details and bank account. This was so obviously fake it baffled me Meta can't detect these automatically. I am also getting multiple message requests per week asking from hookups. These are obviously fake [1]. --- [1] https://imgur.com/a/yZDPh3C https://imgur.com/a/yZDPh3C
- monkeybutton 5y agoSomehow I knew it was going to be a bit.ly link before opening the image
- ebbp 5y agoIt’s a different team, with a different skillset, that would be responsible for that. Big companies can focus on more than one thing at a time.
- deleted 5y ago[deleted]
- zitterbewegung 5y agoI think that large tech companies giving a snapshot of what cool or interesting things they do is great but if there are bigger problems that don’t seem to have that kind of focus it just feels like a marketing / recruiting post (which isn’t that bad). But, the problem would be if they made public antispam systems they can’t give that to spammers which presents as a catch22. Also if you have humans in the loop to evade a spam system it is basically impossible .
- deleted 5y ago[deleted]
- BbzzbB 5y agoThey ban like 1.7B account per quarter ignoring those blocked at registration. Isn't that focus? Subjectively too I also see so much less bot activity on Facebook than I do on any other social media.
- PTOB 5y agoI work on the physical side; building hyperscale datacenters. You guys should try your hand at managing errors in that system. You've got it all: memory leaks, thermal overloads, misallocated heaps, pipes with strong type requirements, dropped packets ... you name it.
- Melatonic 5y agoI would probably be overwhelmed just managing the infrastructure for your monitoring systems and infrastructure is my main thing :-D
- RoboTeddy 5y agoComputational proofs of integrity (STARKs, SNARKs) could detect silent data corruptions (at the cost of a ~1000x slowdown) I wonder if we’ll see them used for large scale applications whose correctness is critical.
- pnw 5y agoThe first thing I noticed about this article is that like all Facebook pages, it silently corrupted my back button.
- mtVessel 5y agoDid you open it in Firefox? If so, that's the fault of Facebook Container, not the page.