4 ms·
I highly assume that they must have reached out and reported these issues to, presumably, Intel as the biggest player here. Likely they're just not disclosing t
by ffff1312 5y ago
I highly assume that they must have reached out and reported these issues to, presumably, Intel as the biggest player here. Likely they're just not disclosing these numbers, and generally not many are talking about these and potentially other CPU issue in public due to NDAs. Either way, it's quite amazing to dig down so deep in the production stack that you must conclude that it's the CPU at fault here. I presume academic research might have a hard time on this given the scale needed to run into these issues, but hopefully we'll see more on this research in future.
- scottlamb 5y ago> Either way, it's quite amazing to dig down so deep in the production stack that you must conclude that it's the CPU at fault here. Google's internal production stack is much more amenable to that kind of digging than public cloud products: * You can easily find out what machine a given borg task was running on. In fact, not just your own borg job but anyone's. You can query live state, or you can use Dremel to look up history. * Similarly, even as a client of Bigtable or Spanner, you can find out the specific tabletservers/spanservers operating on a portion of your database and what machines they're running on. (Not as easy to cross this layer and get to the relevant D servers actually storing the data but I think it's all checksummed here anyway.) If your team has your own partition, you can see tabletserver/spanserver debug logs yourself also. * There's a convenient frontend for looking up a bunch of diagnostic info for the machine, including failures of borg tasks (were other people's tasks crashing at the same time mine did? what was their crash message?), syslog-level stuff, other machine diagnostics like ECC / MCE errors, and repair history (swapped this DIMM, next attempt will swap this CPU). It's not unusual for application teams to suspect a machine and basically vote it off the island (I don't want my jobs running here anymore, I cast a vote for it to be repaired / Office Spaced). It's more rare for them to really take the time to really understand the problem in detail like "core 34 sometimes returns incorrect results on this computation", although there's nothing in particular stopping them from doing so (other than lack of expertise and a long list of other things to do). The platforms team gets involved sometimes and really digs in—iirc in one bug they mentioned sending a CPU back to the vendor to examine with an electron microscope. I'm not sure what lessons that offers for a public cloud where that kind of transparency isn't realistic...