5 ms·
> During a routine update to the critical elements snapshot data, an incomplete snapshot was inadvertently shared which removed several sites from the topology
by ccooffee 4y ago
> During a routine update to the critical elements snapshot data, an incomplete snapshot was inadvertently shared which removed several sites from the topology map.
I wish this went into more detail about how an incomplete snapshot was created and how the incomplete snapshot was valid-enough to sort-of work.
I'm supposing that whatever interchange format was in use does not have any "END" delimiters (e.g. closing quotes/braces), nor any checksumming to ensure the entire message was delivered. I'm mildly surprised that there wasn't a failsafe to prevent automatically replacing a currently-in-use snapshot with one that lacks many of the services. (Adding a "type 'yes, I mean this'" user interaction widget is my preferred approach to avoid this class of problem in admin interfaces.)
- bragr 4y ago>I'm supposing that whatever interchange format was in use does not have any "END" delimiters (e.g. closing quotes/braces), nor any checksumming to ensure the entire message was delivered. Those only ensure you get the whole message, not that the message makes sense.
- ccooffee 4y agoQuite true. I was assuming that the incomplete snapshot was a transmission error or storage error. It's quite possible that the bug was an error of omission inside the data itself (e.g. someone accidentally removed an important key-value mapping from the data generator).
- jgrahamc 4y agoOne of the first things I did at Cloudflare was change the format of a file that contained vital information about the mapping between IPs and zones (no longer used these days) so that it had something like: START <count> <sha1> . . . END because it was just slurped up line by line assuming EOF was good.
- sitkack 4y agoEvery mission critical file eventually gets version numbers and checksums.
- kridsdale1 4y agoAnd eventually a blockchain
- smt88 4y agoFiles can be immutable, cryptographically-verifiable, and distributed across servers without a blockchain. Adding a blockchain to this would slow things down, add surface area for errors, and provide absolutely no value. This is true of every usage of blockchain other than creating tokens.
- cheeselip420 4y agoone reason why JSON is superior to things like TOML or YAML for these use-cases...
- e12e 4y agoNot to worry, JSONL[1] fixes that ;) In all seriousness - just dropping 1gb json file with N million records in one end probably isn't great either. I suppose one could somehow marry js and subresource integrity protection[2] to get a json serialized structure with hash integrity check. It would probably be a terrible idea. [1] https://jsonlines.org/ https://jsonlines.org/ [2] https://developer.mozilla.org/en-US/docs/Web/Security/Subresource_Integrity https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
- ehPReth 4y ago> Text editing programs call the first line of a text file "line 1". The first value in a JSON Lines file should also be called "value 1". I wonder why not zero?
- 4y ago
- omoikane 4y agoEnd markers will help detect truncated files but not other kinds of brokenness, need checksums for those. Also, I think the config files for certain off the shelf network devices don't come with end markers or checksums, and might not be very good at handling corrupted configs. The usual practice is to push updates to a small fraction of those, then wait and see what happens before proceeding further.
- zamnos 4y agoUnfortunately even checksums won't help you if you mistakenly dredge up old but valid config. That doesn't mean they're not worth doing, but we don't know, due to a lack of detail, if Google already has checksums or if they would even have helped in this situation.
- londons_explore 4y agoGoogle internally avoids use of ascii files, so the type of error you are suggesting is unlikley. I suspect it was more a case of incorrect error handling in a loop... Eg. output = [] try: for site in sites: output += process(site) except: print("error!") write_to_file(output)
- jeffbee 4y agoOr it could have just been a recordio file that was being written by something that crashed in the middle of doing it, and the committed offset was at a valid end of record. Really there's 1000 ways for this to happen all of which are going to sound obvious in retrospect but are easy to commit.
- q3k 4y agoEven straight up Proto is vulnerable to this (either with an unlucky crash right between fields being actually written to a file, or when attempting to stream top-level fields).
- deleted 4y ago[deleted]
- amalcon 4y agoIt's more likely that the snapshot was incomplete in that it was based on incomplete precursors, rather than that the message itself was truncated. The details of something like that aren't always appropriate for this kind of report (i.e. they usually create more questions than they answer for a reader unfamiliar with the basics of the system).
- dekhn 4y agoNo. That's not how it would be done at google. they invented a binary protocol to handle things like this reliably. And that protocol goes over an integrity-checking network transport. It's more likely an odd edge case occurred and somehow got past the normal QC checks- say, an RPC returned an error but the error handler didn't do a retry, it just fell through.
- spullara 4y agoI ran into a problem like this with a service that used YAML for their config file. Basically when I edited it and saved the service would automatically pick up the change and load the config. However, the save hadn't completed so it only read a partial file which was still valid because, YAML.
- unxdfa 4y agoI still like XML and XSD. People look at me these days like I'm insane. But in this case a partially loaded XML document would not parse let alone pass schema validation. Again vindicated. YAML needs to go away. It is misery. I'd rather have XSLT than Helm templates.
- CodesInChaos 4y agoIt's crazy how we still don't have the ability to atomically create/replace a file on the popular operating systems and filesystems , despite that being a super common use-case (without workarounds like renaming which have annoying limitations).
- mhandley 4y agoWhat would you want from an atomic replace of the contents of a file that rename() doesn't provide? Surely you've got to write the new contents somewhere before calling your atomic replace, and that's exactly what happens with rename/mv.
- ikiris 4y agoYou're assuming it just stopped abruptly instead of just was missing data.