5 ms·
Hypothesizing about how Backblaze stores data: because they likely keep a series of backups, it's more likely they write each delta out separately. I'd imagine
by prawks 12y ago
Hypothesizing about how Backblaze stores data: because they likely keep a series of backups, it's more likely they write each delta out separately. I'd imagine it's likely they no not have whole copies of your latest backup on their drives, but they could reconstruct one at will if needed. If that were the case, it would indeed be close to write-once.
- brianwski 12y agoBackblaze engineer here. If a file is less than 30 MBytes and it changes even 1 byte we push an entire copy to a new location in our datacenter. We preserve the old copies for 30 days so you can "roll back time" up to 30 days (like if you edit a document but want to revert). For files larger than 30 MBytes we break them into 10 MByte chunks and only transmit the chunks that have changed. So the worst case is you insert 1 byte at the very start of the large file - this effectively changes EVERY 10 MByte chunk and we transmit it all. The best case is you append 1 byte to the end of the large file, because then we only have to transmit that one chunk.
- existencebox 12y agoI'm very curious; has a scheme that attempts to optimize against that worst case been considered? (Assuming you guys have, was it not used because the net tradeoff wasn't worth it?) It seems like if you're already going through the effort to see if chunks have changed, you can readily do some heuristic where you check the end chunk of both sides, and scan inwards until you find a change; something to that respect, further minimize replicated data? I'd assume you might have CPU cycles to spare, but I'm too far into assuming already to feel comfortable so I'll just hope you answer :P (thanks in advance, it's been great to read what you've written thus far.)
- brianwski 12y ago> have you tried a better binary diff algorithm We just never got around to it. In practice, it turns out most well written programs with large data try not to insert one byte at the start of the file. For example, take your large Outlook "pst" file (commonly 1 - 4 GBytes). When you get an email, it seems to append it to the END of the pst file, plus update some internal tables. So a large amount of that won't change. Also, the worst case is it wastes space in our datacenter (and some bandwidth) for 30 days and then is cleaned up anyway, so you can measure the theoretical amount of money to save and it won't profoundly change our business so we put it off. Not to say ANYTHING is off the table, we're just always swamped with some project or other. :-)