7 ms·
Disclaimer: I work at Backblaze. > I wonder if Backblaze does anything similar to provision drives. We do different things for different drives, but ALWAYS ha
by brianwski 7y ago
Disclaimer: I work at Backblaze.
> I wonder if Backblaze does anything similar to provision drives.
We do different things for different drives, but ALWAYS have a "burn in" period for new vaults (group of 20 computers that files are Reed-Solomon encoded across) that come online.
When I say "different things", when we deployed a vault full of recent Toshiba (?) drives it came up "too slow" and we figured out the OS block size and drive block size wasn't lined up "by default". It "works" but essentially a single write required reading two blocks and then writing two blocks. (sigh) I always wonder how many mis-configured PCs exist in the world with little problems like this quietly slowing down some user who never figures it out.
- VectorLock 7y ago>I always wonder how many mis-configured PCs exist in the world with little problems like this quietly slowing down some user who never figures it out. All of them.
- deleted 7y ago[deleted]
- devwastaken 7y agoHow would one go about finding the right block size?
- brianwski 7y ago> How would one go about finding the right block size? From our IT guy who figured out the problem...... "If it's 4K native, then a multiple of 4K is a good starting point. If it's not then a multiple of 512 is ok. If it's SMR based then it's more complicated and depends on whether it's host managed or drive managed and the number of conventional PMR zones. The second part of the equation is alignment. If there's a partition table or logical volume mapping, then the partition or logical volume's offset needs to be a multiple of the native geometry otherwise the block size doesn't really matter because everything will be misaligned but it might be possible to use math to figure out a block size that cancels out the misalignment." Sooooo..... consumers are hopelessly screwed and better hope Dell or Apple got it correct out of the factory. More info from other IT guys....... "These days I think parted gets it right most of the time, at least for individual spinning drives. SSDs are a bit weird because they tend to not want to share the erase block size, but I think the recent-gen ones tend to have enough firmware magic that it's OK. ... parted also has a flag for alignment type (--align optimal) which is supposed to choose an optimal multiple of physical block size."
- BearOso 7y agoI read a while ago that Samsung’s TLC erase unit size was 3MB, so I like aligning to multiples of 12MB. This covers 1.5, 2, 3, and 4MB. No one will miss a few megabytes, and you’re pretty much guaranteed that things are aligned, no matter what the underlying format is.
- AceyMan 7y agoDid not know: thanks for sharing. I may use 12MB going forward, myself (rather than the "1mb covers all ur bases" I cited in my upstream comment).
- AceyMan 7y agoMicrosoft changed the default offset for volume creation to 1Mb (2^20 bytes) a while back, around Win7(? SP1)/2012R2 or so. 2^20 bytes lines up with any known disk sectoring at the expense of "wasting" 1 mb, which wrt today's capacities is trivial. If you want to declare a smaller offset, you can, but absent an explicit value, one (binary) mb is what you get.
- livueta 7y agoArgh, I got bit by a similar issue a while ago when I was expanding a zpool. The pool was originally created using 512b drives. After in-place subbing each lower capacity old drive for a new, bigger (and also 4k) drive and resilvering, I realized that the zpool was running with emulated 512b sectors. Since freebsd zfs doesn't let you modify ashift after pool creation, I was stuck. However, I had some benchmarks laying around and was able to determine that the perf hit wasn't that bad for my workflow (<10%). So, I figure I'll just live with it until I eventually have enough disk to create a new pool with correct ashift and be able to copy all the data on the missized pool without taking anything offline/going under-protection. Definitely not the sort of problem I expected to encounter during that project, though, so I can totally see how even you guys would hit stuff like that.
- dbg31415 7y agoWhen you find issues like this, is it a manual detection process? Or do you have a set of tools? A set of tools that could be distributed to the public? (= You could save the world! Ha.