4 ms·
I used to use 2048, that's probably the wrong value to use. want to use something that corresponds to the ECC block size of the underlying media. but yes, I d
by compsciphd 6y ago
I used to use 2048, that's probably the wrong value to use. want to use something that corresponds to the ECC block size of the underlying media. but yes, I do that as well (came to say mostly the same comtent, but I'd do r -1)
- colejohnson66 6y agoIs the ECC available to the OS? I thought the drive handled the ECC and just reported errors?
- wtallis 6y agoIf you're trying to read with a smaller block size than the media's native block size, then you'll make two or more attempts to read data from each corrupted block that's unrecoverable—making your recovery process much longer. If you try to read with a larger block size than the media's native block size, you'll get errors for chunks of that size even when part of the data may have been recoverable. The above is true whether or not the OS has access to the media's raw ECC data.
- compsciphd 6y agoit's not just that. Reading scratched blocks means sometimes it might work, sometimes it might not. lets say 1% of the time you get a good read. If the ECC block is 16k and you always read in 16k blocks, if you get that 1% magic time that the read succeeds, you got the whole block. If however, you read 2k blocks, you need to have the same luck 8x times. From experience of recovering bad DVDs and BRs, there were discs I had to pass through 100+ times to get a valid read on all blocks. This is also because, the underlying hardware will always do a full ecc block read (only way for it to determine that it read the block correctly, to read whole block and verify it), so any smaller reads are pointless.
- compsciphd 6y agojust in case you didn't read my response below, its not about the ECC being available to the OS, the OS doesn't need to see it. the way optical media works is that the optical media reads bits from the drive in ECC block size and then verifies / fixes the block and passes that back to the OS if its has a valid block, otherwise returns an error to the OS. hence, my logic that I describe below. optical media is an unreliable medium in general and hence depends on the ECC codes to ensure blocks are read correctly and they are used a lot. back in the day of CD and DVD burning there were fancier burners that provided apis for reading the error correcting stats into user space (i.e. how many of different types of errors were corrected), dont know if they still exist. It was never 0 across the board, but that's how the medium was designed, not to require it be 0 across the board.
- colejohnson66 6y agoWhat I guess I don’t understand is: why does the OS need to know the internal ECC block size if it doesn’t even see the ECC (or even know of its existence)? When I want a sector from my hard drive, I ask for 4096 bytes, not 4096+$ECC.[0] If I asked for 4096+$ECC, it would actually give me 4096 from the sector I requested, and $ECC from the next one. So why ask for 2048+$ECC and not just 2048? As you said, the nature of the medium requires ECC (side note: modern hard drives do too). So if I ask for a 2048 byte sector, the drive has to read the ECC. So why ask for more than that? It already knows the sector boundaries. In other words, if I tell `dd` to use a block size of 2048+$ECC, won’t that actually work a sector and a half (well, 1 + $ECC/2048) at a time? [0]: In fact, unlike CDs, I don’t even know or have any way of finding out how many ECC “bytes” there are in my hard drives' sectors
- compsciphd 6y agobecause the OS did't cache the full ecc block (it views the block size as 2048 or 4096 bytes), and with scratched media 2 reads of the same ecc block aren't going to necessarily both succeed. simplistic case, imagine we have 1 ECC block of 16k, but we read at 2k, so we'll number the 2k blocks 0-7 T0 - read block 0, fails T1 - read block 1, succeeds! T2 - read block 2, fails T3-T7 repeat for blocks 3-7, all fail in practice if we read a 16k bock at T1, we would be golden (and finished). Instead we did 8 steps, and only got 1/8 of the data. This is becaue the OS doesn't have a concept of the hardware's ECC block size, so the optical hardware in a sense virtualizes it, and the OS will just keep on rereading the same ECC block on the media and possibly continue to get errors.
- compsciphd 6y agoalso, you don't ask for 2488+ECC (or +1). the ECC is not visible to the OS or user. you just care about the size of data the ECC is protecting. if the ECC protects 16K or 32K of data, you want to read on those physical boundaries. as then you'll read a whole ECC block and it will either pass or fail. If it passes, you never have to try to read that ECC protected block again (and maybe fail). Of course, there is one hitch to my scheme. Ensuring that you always read on ECC protected block boundaries. I'm pretty sure if you use ddrescue to always read the right block size it will, but not 100% (why not? perhaps the ECC protects data not visible to the end user in some way (say the first block is only 8kb, not 16kb in practice). on the issue of hard drives, there is a lot more going on that puts you at the mercy of the firmware (relocatable sectors and the like). I did lose a RAID5 once (1 drive totaly died, and then in rebuild, another drive threw and error) , and I was able to use ddrescue to recover all but 4K block on it. as I was using a 128KB stripe size, that meant I probably lost somewhere between half MB and a MB of data - if the 4K was contained within a single stripe or not (probable it was). I was content with that. never did discover what data if at all was corrupted, but I was able to recover the raid5.