4 ms·
The Shingled magnetic recording NAS debacle
- sokoloff 6y agoIn the late 90s, Western Digital drives had a terrible reputation (well-deserved). Just 15 years later, people were very positively inclined towards WD. I think WD bean counters are well justified in thinking “even if we get caught here, this too shall pass...” Shortly after WD starts shipping competitive products for 5% less than anyone else, the market will adapt and most will forget the SMR debacle, just like we forgot the tragically poor quality drives of the late 90s.
- AtlasBarfed 6y agoExcept spinning rust is staring at the prospect of flash storage eating their lunch. Flash isn't THAT far away. If these increased areal density techs don't pan out, they are done.
- zozbot234 6y agoFlash storage is not nearly as reliable as spinning rust, given comparable density. The spinning rust niche is pretty much unassailable, and now SMR lets the industry expand into tape-like use cases. What WD did was a huge screwup for the reasons OP talks about, and they had to backpedal quickly.
- bcrl 6y agoI disagree. Flash storage used within the range of its capabilities is more reliable than spinning rust. Failure of flash is quite predictable and can be managed through the use of over provisioning and advanced error correction codes. Flash failures tend to be localized and directly related to wear caused by writes. This can be managed by properly written firmware. In contrast, traditional HDDs are mechanical devices. Total failure of a hard drive tends to happen unpredictably. A poorly time impact can result in damage that is fatal for the entire device. Once debris starts bouncing around inside of an HDD, damage can be catastrophic. I will agree that early flash drives were not as reliable as they should have been, a problem is that mostly solved in SSDs produced today. Software issues in SSDs were a problem in the early days -- I even had an early Intel 32GB SLC SSD that bricked itself even though it was nowhere near the end of its rated life. It takes time for firmware developers to develop an understanding of the problem space and create a sufficiently advanced test bed to exercise some of the more complicated corner cases associated with wear related issues near the end of life. SSDs today are a very different beast than those on the market 10 years ago. Lately I've been getting some exposure to the archival problem space. Event today, tapes are still an important part of the storage hierachy, especially when it comes to legal compliance by certain large businesses.
- zozbot234 6y agoYou can choose SLC (or limited MLC) and overprovisioning to make flash storage quite reliable, and perhaps even preferable to disk in some respects. But that comes with severe compromises in density and you're still vulnerable to bugs in the opaque, proprietary Flash Translation Layer, which has no equivalent in (non-SMR) hard disks.
- toast0 6y ago> Software issues in SSDs were a problem in the early days -- I even had an early Intel 32GB SLC SSD that bricked itself even though it was nowhere near the end of its rated life. They're still a problem. HP had two separate instances of branded SSDs fail when the lifetime counter rolled over (I guess they didn't test everything the first time). My experience with running around 10k SSDs and spinning drives in a datacenter was that spinning drives would fail more often, but would provide clear warning signs if you check SMART stats regularly, providing opportunity to replace failing drives without losing redundancy if your storage systems are flexible. SSDs would fail significantly less often, but there would rarely be a warning --- disks would just disappear, never to be seen again. Our use was relatively low writes, we weren't anywhere near the lifetime write counts. Some SSDs would really tank performance if they declared bad sectors, while they were moving things around, host driven I/O was very slow. I don't remember which brand that was.
- bcrl 6y agoIf I were buying SSDs, HP is not at the top of my list of well known SSD vendors that have an excellent data integrity record, probably owing to an experience many years ago with HP branded MegaRAIDs that only had a known data corruptor in the HP "blessed" firmware. My personal theory is that consumer gear tends to get better tested over the long run since millions of units out in the field will result in a lot of weird corner cases actually happening. The electrical engineers I know love it when their products are in high volume production, as it makes it a lot easier to find patterns and the underlying causes of failures. In contrast, low volume niche products will always have lurking bugs deep in obscure corners that haven't been exercised.
- rasz 6y ago
- tssva 6y agoUnless they have changed their response in the last couple of weeks WD hasn't done any backpedaling. They instead just clarified which drives use SMR and then issued a statement trying to gaslight buyers of their NAS drives that this issue occured because they purchased the wrong type of drive for their use case. You know purchasing a drive marketed as a NAS drive in a NAS.
- derekp7 6y agoAlso, the other two manufactures are apparently doing similar -- shipping SMR drives without any indication of what they are. At least one of them released a statement that said they weren't doing SMR on drives labeled as enterprise san or something similar, but you still take a chance if getting one of the consumer drives (and for home storage, a consumer drive may make sense, as long as it isn't SMR).
- nicolaslem 6y ago> So it came as a surprise when sysadmins began noticing their new Western Digital Red NAS drives were dropping out of NAS RAIDs and ZFS pools owing to random write timeouts and failures. I experienced exactly that a few days ago. I have some of those SMR WD drives in my NAS and one drive just disappeared from the ZFS pool while importing a large database. I rebooted and the drive was back. ZFS worked its magic and everything seems fine now, but it doesn't inspire confidence.
- phire 6y agoIt's quite clear that WD never properly tested these SMR drives in actual 3-8 bay NAS workloads. Certainly not a rebuild. Based on their initial response to this scandal, it seems at some point they defined the "small NAS" Red spec as 180 TB/year read/written per year. Then they later certified the new SMR drives as meeting this 180 TB/year spec, probably by testing they were good for 500gb/day. They completely ignored the fact that larger home NAS installs don't use their 180 TB/year spec at a uniform rate, but instead tear though several TB doing a rebuild then sit near idle for most of the rest of the year.
- zozbot234 6y agoNo NAS uses that spec at a uniform rate. Practically every NAS these days uses RAID, which implies rebuilds.
- mrlonglong 6y agoGoddamn MBA types making a quick buck to get their bonus. This never works out too well for the company involved. Boeing for example.
- tambeb 6y agoTook the words right out of my mouth.
- genr8 6y agoThis is a warning sign that we have 0 agency over hard drives we think we own. The corporations control the actual tech being pushed, to pull a bait and switch. They also control the firmware, which is super proprietary and hidden, and can be subverted and controlled without your knowledge. The end goal is to push consumers to end up storing all their data in the cloud.
- donavanm 6y ago> The end goal is to push consumers to end up storing all their data in the cloud. Do you really think seagate, toshiba, and WD want you to move your data to the cloud? They dont have any particular commercial interest there. Theyre just another set of commodity input providers to the Amazon/Google/Microsoft supply chains. The same as they are to dell/synnex/foxconn/lenovo when the drive is put in a different chassis and sold to you. Wheres the commercial motivation? Personally, if I was a commodity supplier, Id much rather sell to you with a wholesale markup than negotiate with Amazon grinding out pennies from the price.
- rasz 6y agoScam, not debacle.
- LargoLasskhyfv 6y agoRelevant material to think about here: Sprites mods HDD-hack [1] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=Sprites%20mods%20hdd&sort=byPopularity&type=all https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... Andrew "bunnie" Huang's presentation at the 30c3 Chaos Communication Congress from 2013 about the vagaries of flash memory of all sorts... [2] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=30c3%20-%20The%20Exploration%20and%20Exploitation%20of%20an%20SD%20Memory%20Card&sort=byPopularity&type=all https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... The OpenSSD Project [3] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=Open-SSD&sort=byPopularity&type=all https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... and just from a few days ago about the remake of their website? [4] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=http%3A%2F%2Fwww.openssd.io%2F&sort=byPopularity&type=all https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... and then remembering this: [5] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=western%20digital%20risc-v&sort=byPopularity&type=all https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... and this [6] https://www.westerndigital.com/company/innovations/risc-v https://www.westerndigital.com/company/innovations/risc-v
- bestham 6y agoThe funny thing is that SMR is perfect for storing large immutable files like photos, music, videos or game resources. What we need is a controlled way of relaying the SMR properties further up the chain. Anything that is “drive managed” is bound to fall apart sooner rather than later. I have the first Seagate 8TB SMR “Archive” HDD, and I have yet to find a way to tell the controller to “TRIM” the complete drive and start over. Without this SMR-drives essentially become WORM-drives. The “read-modify-write” nature of copy-on-write file systems are a perfect marriage with “host managed SMR”. There was a couple of sessions on SMR during the OpenZFS summit: “Host aware SMR” [1] and “HGST SMR” [2] [1] : https://openzfs.org/w/images/2/2a/Host-Aware_SMR-Tim_Feldman.pdf https://openzfs.org/w/images/2/2a/Host-Aware_SMR-Tim_Feldman... and https://youtu.be/b1yqjV8qemU https://youtu.be/b1yqjV8qemU [2] : https://youtu.be/a2lnMxMUxyc https://youtu.be/a2lnMxMUxyc
- projektfu 6y agoThe manufacturers could have ushered in a revolution in filesystems that would have taken full advantage of their new drives. If they had taken an approach like in "Venti" (Plan 9), and provided an underlying log-structured block archiving service, while placing SSDs in the front, they could have sold new appliances that never forget a change and respond immediately. Running out of archival space? Just plug in another SMR drive. A machine gets encrypted by malware? No problem, all the prior data is still on the archival drive. Need to compare this file to a version from 3 years ago? The backup was automatically saved. A lot of the benefits are there in something like Time Machine, but this system is even more automated and doesn't require explicit backup steps every hour or day, nor does it need to remove backups as they get old.