3 ms·
This is a typical case of poor administration at a company blaming a vendor for their problems: 1. Yes, any storage array should work to first protect the inte
by mike_mccracken 15y ago
This is a typical case of poor administration at a company blaming a vendor for their problems:
1. Yes, any storage array should work to first protect the integrity of the data
2. Yes, a good DR solution is only as good as the testing and tuning that the administrators have done to ensure it is working properly
Here are some questions:
- How many spare drives were in the system when the first and second drive failed? Netapp does not shut down volumes of storage if spares are in the system to take over for the failed drives.
- How long was it really before those drives were replaced with a spare that could take over and rebuild?
- Why don't you publish your DR plan and explain exactly where it didn't go as tested and planned? Point this out to show where the issue occurred that had been previously tested and shown to work properly.
I am not an employee of NetApp and not even a customer of NetApp. I have used it in the past and like all technologies it requires the care and feeding that is well documented in manuals they provide. And it requires good administrators to do their jobs to test and monitor things and ensure the resources are available for the system to work as designed.
Once you have heard both sides of the story the only thing you learn is that there are more the 2 sides to the story.
When these types of things happen the folks closest to them always leave out details or cannibalize the story so that they can be found blameless. I have been in the IT field for 15 years and seen it time and time again (with many technologies).
I read an email from Grasshopper this AM detailing to their customers what this issue was. It was so vague and left so much to interpretation that it really came across as whiny and misinformed. It was extremely unprofessional to apologize and then blame (without full explanation or root cause).
There's more to this than meets the eye. Believe me.
Since you all are replacing NetApp, I would suggest paying for a full time engineer from the next storage company you buy from. They can manage the array for you (properly) and ensure these things don't happen. Otherwise you'll need your sysadmins to start reading product documentation, following best practices, and testing procedures.
- dh 15y agoIf you want to talk about any of the details I am happy to discuss. I never said that the array went offline because of the failure but the head was under very heavy load trying to recover from it. The email is a careful balance of information that 90% of people will find useful and not too much information that no one understands it. Never once did we say we are not to blame, actually the opposite, it is our responsibility no matter the vendor or what we replace the hardware with.