3 ms·
There is no such thing as a 'large commodity cluster.' You can order a pile of commodity parts and spend the next five years fiddling with it as you watch your
by khm 11y ago
There is no such thing as a 'large commodity cluster.' You can order a pile of commodity parts and spend the next five years fiddling with it as you watch your metrics fall into the toilet, or you can contract with a systems integrator with domain experience to make sure that you can actually accomplish the work you set out to do. When you're moving an established set of software from a known environment into a new cluster, it helps to have a wide array of technical experts and SLAs in place to ensure that the environment -- here meaning the physical compute infrastructure, storage service, and software installation -- is reliable and maintainable, so that your team can focus on their core mission -- in this case, the modeling software.
"Chaos monkey" works fine when you're dealing with random cloudstuff and you are in a position to duck-punch your production code. When you have a large dataset that you need to process reliably and predictably, the mantra shifts to "fail never", and that requires engineering up front in combination with carefully-planned maintenance.
Full disclosure: I've worked on this specific supercomputer. While I'm not on the admin team in question, I do work with them occasionally and I have a lot of respect for them -- their job is not an easy one.