4 ms·
The real mismatch is handling failures. The HPC world has grown up handling failures by failing the entire computation. As a defense against failure, they usual
by epaulson 12y ago
The real mismatch is handling failures. The HPC world has grown up handling failures by failing the entire computation. As a defense against failure, they usually checkpoint the entire computation.
Part of this is the fault of the MPI standardization process: MPI didn't even try to deal with dynamic-but-expected changes to the computation (much less changes caused by "node" failure) because not all of the initial systems MPI was targeting could handle nodes coming or going. In this way, MPI was sort of a step backwards - PVM was much more flexible than MPI from the get-go.
I haven't paid much attention to MPI in a few years, so maybe some of the MPI-FT approaches have caught on.
One of the initial targets for YARN was MPI, but it was never clear to me that any MPI users actually wanted to use YARN.
With just a few tweaks, PVM and AFS could have given us a pretty good Big Data stack 15 years ago (And in the case of security, a better stack than what we've got today.) The hardware wasn't quite cheap enough to make the case as compelling as it is today, and the MapReduce programming model helped sell the entire approach, but it was a real missed opportunity.
- TallGuyShort 12y agoI've discussed the MPI-on-YARN thing with quite a few people from both communities and really only found one use-case: MPI jobs you want to run with data in your Hadoop cluster, where it will be done so rarely it's not worth moving the data from your cluster to your supercomputer. Even then, I can't name an instance of this type of use case I've actually seen in the field.
- anonymousDan 12y agoOut of interest, what are the security issues you see with Big Data stacks?
- TallGuyShort 12y agoMost of the tools are written under the assumption that you're behind a firewall or airgap that will do all your security for you, and the controls are very weak and poorly enforced in the younger projects. Even in the mature ones, there's hasn't really been a cohesive model of security across the stack that meets the needs of stricter organizations. That's improving - I'm doing some work with Apache Sentry (incubating) that lets you apply more sophisticated security policies to fine-grained data in Hive, Impala and Solr (and in the future, hopefully many others). There are also projects like Apache Knox that, as I understand it, is a gateway for your cluster - so taking the idea of having an external tool protect your cluster, but at least one that understands what's happening in the cluster and handles authentication, etc. Multi-tenancy is also in it's infancy for Hadoop - but I'd expect as more of the projects start integrating better with YARN's resource management this will also be improving in the very near future.