4 ms·
I agree but sadly we don't have any plug-in solutions that operate in similar manner. How do you allocate capacity? What if your node has hard upper limit? How
by aub3bhat 9y ago
I agree but sadly we don't have any plug-in solutions that operate in similar manner. How do you allocate capacity? What if your node has hard upper limit? How do you ensure security e.g. secomp, jails etc. The moment data crosses the DB boundry you need significant investment in ensuring compliance.
The reason Spark ecosystem has been so popular is because it enables these types of computations without breaking the model.
- Terretta 9y agoThis line of discussion worries me, in case others think your security assertions are true and could be lead astray. Today’s OSS big data query execution environments are BnL neither secure nor compliant. This is not a knock. They were simply not designed for those constraints. They were built for academic or single tenant use cases without separation of duties / control. This also applies to commercial products such as Splunk. There is tremendous investment going on the last couple years in raising the security and compliance bar. We’ve worked on this with the usual suspects. But barring a handful of proprietary stacks (the big three CSPs, and a couple enterprise on prem bare metal distros) getting close, we are not there yet. Trying to land with your security or risk teams or regulators that the CSP pulled it off will likely take you longer than provisioning a compliant laptop you’d keep locked up with two keys. Unless you have several tens of man years invested in in-depth security wrapping these environments, or can choose a big three CSP w/o answering to anyone, today I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations.
- aub3bhat 9y agotoday I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations. You are utterly wrong. Algorithms and computations (especially ML kind) are never sensitive, its the data which is always sensitive. And that ALREADY exists on the cluster. If an Organization already has a Hadoop cluster containing data. You are suggesting that somehow having it downloaded to a secured laptop is better? Than say Spark running on top of the cluster? I think you are deeply mistaken. The cluster instances are already protected (if not you have a bigger problems). Also while an organization might not have Hadoop, they surely have an RDBMS, in which case the algorithms are even more useful.
- Terretta 9y agoThe code can certainly be more sensitive than the data. Consider codified expertise applied to public data.