3 ms·
> However having something available as part of DB or Hadoop ecosystem enables reproducible execution while taking care of compliance (e.g. privacy), security,
by lambda 9y ago
> However having something available as part of DB or Hadoop ecosystem enables reproducible execution while taking care of compliance (e.g. privacy), security, access control etc.
You are conflating a couple of things here.
Of course for real analyses, you want to run them on standardized environments with appropriate access controls.
But nothing says that "standardized environment" has to use a distributed database that, from all accounts, seems to harm more than help.
You could just have a RDBMS or even filesystem, which is appropriately backed up or replicated, on a single beefy VM or physical machine.
> The real opportunity lies in ability to add arbitrary compute/memory capacity to DB query execution flows.
Only if the tools used to provide that ability actually allow you to take advantage of it. If scaling up to 160 nodes doesn't allow you to be any faster than 1 single laptop, you probably could have spent a lot less money on just adding more RAM to a single beefy server than a 160 node cluster.