4 ms·
Why do you say it'd be a bad idea to store frequently accessed experimental data in Google Drive?
by trunnell 5y ago
Why do you say it'd be a bad idea to store frequently accessed experimental data in Google Drive?
- g_p 5y agoIndeed - given much of Google drive marketing seems to be trying to encourage it's use for routine frequently accessed data, this is an interesting question. Much of the marketing for drive seems to still be (in sentiment at least) about dumping it all into Google drive and never having to worry about storage again. I assume they mean storing huge files in Google drive, but actually that's probably a significant use case for many users of cloud storage.
- tuckerman 5y agoI think the biggest issue is provenance and access control/logging—Drive doesn’t have a lot of great tools for that. Backups as well. That said, I’m struggling to imagine how you could even efficiently process on the order of 1 PiB of data from Drive. Having that in a data warehouse or some object storage close to your compute seems required to make any good use of it.
- agentdrtran 5y agoLogging in Drive is pretty robust (backups are not). There's a few tools that can move that level of data, but google's now-first-party migrate product could likely handle it.
- tuckerman 5y agoThat logging will give you who/when but won’t have things like field level access logging or information about the analysis pipeline itself. With regards to data locality, I was more thinking about running a map reduce or similar ETL pipeline.