5 ms·
We are struggling with reliability when using mounting solutions for big data in S3. Would this help?
by RocketSyntax 7y ago
We are struggling with reliability when using mounting solutions for big data in S3. Would this help?
- iRobbery 7y agodepends how you would use it i guess. And it seems quite bound to certain data formats. I initially thought after reading the headline, data as in any kinds of bytes to replicate or something. But it is something else, mainly by reading "is focused on optimized transport of the Arrow columnar format (i.e. “Arrow record batches”) over gRPC"
- khc 7y agoDefine big data? Have you tried https://github.com/kahing/goofys/ https://github.com/kahing/goofys/ ? Disclaimer: I am the author
- RocketSyntax 7y agoYes. 1PB. Although I don't remember the specifics about reliability; something about having to remount the entire fs if it wasn't 100% there.
- khc 7y agoWhat do you mean by not 100% there?
- luizfelberti 7y agoProbably the fact that S3 is eventually consistent and IO is not immediately visible to other clients? I'd suggest that OP use something like https://aws.amazon.com/fsx/lustre/ https://aws.amazon.com/fsx/lustre/ if they really need to mount S3 as a filesystem (do read the documentation regarding distributed consistency though), but other than that using AWS solutions in general (Athena, Presto on EMR, Spark on EMR) will tend to use specialized S3 committers such as https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-spark-s3-optimized-committer.html https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-spar... and are quite good at minimizing this problem. Other than that, read the S3 docs for this behavior.
- khc 7y agoAgreed that using native s3 solutions is better than using something that emulates posix on s3. Unfortunately the former isn't always possible.
- StreamBright 7y ago>> mounting solutions Could you elaborate?