4 ms·
I do scientific computing in google cloud. When I first got started, I heavily relied on GCSFuse. Over time, I have encountered enough trouble that I no longer
by carbocation 3y ago
I do scientific computing in google cloud. When I first got started, I heavily relied on GCSFuse. Over time, I have encountered enough trouble that I no longer use it for the vast majority of my work. Instead, I explicitly localize the files I want to the machine that will be operating on them, and this has eliminated a whole class of slowdown bugs and availability bugs.
The scale of data for my work is modest (~50TB, ~1 million files total, about 50k files per "directory").
- nyc_pizzadev 3y agoDid you use a local caching proxy like Varnish or Squid? Would that have helped?
- dekhn 3y agoThese codes aren't talking HTTP. They are talking POSIX to a real filesystem. The problem is that cloud-based FUSE mounts are never as reliable (they will "just hang" at random times and you need some sort of external timeout to kill the process and restart the job and possible the host) as a real filesystem (either a local POSIX one or NFS or SMB). I've used all the main FUSE cloud FS (gcsfuse, s3-fuse, rclone, etc) and they all end up falling over in prod. I think a better approach would be to port all the important science codes to work with file formats like parquet and use user-space access libraries linked into the application, and both the access library and the user code handle errors robustly. This is how systems like mapreduce work, and in my experience they work far more reliably than FUSE-mounts when dealing with 10s to 100s of TBs.
- markstos 3y agoI had a similar experience with S3 Fuse. It was slower, more complex and expensive than using S3 directly. I had feared refactoring my code to use the API, but it went quickly. I’ve never gone back to using or recommending a cloud filesystem like that for a project.
- laurencerowe 3y agoThese file systems are not a good fit for large numbers of small files. Their sweet spot is working with large (~GB+) files which are mostly read from beginning to end. I’ve mostly used them for bioinformatics stuff.
- paulddraper 3y ago> The scale of data for my work is modest (~50TB, ~1 million files total, about 50k files per "directory"). Then my work must be downright embarassing.
- ashishbijlani 3y agoFUSE does not work well with a large number of small files (due to high metadata ops such as inode/dentry lookups). ExtFUSE (optimized FUSE with eBPF) [1] can offer you much higher performance. It caches metadata in the kernel to avoid lookups in user space. Disclaimer: I built it. 1. https://github.com/extfuse/extfuse https://github.com/extfuse/extfuse
- laurencerowe 3y agoExtFUSE seems really cool and great for implementing performant drivers in userspace for local or lower latency network filesystems, but I doubt FUSE is the bottleneck in this case since S3/GCS have 100ms first byte latency. https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimizing-performance.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimi...
- dekhn 3y agoI am curious to hear what solutions people have found for this- for example, does anybody cache S3 in cloudfront and then point S3 clients at cloudfront? This comes up because we store a lot of image data where time-to-first-byte affects the user experience but the access patterns preclude caching unless we are willing to spend $$$.
- eblume 3y agoYou may wish to investigate cloudflare's image API: https://developers.cloudflare.com/images/cloudflare-images/ https://developers.cloudflare.com/images/cloudflare-images/ If the reason you were unable to use a CDN cache was because your access patterns require a lot of varying end serializations (due to things like image manipulation, resizing, cropping, watermarking, etc.), then this API could be a huge money saver for you. It was for me. OTOH if the cost was because compute isn't free and the corresponding cloudflare worker compute cost is too much, then yeah, that's a tough one... I don't have a packaged answer for you, but I would investigate something like ThumbHash: https://evanw.github.io/thumbhash/ https://evanw.github.io/thumbhash/ - my intuition is that you can probably serve some highly optimized/interlaced/"hashed" placeholder. The advantage of thumbhash here could be that you can optimize the access pattern to be less spendy by simply storing all of your hashes in an optimized way, since they will be extremely small, like small enough to be included in an index for index-only scans ("covering indexes"). (I have not actually tried this.)
- objectivefs 3y agoFor workloads with many small files, it usually is better to store many files in a single object. Filesystems with regular POSIX semantics, such as atomic directory renames etc, also makes it easier to integrate with existing software. We have seen a lot of scientific computing usage of our filesystem (https://objectivefs.com https://objectivefs.com) and as you mentioned localized caching of the working set is key to great performance.
- carbocation 3y agoVery strongly agree with your point. In my case the real number of files is in the billions, and they are already aggregated into grouped files to reduce that overall size. But for a period of time I tried to operate on the individual unaggregated files, and that was totally untenable. (Also expensive due to the normally-negligible cost of fetching operations.)