7 ms·
A Distributed File System in Go Cut Average Metadata Memory Usage to 100 Bytes
- invalidator 3y agoheadline.append(" Per File")
- teaearlgraycold 3y agoUncaught TypeError: headline.append is not a function
- thwarted 3y agoThis is implied by "average", assuming it's understood that a filesystem's primary unit is the "file".
- LittleCat38 3y agoAgree
- rat9988 3y agoActually, the average here just implies that tested in different configurations and scenarios, it should cut by 100 bytes in averages. Net per file. Adding "per file" does have a benefit.
- tonyhb 3y agoI'd like to learn more about JuiceFS, but from their architecture diagrams I'm struggling to see what benefit they provide if they're a layer over blob store systems like Ceph, MinIO, etc. It looks as though you need to set up the underlying storage engine that has its own recordkeeping (one point of failure), then layer JuiceFS on top — another point of failure, and I don't know what it gives you. It's good that it's fast, but if it's pointing to another blob store you'd... expect it to be fast? Is this only needed if you want properly faked file system primitives over blob stores, if you can't use blob stores directly? Definitely need to spend some time reading.
- daviesliu 3y agoJuiceFS is similar to HDFS/CephFS/Lustre, so it MUST has a component to manage metadata, similar to NameNode of HDFS or MDS of CephFS, this point of failure is the problem we have to address. The underlying blob store systems is similar to DataNode or OSD in other distributed file system, could be slower than them a little bit because of the middle layers, the overall performance is determined by the disks. So we can expect similar performance comparing to HDFS/CephFS, the benchmark results also confirm that.
- thinxer 3y agoFor “cloud-native” apps, JuiceFS is not needed. S3 is not designed for intensive metadata operations, like listing, renaming etc. For these operations, you will need a somewhat POSIX-complaint system. For example, if you want to train on ImageNet dataset, the “canonical” way [1] is to extract the images and organize them into folders, class by class. The whole dataset is discovered by directory listing. This where JuiceFS shines. Of course, if the dataset is really massive, you will mostly end-up with in-house solutions. [1]: https://github.com/pytorch/examples/blob/main/imagenet/extract_ILSVRC.sh https://github.com/pytorch/examples/blob/main/imagenet/extra...
- snerbles 3y ago> Is this only needed if you want properly faked file system primitives over blob stores, if you can't use blob stores directly? Semiconductor EDA comes to mind, where users are at the mercy of their tool vendors and they expect something that really hasn't evolved much past an office network of Unix workstations from 1999. Object storage is almost completely alien to the tooling, and despite significant file & block storage costs there is little interest from EDA tool vendors to adapt their tools to object storage. This is a challenge for semiconductor firms operating in or moving to a cloud environment. S3FS/rclone can of course act as a shim, but are very slow when it comes to metadata operations in a typical shell. But if you were to move your metadata away from the distant object store and closer to to your compute environment things actually start becoming usable - this is the case with JuiceFS. Of course there are also tiered storage systems like Weka would have better overall performance, but is more complicated to set up and more expensive to operate than JuiceFS.
- 3y ago
- alimiracle 3y ago[dead]
- seungwoolee518 3y agoI've put this on table for distributed storage on Kubernetes. It's quite simple in front, simple to explain that split object and metadata storage. but it's too hard to select and tune each storage. So I (or We)'ve decided to move to Ceph.
- daviesliu 3y agoAgreed, the all-in-one solution (Ceph) should be better, if you have to setup all the components. If you already have the infra (databases and object stores), then JuiceFS is the easiest solution to have a distributed file system.
- KaiserPro 3y agoEx HPC person here. 1) never cluster unless you have to. 2) never use distributed Filesystems unless you need a global namespace 3) distributed filesystems mean that you have a single point of failure, its just more complex and subtle form of failure. 4) You have to think about how partition works, and how you want to recover. Sharding standard NFS fileservers is by far the easiest way to scale, have redundancy, and partition your failure zones. using autofs on a custom directory you can mount servers ondemand. /mnt/0 /mnt/1 /mnt/2 etc etc etc. You then have the ability to chose your sharding scheme for speed, redundancy or some other constraint. Nowadays the performance that distributed systems allowed are commodity. A single 4u storage server can saturate a 100gig network link with random disk IO.
- seungwoolee518 3y agoWe've used a some of Distributed storage (Longhorn, OpenEBS) things for Kubernetes and had no luck. I agree with you at ALL the things you've listed. When we've started to used Ceph for the first time - (At that time I was an user, not a Manager) all the thing was a total mess. All the things are set to default and Use all JBODs on Node. (And I've started to face Corrupted WALs and OSDs came from nowhere.) The handful of documentations was very helpful to cleans up and make it to run correctly. Maybe, when I go back to the starting point, I'd like to say "move to cloud" or "do nothing" or following your guidance.
- mike_d 3y agoI'll never understand the "we are open source but lock $cool_but_essential_features behind an enterprise license." Unless you plan to charge less than what a single developer makes in a quarter or two, any of your potential customers could at any point patch in the feature and open source the changes for everyone. (Traefik and Riak are the two biggest offenders I've ran into)
- MoOmer 3y agoThose patches are unlikely to be officially supported, though. I really support this mode of monetization, because community/free users get something, well, free, while others get something more for a bit more. It’s [hopefully] sustainable for the engineers and team behind the product, too!
- mike_d 3y agoSo sell support? Seems to work really well for a lot of companies.
- rapsey 3y agoNo it does not. It is a garbage business model that killed the vast majority of companies pursuing it. The rest barely make a living. https://techcrunch.com/2014/02/13/please-dont-tell-me-you-want-to-be-the-next-red-hat/ https://techcrunch.com/2014/02/13/please-dont-tell-me-you-wa...
- panta 3y agoWhile the underlying idea that selling services on open source products is not sustainable from a purely business perspective may be true, the linked article is not particularly convincing to me, it doesn't provide strong evidences and it's based on an embarrassingly small sample size. Besides, maybe we should start to consider also other metrics when we evaluate a Business success, not only the mere economic profit: there are externalized societal benefits (and damages) that are very important and nonetheless poorly accounted for.
- amluto 3y ago> JuiceFS, written in Go, can manage tens of billions of files in a single namespace. At that scale, I care about integrity. Can someone working in this space please have a real integrity story as part of the offering? Give each object (object, version pair, perhaps) a cryptographic hash of the contents, and make that hash be part of the inventory. Allow the entire bucket to opt in to mandatory hashing. Let me ask the system to do a scrub in which it verifies those cryptographic hashes. If this blows up metadata to 164 bytes per object, so be it. But the hash can probably get away with being stored with data, not metadata, as long as there is a mechanism to inventory those hashes. Keeping them in memory doesn’t seem necessary. Even S3 has only desultory support for this. A lot of competitors have nothing of the sort.
- jorticka 3y agoRocksDB supports hashing at multiple levels (key, value, files) because Meta also realized the importance of integrity. It also supports verifying them in bulk. Presumably filesystems built over rocksdb also support this.
- snissn 3y agodo you use the hash of the items you're storing as the key/filename?
- daviesliu 3y agoJuiceFS relies on the object store to provide integrity for data. Besides that, JuiceFS stores the checksum of each object as tags in S3, and verifies that when downloading the objects. Inside the metadata service, it uses merkle tree (hash of hash) to verify the integrity of whole namespace (including id of data blocks) between RAFT replicas. Once we store the hash (4 bytes) of each objects into metadata, it should provide the integrity of the whole namespace.
- amluto 3y agoDoes JuiceFS allow the user to specify the hash of a file when uploaded? And then to read that hash back later? Otherwise there’s no end-to-end integrity check.
- styluss 3y ago> Techniques like memory pools, manual memory management, directory compression, and compact file formats reduced metadata memory usage by 90%. It is interesting that high performance systems in Go end up needing these kind of performance optimizations.
- Intermernet 3y agoHigh performance systems in any language require these optimisations. Go provides a solid foundation to create the platform before you start optimising. Not saying it's the best or only choice, but it's not a bad choice.
- Karellen 3y agoHigh performance systems are where these kinds of optimisations make the most difference. If most of your overhead is due to the low performance of the language/runtime you're using, these kinds of optimisations won't make as much of a difference. That is, if you can even implement them at all. I mean, good luck trying to use memory pools and manual memory management in Python or Javascript.
- afr0ck 3y agoYou're implementing a Slab allocator. That's exactly what any Slab allocator does, including gmalloc, tcmalloc, jemalloc and SLUB (Linux kernel). But, there is probably much much more you can do on a modern machine to squeeze even more performance. Some of the things that comes into my mind are: reducing tlb stalls with hugepage awareness[1], and reducing false cache sharing on SMPs [2]. For the compression part, have you thought about OS-managed memory compression [3][4]? [1] https://google.github.io/tcmalloc/temeraire.html https://google.github.io/tcmalloc/temeraire.html [2] https://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemalloc.pdf https://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemal... [3] https://dl.acm.org/doi/pdf/10.1145/3297858.3304053 https://dl.acm.org/doi/pdf/10.1145/3297858.3304053 [4] https://www.kernel.org/doc/Documentation/vm/zswap.txt https://www.kernel.org/doc/Documentation/vm/zswap.txt
- gnarlouse 3y ago> In production, 10 metadata nodes each with 512 GB of memory collectively manage over 20 billion files. That doesn’t sound all that impressive—is it impressive?
- hansvm 3y agoThey're only using 36% of that (so they have slack capacity). It's in-memory, which has some perf wins (though I'd question if distributed systems are worth it when you can get a PCIe splitter and just mmap some 4TB SSDs on whatever cheap desktop computer somebody has lying around). 100 bytes per file is less than you would expect to support all that data if you coded it "naturally" without an eye toward space efficiency. You'll need 16-32 bytes or so just to handle a tree structure with internal string indices to be able to support finding any given file's location quickly. Toss in created times, modified times, permissions, attributes, the fact that they have a chunking structure to support both random seek and efficient/easy allocation, .... It adds up to more than 100B if you're not pretty careful. They hit that 100B mark. Is that impressive? I dunno. Their competitors didn't do it, and they did, so that's something.
- biomcgary 3y agoAs a computational biologist, I have a lot of embarrassingly parallel problems, so the lightweight concurrency of Go is really easy. In addition, my algorithms and data perform comfortably in GC-land 99% of the time, reducing my mental overhead (for science!). I had to reach for unsafe.Pointer() only once and that was to use Other People's Code (TM). I consistently find my code runs 20-80x faster than the equivalent in Python and I'm not about to cram C++ or Rust into my head along with all the biology I've memorized for ~2x performance improvement and >5x learning curve. On HN, I shouldn't need to explain why I don't use Java despite a similar set of tradeoffs to Go (other than memory use, historically).
- ramses0 3y agoCompetitive in this space is also SeaweedFS, which I've taken out for a local spin. It feels kindof like "memcached-for-fs/s3"... you add 30GB chunks of storage and it coalesces them for you without much ceremony. ~20 bytes storage overhead per file and performance seems pretty decent.
- tehcopec 3y agoI’ve been using community edition JuiceFS with Percona XtraDB cluster for Metadata and MinIO multi-node for a couple of years for large archival and backup data storage. That setup has worked really well.