4 ms·
> Just creating a filesystem on an 8TB disk takes hours. Am I being dense here, but why would you create a filesystem on the hard drive? The hard drive should
by bsdetector 11y ago
> Just creating a filesystem on an 8TB disk takes hours.
Am I being dense here, but why would you create a filesystem on the hard drive? The hard drive should only store file contents.
There's all kinds of seeking and syncing and random access needed for the filesystem. Seeks from file contents can't be avoided, but ones due to the filesystem can be.
If there's no metadata-on-ssd + data-on-disk filesystem for Linux, there should be.
- winter_blue 11y agoSo what you're advocating is that instead of creating an ext4/zfs/xfs file system on each drive, have a sort of distributed filesystem that has raw access to each drive. And the i-node of the distributed filesystem could be always-available on RAM with SSD backing. My guess is, Google probably already does this.
- woodman 11y agoZFS allows for that level of control: assign a cache device to the pool, then set your secondarycache flag to metadata. Yes the metadata will hit the primary cache first (memory), but it will spill onto the secondary as it is pressured out of ARC. I've got a SD card playing that role in my build machine, it works really well for keeping track of all the tiny files that make up the FreeBSD base and ports trees.
- pjc50 11y agoInteresting, I don't think much work has been done in this area. It's not a common way of dividing the workload, even though it sounds obvious.
- darkr 11y agoLots of distributed filesystems (I guess I'm talking about distributed object stores here rather than an actual clustered network filesystem) rely on a regular POSIX filesystem for the storage of blobs. A huge amount of work has gone into making ext4/zfs/xfs etc fast and reliable, plus you get other benefits like journaling, filesystem caching, fsync(), metadata caching (a la ZFS), so there are credible arguments for not going down the path of NIH.