4 ms·
Somewhat related, they discussed why they chose to use ZFS for their storage backend as opposed to (say) Ceph in a podcast episode: * https://www.youtube.com/w
by throw0101b 2y ago
Somewhat related, they discussed why they chose to use ZFS for their storage backend as opposed to (say) Ceph in a podcast episode:
* https://www.youtube.com/watch?v=UvEKSqBBcZw https://www.youtube.com/watch?v=UvEKSqBBcZw
Certainly they already had experience with ZFS (as it is built into Illumos/Solaris), but as it was told to them by someone they trusted who ran a lot of Ceph: "Ceph is operated, not shipped [like ZFS]".
There's more care-and-feeding required for it, and they probably don't want that as they want to treat product in a more appliance/toaster-like fashion.
- pclmulqdq 2y agoCeph is sadly not very good at what it does. The big clouds have internal versions of object store that are far better (no single point of failure, much better error recovery story, etc.). ZFS solves a different problem, though. ZFS is a full-featured filesystem. Like Ceph it is also vulnerable to single points of failure.
- throw0101b 2y ago> The big clouds have internal versions of object store that are far better (no single point of failure, much better error recovery story, etc.). There are different levels of scalability needs. CERN has over a dozen (Ceph) clusters with over 100PB of total data as of 2023: * https://www.youtube.com/watch?v=bl6H888k51w https://www.youtube.com/watch?v=bl6H888k51w Certainly there are some number of folks that need more than that, but I don't there are many. > Like Ceph it is also vulnerable to single points of failure. The SPOF for ZFS is the host (unless you replicate, e.g., zfs send). What is SPOF of Ceph? You can have multiple monitors, managers, and MDSes.
- pclmulqdq 2y agoSingle-monitor is a common way to run Ceph. On top of that, many cluster configurations cause the whole thing to slow to a crawl when a very small minority of nodes go down. Never mind packet loss, bad switches, and other sorts of weird failure mechanisms. Ceph in general is pretty bad at operating in degraded modes. ZFS and systems like Tectonic (FB) and Colossus (Google) do much better when things aren't going perfectly. Do you know how many administrators CERN has for its Ceph clusters? Google operates Colossus at ~1000x that size with a team of 20-30 SREs (almost all of whom aren't spending their time doing operations).
- antongribok 2y agoThis is complete nonsense. No one running business critical installs of Ceph runs single-monitor. You can also tell Ceph to use a single disk as your failure domain. No one does that either. Homelabbers maybe, but then why are you comparing such setups with Google? We run Ceph with a failure domain of an entire rack. We can literally take down (scheduled or unscheduled) an entire rack of 40 servers, and continue to serve critical, latency sensitive applications, with no noticeable performance loss. We have a Ceph footprint 5x larger than CERN run by a team of 4-5 people.
- throw0101b 2y ago> Single-monitor is a common way to run Ceph. What? > A Ceph cluster must contain a minimum of three running monitors in order to be both redundant and highly-available. * https://docs.ceph.com/en/latest/glossary/#term-Ceph-Monitor https://docs.ceph.com/en/latest/glossary/#term-Ceph-Monitor > Our Configuring ceph section provides a trivial Ceph configuration file that provides for one monitor in the test cluster. A cluster will run fine with a single monitor; however, a single monitor is a single-point-of-failure. To ensure high availability in a production Ceph Storage Cluster, you should run Ceph with multiple monitors so that the failure of a single monitor WILL NOT bring down your entire cluster. * https://docs.ceph.com/en/latest/rados/configuration/mon-config-ref/#monitor-quorum https://docs.ceph.com/en/latest/rados/configuration/mon-conf...
- anonfordays 2y agoZFS and Ceph is apples to oranges. ZFS is scoped to a single host, Ceph can span data centers.
- ComputerGuru 2y agoIt’s very possible to run a light/small layer on top of ZFS (either userspace daemon or via FUSE) to get you most of the way to scaling ZFS-backed object storage within or across data centers depending on what specific availability metrics you need.
- anonfordays 2y agoThat's true for any filesystem, not specific to ZFS. ZFS is not a clustered or multi-host filesystem.
- seabrookmx 2y agoWhat does this light/small layer look like? In my experience you need something like GlusterFS which I wouldn't call "light".
- throw0101b 2y ago> ZFS and Ceph is apples to oranges. Oxide is shipping an on-prem 'cloud appliance'. From the customer's/user's perspective of calling an API asking for storage, it does not matter what the backend is—apple or orange—as long as "fruit" (i.e., a logical bag of a certain size to hold bits) is the result that they get back.
- anonfordays 2y agoYes, it could be NTFS behind the scenes, but this is still an apples to oranges comparison because the storage service Oxide created is Crucible[0], not ZFS. Crucible is more of an apples to apples comparison with Ceph. [0] https://github.com/oxidecomputer/crucible https://github.com/oxidecomputer/crucible
- wmf 2y agoYou mean they use Crucible instead of Ceph?