6 ms·
Comparing Ceph, Linstor, Mayastor and Vitastor Storage Performance in Kubernetes
- polskibus 4y agoThis is great. Does anyone have experience with building on premise scalable storage based on open source components ? I mostly read about problems and edge cases with Ceph in production. I would be grateful for sharing your resource, experience and recommendations on how to tackle such task.
- ibotty 4y agoI did have Problems with ceph many years ago but have many hyper-converged clusters that run ceph without any problems. In the last years I am running rook-managed ceph without any real problems.
- rjzzleep 4y agoWhen you use Ceph in conjunction with Kubernetes, most people probably use rook. It's supposedly been production grade for a few years, but I keep running into random corner cases. In the beginning it would have problems when you ran it with XFS(the default at the time) causing the kernel to lock up. A lot of the corner cases nowadays are related to not being able to scale the cluster, odd behavior with CRD's of users and other components that require you to run ceph admin commands directly through a proxy. The annoying thing about a lot of the "Cloud Native" projects is that you will spend a lot of time going through github issue lists. There is a Taiwanese company that provides extremely cheap OEM ceph racks though. I've noticed that in terms of maintenance and performance something like that or a cheap TrueNAS almost certainly ends up cheaper than production workloads, unless you don't care about the data integrity of your ceph workloads.
- encryptluks2 4y agoWhen you say open source components, do you mean more commodity consumer-grade hardware or just the software components?
- TruthWillHurt 4y agoI'm so glad we have cloud managed storage like EFS / EBS so we don't have to deal with this stuff...
- nix23 4y agoI love to deal with stuff like that.
- jknoepfler 4y agoSure, if you enjoy paying $80/TB-month (before replication or backups). If you're operating at a scale that justifies running a datacenter, EBS is basically a non-starter.
- TruthWillHurt 4y ago$80 is nothing compared to an infra engineer salary of +$5000 a month. And you will need one. permamently. because this stuff breaks in so many spectacular ways... i.e splitbrain. a joy to resolve.
- jknoepfler 4y agoPer petabyte, we are talking $80,000 per month.
- Andys 4y agoI love baremetal but I agree, distributed storage is hell. I found it slightly harder to sleep at night managing a prod Ceph cluster.
- vbezhenar 4y agoWhat are the cases where NFS is not enough?
- hardwaresofton 4y agoThe usecases are related but difference -- these are block storage solutions, not filesystem level solutions. Some technologies work on multiple levels (Ceph), but others do not.
- Proven 4y ago
- paol 4y agoAny case where you need horizontal scaling or redundancy though data replication. So lots of cases.
- yjftsjthsd-h 4y agoIsn't NFS generally served by a single machine? I don't have anything inherently against NFS, but creating a SPoF / a box that can't rebooted doesn't sit well with me.
- vbezhenar 4y agoThere are guides for HA NFS servers. I don’t have experience though. I assume that it requires shared block storage.
- magicalhippo 4y agoFrom personal experience, if you want to run some software that uses SQLite or similar...
- somat 4y agoceph(and the others I assume) solves the same problem that raid solved, but using the network as a transport layer(rather than the computer for raid). this generally means it is very easy to scale. So the answer to your question is that your question was wrong, you are comparing different layers. you can actually access ceph over nfs, however usually you will try to use a ceph aware access method, ether the ceph filesystem driver, or the S3-like library depending on what your application wants.
- lukaslalinsky 4y agoI wonder if the Linstor integration with Kubernetes got better. A few years back, the CSI driver was full bugs and you couldn't really use it from K8s alone, you had to use Linstor tools on the servers directly to get things unstuck or get orphaned volumes to actually get deleted.
- tbronchain 4y agoA lot better. A couple of years ago I could see volumes on Linstor getting completely stuck and unrecoverable whenever the network was getting busy or unstable. Nodes reboot were a nightmare too. Have a setup now with their Piraeus operator[1], Kubernetes >= 1.20, rancher and calico, and it seems to be very stable. XFS have been giving better results too. Still, better not to try too many reboot loops on the nodes. 1: https://github.com/piraeusdatastore/piraeus-operator/ https://github.com/piraeusdatastore/piraeus-operator/
- hardwaresofton 4y agoNote, kvaps actually wrote one of the LINSTOR operators (which is no longer maintained): https://github.com/kvaps/kube-linstor https://github.com/kvaps/kube-linstor
- kvaps 4y agoYeah, because I have joined to Piraeus project :)
- daper 4y agoNote that they are using 10GbE network when even a single NVMe disk used in the test has few times more bandwidth so the bandwidth results are constrained by the network hardware. Creating on-premise infrastructure a year ago we went for 2x25Gbps network. This or even 100Gbps seems to be current "mainstream" and 10GbE is definitely not enough for NVMe speeds available today.
- neverartful 4y ago3 servers with 2 disks each for a grand total of 6 disks?? This configuration is not even close to a suitable size for a tiny POC. How about 4 or 5 servers with a minimum of 8 disks each for a MINIMAL total of 32-40 disks? A more realistic size would be 8-10 servers with 12-16 disks each (total of 100-160 disks).
- tlamponi 4y agoTheir ceph usage is quite odd IMO, that few disks per node and such a (for storage) slow network interconnect cannot give you good results. We have a two-year-old (time really does fly) ceph benchmark with 100G Ethernet interconnect for cluster network[0] if anybody is interested. IMO using NVMe's flash and then connecting that with anything below 25G makes not much sense and is a waste of money, as network isn't that expensive anymore; especially with three to five nodes you can set up a full mesh to avoid an often more expensive switch. IME Ceph really starts to shine on a) slightly bigger setups and b) maintenance like upgrading nodes and swapping faulty disks, that's just a bliss if one doesn't violate a few basic rules (like maybe don't use 2/1 replica (two copies but return already after only one got written)). But also smaller cluster can work fine and can be quite performant, well at least if the interconnect bandwidth isn't one that was specified at the end of the 90s.. Also, ZFS 0.8.6 is from 2020, lots of stuff happened since then, a bit odd to post a new benchmark using almost two-year-old software. [0]: https://www.proxmox.com/en/downloads/item/proxmox-ve-ceph-benchmark-2020-09 https://www.proxmox.com/en/downloads/item/proxmox-ve-ceph-be...
- yjftsjthsd-h 4y ago> Also, ZFS 0.8.6 is from 2020, lots of stuff happened since then, a bit odd to post a new benchmark using almost two-year-old software. That's probably a side effect of using Ubuntu 20.04, just like having a 5.4 kernel (which is likewise getting old). Although yes, that probably isn't helping, and it would be interesting to do the same test and change nothing but moving to Ubuntu 22.04 and the corresponding kernel and ZFS versions, both of which should help.
- hardwaresofton 4y agoIt's absolutely a side effect of using Ubuntu -- I had to build ZFS (was quite easy) to get newer versions on Ubuntu and do a bit to prevent the system from picking up the default version (pakages & *-dkms packages)
- IcePic 4y ago"maybe don't use 2/1 replica (two copies but return already after only one got written)" I don't think ceph ever does that.
- derefr 4y agoTo step back a level of consideration: if you're already running on an IaaS platform that has its own (expensive) object storage, can it ever be worth it to run your own object-storage system using that same IaaS provider's (expensive) compute? Or to run your object-storage system on some other bare-metal host, and then PUT things into it using your IaaS provider's (expensive) egress bandwidth?
- ddorian43 4y agoThat is the reason why bandwidth is expensive on cloud. So youll have no choice.