18 ms·
Ceph: A Journey to 1 TiB/s
- riku_iki 3y agoWhat router/switch one would use for such speed?
- KeplerBoy 3y ago800Gbps via OSFP and QSFP-DD are already a thing. Multiple vendors have NICs and switches for that.
- _zoltan_ 3y agocan you show me a 800G NIC? the switch is fine, I'm buying 64x800G switches, but NIC wise I'm limited to 400Gbit.
- KeplerBoy 3y agofair enough, it seems I was mistaken about the NIC. I guess that has to wait for PCIe 6 and should arrive soon-ish.
- CyberDildonics 3y ago16x PCIe 4.0 is 32GB/s 16x PCIe 5.0 should be 64 GB/s, how is any computer using 100 GB/s ?
- KeplerBoy 3y agoI was talking about Gigabit/s, not Gigabyte/s. The article however actually talks about Terabyte/s scale, albeit not over a single node.
- CyberDildonics 3y ago800 gigabits is 100 gigabytes which is still more than PCIe 5.0 16x 64 gigabyte per second bandwidth. You said there were 800 gigabit network cards, I'm wondering how that much bandwidth makes it to the card in the first place. The article however actually talks about Terabyte/s scale, albeit not over a single node. This does not have anything to do with what you originally said, you were talking about 800gb single ports.
- KeplerBoy 3y agoYes, apparently I was mistaken about the NICs. They don't seem to be available yet. But it's not a PCIe limitation. There are PCIe devices out there which use 32 lanes, so you could achieve the bandwidth even on PCIe5. https://www.servethehome.com/ocp-nic-3-0-form-factors-quick-guide-intel-broadcom-nvidia-meta-inspur-dell-emc-hpe-lenovo-gigabyte-supermicro/ https://www.servethehome.com/ocp-nic-3-0-form-factors-quick-...
- NavinF 3y agoI'm not aware of any 800G cards, but FYI a single Mellanox card can use two PCIe x16 slots to avoid NUMA issues on dual-socket servers: https://www.nvidia.com/en-us/networking/ethernet/socket-direct/ https://www.nvidia.com/en-us/networking/ethernet/socket-dire... So the software infra for using multiple slots already exists and doesn't require any special config. Oh and some cards can use PCIe slots across multiple hosts. No idea why you'd want to do that, but you can.
- epistasis 3y agoGiven their configuration of just 4U spread across 17 racks, there's likely a bunch of compute in the rest of the rack, and 1-2 top of rack switches like this: https://www.qct.io/product/index/Switch/Ethernet-Switch/T7000-Series/QuantaMesh-T7032-IX7D https://www.qct.io/product/index/Switch/Ethernet-Switch/T700... And then you connect the TOR switches to higher level switches in something like a Clos distribution to get the desired bandwidth between any two nodes: https://www.techtarget.com/searchnetworking/definition/Clos-network https://www.techtarget.com/searchnetworking/definition/Clos-...
- NavinF 3y agoLinked article says they used 68 machines with 2 x 100GbE Mellanox ConnectX-6 cards. So any 100G pizza box switches should work. Note that 36 port 56G switches are dirt cheap on eBay and 4tbps is good enough for most homelab use cases
- riku_iki 3y ago> So any 100G pizza box switches should work. but will it be able to handle combined TB/s traffic?
- baq 3y agoany switch which can't handle full load on all ports isn't worthy of the name 'switch', it's more like 'toy network appliance'
- birdman3131 3y agoI will forever be scarred by the "Gigabit" switches of old that were 2 gigabit ports and 22 100mb ports. Coworker bought it missing the nuance.
- bombcar 3y agoStill happens, gotta see if the top speed mentioned is an uplink or normal ports.
- aaronax 3y agoYes. Most network switches can handle all ports at 100% utilization in both directions simultaneously. Take for example the Mellanox SX6790 available for less than $100 on eBay. It has 36 56gbps ports. 36 * 2 * 56 = 4032gbps and it is stated to have a switching capacity of 4.032Tbps. Edit: I guess you are asking how one would possibly sip 1TiB/s of data into a given client. You would need multiple clients spread across several switches to generate such load. Or maybe some freaky link aggregation. 10x 800gbps links for your client, plus at least 10x 800gbps links out to the servers.
- matheusmoreira 3y agoDoes anyone have experience running ceph in a home lab? Last time I looked into it, there were quite significant hardware requirements.
- nullwarp 3y agoThere still are. As someone who has done both production and homelab deployments: unless you are specifically just looking for experience with it and just setting up a demo - don't bother. When it works, it works great - when it goes wrong it's a huge headache. Edit: As just an edit, if distributed storage is just something you are interested in there are much better options for a homelab setup: - seaweedfs has been rock solid for me for years in both small and huge scales. we actually moved our production ceph setup to this. - longhorn was solid for me when i was in the k8s world - glusterfs is still fine as long as you know what you are going into.
- reactordev 3y agoI'd throw minio [1] in the list there as well for homelab k8s object storage. [1] https://min.io/ https://min.io/
- speedgoose 3y agoAlso garage. https://garagehq.deuxfleurs.fr/ https://garagehq.deuxfleurs.fr/
- BlackLotus89 3y agoGarage seems to only to duplication https://garagehq.deuxfleurs.fr/documentation/design/goals/ https://garagehq.deuxfleurs.fr/documentation/design/goals/ > Storage optimizations: erasure coding or any other coding technique both increase the difficulty of placing data and synchronizing; we limit ourselves to duplication. This is probably a nogo for most use cases where you work with large datasets....
- 3y ago
- stuff4ben 3y agoI used to love doing experiments like this. I was afforded that luxury as a tech lead back when I was at Cisco setting up Kubernetes on bare metal and getting to play with setting up GlusterFS and Ceph just to learn and see which was better. This was back in 2017/2018 if I recall. Good ole days. Loved this writeup!
- knicholes 3y agoI had to run a bunch of benchmarks to compare speeds of not just AWS instance types, but actual individual instances in each type, as some NVME SSDs have been more used than others in order to lube up some Aerospike response times. Crazy.
- j33zusjuice 3y agoAd-tech, or?
- knicholes 3y agoYeah. Serving profiles for customized ad selection.
- redrove 3y agoA Heketi man! I had the same experience around the same years, what a blast. Everything was so new..and broken!
- CTrox 3y agoSame here, still remember that time our Heketi DB partially corrupted and we had to fix it up by exporting it to a massive json file, fix it up by looking at the Gluster state and importing it again. I can't quite remember the details but I think it had to do with Gluster snapshots being out of sync with the state in the DB.
- amluto 3y agoI wish someone would try to scale the nodes down. The system described here is ~300W/node for 10 disks/node, so 30W or so per disk. That’s a fair amount of overhead, and it also requires quite a lot of storage to get any redundancy at all. I bet some engineering effort could divide the whole thing by 10. Build a tiny SBC with 4 PCIe lanes for NVMe, 2x10GbE (as two SFP+ sockets), and a just-fast-enough ARM or RISC-V CPU. Perhaps an eMMC chip or SD slot for boot. This could scale down to just a few nodes, and it reduces the exposure to a single failure taking out 10 disks at a time. I bet a lot of copies of this system could fit in a 4U enclosure. Optionally the same enclosure could contain two entirely independent switches to aggregate the internal nodes.
- deleted 3y ago[deleted]
- jeffbee 3y agoI think the chief source of inefficiency in this architecture would be the NVMe controller. When the operating system and the NVMe device are at arm's length, there is natural inefficiency, as the controller needs to infer the intent of the request and do its best in terms of placement and wear leveling. The new FDP (flexible data placement) features try to address this by giving the operating system more control. The best thing would be to just hoist it all up into the host operating system and present the flash, as nearly as possible, as a giant field of dumb transistors that happens to be a PCIe device. With layers of abstraction removed, the hardware unit could be something like an Atom with integrated 100gbps NICs and a proportional amount of flash to achieve the desired system parallelism.
- booi 3y agoIs that a lot of overhead? The disk itself uses about 10W and high speed controllers use about 75W leaves pretty much 100W for the rest of the system including overhead of about 10%. Scale up the system to 16 disks and there’s not a lot of room for improvement
- kbenson 3y agoThere probably is a sweet spot for power to speed, but I think it's possibly a bit larger than you suggest. There's overhead from the other components as well. For example, the Mellanox NIC seems to utilize about 20W itself, and while the reduced numbers of drives might allow for a single port NIC which seems to use about half the power, if we're going to increase the number of cables (3 per 12 disks instead of 2 per 5), we're not just increasing the power usage of the nodes themselves put also possible increasing the power usage or changing the type of switch required to combine the nodes. If looked at as a whole, it appears to be more about whether you're combining resources at a low level (on the PCI bus on nodes) or a high level (in the switching infrastructure), and we should be careful not to push power (or complexity, as is often a similar goal) to a separate part of the system that is out of our immediate thoughts but still very much part of the system. Then again, sometimes parts of the system are much better at handling the complexity for certain cases, so in those cases that can be a definite win.
- mrb 3y agoI wanted to see how 1 TiB/s compares to the actual theoretical limits of the hardware. So here is what I found: The cluster has 68 nodes, each a Dell PowerEdge R6615 (https://www.delltechnologies.com/asset/en-us/products/servers/technical-support/poweredge-r6615-technical-guide.pdf https://www.delltechnologies.com/asset/en-us/products/server...). The R6615 configuration they run is the one with 10 U.2 drive bays. The U.2 link carries data over 4 PCIe gen4 lanes. Each PCIe lane is capable of 16 Gbit/s. The lanes have negligible ~3% overhead thanks to 128b-132b encoding. This means each U.2 link has a maximum link bandwith of 16 * 4 = 64 Gbit/s or 8 Gbyte/s. However the U.2 NVMe drives they use are Dell 15.36TB Enterprise NVMe Read Intensive AG, which appear to be capable of 7 Gbyte/s read throughput (https://www.serversupply.com/SSD%20W-TRAY/NVMe/15.36TB/DELL/182NW_356114.htm https://www.serversupply.com/SSD%20W-TRAY/NVMe/15.36TB/DELL/...). So they are not bottlenecked by the U.2 link (8 Gbyte/s). Each node has 10 U.2 drive, so each node can do local read I/O at a maximum of 10 * 7 = 70 Gbyte/s. However each node has a network bandwith of only 200 Gbit/s (2 x 100GbE Mellanox ConnectX-6) which is only 25 Gbyte/s. This implies that remote reads are under-utilizing the drives (capable of 70 Gbyte/s). The network is the bottleneck. Assuming no additional network bottlenecks (they don't describe the network architecture), this implies the 68 nodes can provide 68 * 25 = 1700 Gbyte/s of network reads. The author benchmarked 1 TiB/s actually exactly 1025 GiB/s = 1101 Gbyte/s which is 65% of the maximum theoretical 1700 Gbyte/s. That's pretty decent, but in theory it's still possible to be doing a bit better assuming all nodes can concurrently truly saturate their 200 Gbit/s network link. Reading this whole blog post, I got the impression ceph's complexity hits the CPU pretty hard. Not compiling a module with -O2 ("Fix Three": linked by the author: https://bugs.launchpad.net/ubuntu/+source/ceph/+bug/1894453 https://bugs.launchpad.net/ubuntu/+source/ceph/+bug/1894453) can reduce performance "up to 5x slower with some workloads" (https://bugs.gentoo.org/733316 https://bugs.gentoo.org/733316) is pretty unexpected, for a pure I/O workload. Also what's up with OSD's threads causing excessive CPU waste grabbing the IOMMU spinlock? I agree with the conclusion that the OSD threading model is suboptimal. A relatively simple synthetic 100% read benchmark should not expose a threading contention if that part of ceph's software architecture was well designed (which is fixable, so I hope the ceph devs prioritize this.)
- wmf 3y agoI think PCIe TLP overhead and NVMe commands account for the difference between 7 and 8 GB/s.
- MPSimmons 3y agoThe worst problems I've had with in-cluster dynamic storage were never strictly IO related, and were more the storage controller software in kubernetes having problems with real-world problems like pods dying and the PVCs not attaching until after very long timeouts expired, with the pod sitting in ContainerCreating until the PVC lock was freed. This has happened in multiple clusters, using rook/ceph as well as Longhorn.
- einpoklum 3y agoWhere can I read about the rationale for ceph as a project? I'm not familiar with it.
- jseutter 3y agohttp://www.45drives.com/blog/ceph/what-is-ceph-why-our-customers-love-it/ http://www.45drives.com/blog/ceph/what-is-ceph-why-our-custo... is a pretty good introduction. Basically you can take off-the-shelf hardware and keep expanding your storage cluster and ceph will scale fairly linearly up through hundreds of nodes. It is seeing quite a bit of use in things like Kubernetes and OpenShift as a cheap and cheerful alternative to SANs. It is not without complexity, so if you don't know you need it, it's probably not worth the hassle.
- jacobwg 3y agoNot sure how common the use-case is, but we're using Ceph to effectively roll our own EBS inside AWS on top of i3en EC2 instances. For us it's about 30% cheaper than the base EBS cost, but provides access to 10x the IOPS of base gp3 volumes. The downside is durability and operations - we have to keep Ceph alive and are responsible for making sure the data is persistent. That said, we're storing cache from container builds, so in the worst-case where we lose the storage cluster, we can run builds without cache while we restore.
- alberth 3y agoCeph has an interesting history. It was created at Dreamhost (DH), for their internal needs by the founders. DH was doing effectively IaaS & PaaS before those were industry coined words (VPS, managed OS/database/app-servers). They spun Ceph off and Redhat bought it. https://en.wikipedia.org/wiki/DreamHost https://en.wikipedia.org/wiki/DreamHost
- epistasis 3y agoA bit more to the story is that it was created also at UC Santa Cruz, by Sage Weil, a Dreamhost founder, while he was doing graduate work there. UCSC has had a lot of good storage research.
- dekhn 3y agothe fighting banana slugs
- AdamJacobMuller 3y agoI remember the first time I deployed ceph, would have been around 2010 or 2011, had some really major issues which would nearly resulted in data loss and due to someone else not realizing what "this cluster is experimental, do not store any important data here" meant, the data on ceph was the only copy of the irreplaceable data in the world, loosing the data would have been fairly catastrophic for us. I ended up on the ceph IRC channel and eventually had Sage helping me fix the issues directly, helping me find bugs and writing patches to fix them in realtime. Super amazingly nice guy that he was willing to help, never once chastised me for being so stupid (even though I was), also wicked smart.
- antongribok 3y agoSage is one of the nicest, down to earth, super smart individuals I've met. I've talked to him at a few OpenStack and Ceph conferences, and he's always very patient answering questions.
- artyom 3y agoYeah, as a customer (still one) I remember their "Hey, we're going to build this Ceph thing, maybe it ends up being cool" blog entry (or newsletter?) kinda just sharing what they were toying with. It was a time of no marketing copy and not crafting every sentence to sell you things. I think it was the university project of one of the founders, and the others jumped in supporting it. Docker has a similar origins story as far as I know.
- louwrentius 3y agoRemember, random IOPs without latency is a meaningless figure.
- peter_d_sherman 3y agoCeph is interesting... open source software whose only purpose is to implement a distributed file system... Functionally, Linux implements a file system (well, several!) as well (in addition to many other OS features) -- but (usually!) only on top of local hardware. There seems to be some missing software here -- if we examine these two paradigms side-by-side. For example, what if I want a Linux (or more broadly, a general OS) -- but one that doesn't manage a local file system or local storage at all? One that operates solely using the network, solely using a distributed file system that Ceph, or software like Ceph, would provide? Conversely, what if I don't want to run a full OS on a network machine, a network node that manages its own local storage? The only thing I can think of to solve those types of problems -- is: What if the Linux filesystem was written such that it was a completely separate piece of software, and a distributed file system like Ceph, and not dependent on the other kernel source code (although, still complilable into the kernel as most linux components normally are)... A lot of work? Probably! But there seems to be some software need for something between a solely distributed file system as Ceph is, and a completely monolithic "everything baked in" (but not distributed!) OS/kernel as Linux is... Note that I am just thinking aloud here -- I probably am wrong and/or misinformed on one or more fronts! So, kindly take this random "thinking aloud" post -- with the proverbial "grain of salt!" :-)
- wmf 3y agowhat if I want a Linux ... that doesn't manage a local file system or local storage at all [but] operates solely using the network, solely using a distributed file system Linux can boot from NFS although that's kind of lost knowledge. Booting from CephFS might even be possible if you put the right parts in the initrd.
- lmz 3y agoNFS root docs here https://www.kernel.org/doc/Documentation/filesystems/nfs/nfsroot.txt https://www.kernel.org/doc/Documentation/filesystems/nfs/nfs...
- peter_d_sherman 3y ago
- chx 3y agoThere was a point in history when the total amount of digital data stored worldwide reached 1TiB for the first time. It is extremely likely this day was within the last sixty years. And here we are moving that amount of data every second on the servers of a fairly random entity. We not talking of a nation state or a supranatural research effort.
- qingcharles 3y agoThat reminds me of a calculation I did which showed that my desktop PC would be more powerful than all of the computers on the planet combined in like 1978 :D
- plagiarist 3y agoMy phone has more computation than anything I would have imagined owning, and I sometimes turn on the screen just to use as a quick flashlight.
- qingcharles 3y agoHaha.. imagine taking it back to 1978 and showing how it has more computing power than the entire planet and then telling them that you mostly just use it to find that thing you lost under the couch :D
- fiddlerwoaroof 3y agoIt’s at least 20ish years ago: I remember an old sysadmin talking about managing petabytes before 2003
- aspenmayer 3y agoThose numbers seem reasonable in that context. I first started using BitTorrent around that time as well, and it wasn't uncommon to see many users long-term seeding multiple hundreds of gigabytes of Linux ISOs alone. Here’s another usage scenario with data usage numbers I found a while back. > A 2004 paper published in ACM Transactions on Programming Languages and Systems shows how Hancock code can sift calling card records, long distance calls, IP addresses and internet traffic dumps, and even track the physical movements of mobile phone customers as their signal moves from cell site to cell site. > With Hancock, "analysts could store sufficiently precise information to enable new applications previously thought to be infeasible," the program authors wrote. AT&T uses Hancock code to sift 9 GB of telephone traffic data a night, according to the paper. https://web.archive.org/web/20200309221602/https://www.wired.com/2007/10/att-invents-pro/ https://web.archive.org/web/20200309221602/https://www.wired...
- one_buggy_boi 3y agoIs modern Ceph appropriate for transactional database storage, how is the IO latency? I'd like to move to a cheaper cfs that can compete with systems like Oracle's clustered file system or DBs backed by something like Veritas. Veritas supports multi-petabyte DBs and I haven't seen much outside of it or ocfs that similarly scales with acceptable latency
- samcat116 3y agoLatency is quite poor, I wouldn't recommend running high performance database loads there.
- louwrentius 3y agoFrom my dated experience, Ceph is absolutely amazing but latency is indeed a relative weak spot. Everything has a trade-off and for Ceph you get a ton of capability but latency is such a trade-off. Databases - depending on requirements - may be better off on regular NVMe and not on Ceph.
- yencabulator 3y agoIt's pretty unfair to compare latency of a local NVMe SSD to over-the-network 3x replicated storage. "It's faster if I do less." [Disclaimer: ex-Inktank employee]
- louwrentius 3y agoI don’t think it’s unfair, there are applications that still are ok with Ceph latencies: I bet it’s good enough for a ton of things. But not all things.
- e12e 3y agoNo, it's important when planning - eg: one big database cluster that provides db-as-a-service (but maybe needs some dedicated ops resources) vs smaller DBs with virtualized storage on ceph (ops resources for ceph cluster and vm tools like k8s). If the latter is too slow for your typical usage...
- rafaelturk 3y agoI'm playing a lot with MicroCeph. Its aopinionated low TOS, friendly setup of Ceph. Looking forward additional comments. Planning to use it in production and replace lots of NAS servers.
- louwrentius 3y agoI think Ceph can be fine for NAS use cases, but be wary of latency and do some benchmarking. You may need more nodes/osds than you think to reach latency and throughput targets.
- hinkley 3y agoSure would be nice if you defined some acronyms.
- nghnam 3y agoMy old company ran public and private cloud with Openstack and Ceph. We had 20 Supermicro (24 disks per server) storage nodes and total capacity was 3PB. We learnt some experiences, especially a flapping disk made whole system performance degraded. Solution was removing bad sector disk as soon as possible.
- amadio 3y agoNice article! We've also recently reached the mark of 1TB/s at CERN, but with EOS (https://cern.ch/eos https://cern.ch/eos), not ceph: https://www.home.cern/news/news/computing/exabyte-disk-storage-cern https://www.home.cern/news/news/computing/exabyte-disk-stora... Our EOS clusters have a lot more nodes, however, and use mostly HDDs. CERN also uses ceph extensively.
- theyinwhy 3y agoGreat! What's your take on ceph? Is the idea to migrate to EOS long term?
- amadio 3y agoEOS and ceph have different use cases at CERN. EOS holds physics data and user data in CERNBox, while ceph is used for a lot of the rest (e.g. storage for VMs, and other applications). So both will continue to be used as they are now. CERN has over 100PB on ceph.
- ComputerGuru 3y agoIs there a reason you run both and don't converge on one or the other?
- amadio 3y agoYes, let me expand a bit on the answer above. EOS was designed and developed with the unique needs of the LHC experiments in mind. The advantages it has are features used by the experiments, like support for remote access via the XRootD protocol, which is used for data analysis with ROOT (only the parts of files needed by an analysis are downloaded); rich support for client authentication methods (Kerberos, X509, OIDC, etc); and support to also FUSE mount everything to give a convenient POSIX-like view of the data. EOS needs to sustain ingestion of data at high rates from experiments (10s of GB/s each) for several months at a time during data taking without any downtime, while at the same time having tens of thousands of clients connected reading data as well. It's also integrated with the CERN Tape Archive (CTA) and File Tranfer Service (FTS), used for long term archival and data management across sites, respectively. In the cases where block/object storage is needed, like storage for VMs, S3 storage for various uses, etc, then ceph is better suited. It has lower latency and EOS does not offer block-level access. In addition to providing storage services for OpenStack/Openshift, ceph is used to provide storage to back AFS and CVMFS, for example. CVMFS is another interesting piece of CERN's infrastructure, it's a read-only, HTTP-based FUSE filesystem used to distribute the software used by the experiments to grid sites around the world. Dan van der Ster, mentioned in the article above, has a good overview of ceph usage at CERN here: https://youtu.be/2I_U2p-trwI?si=Tsq4h8NIu4vSZQwt https://youtu.be/2I_U2p-trwI?si=Tsq4h8NIu4vSZQwt If you are interested in EOS, we have the EOS workshop coming up in March: https://indico.cern.ch/event/1353101/ https://indico.cern.ch/event/1353101/
- brobinson 3y agoI'm curious what the performance difference would be on a modern kernel.
- PiratesScorn 3y agoFor context, I’ve been leading the work on this cluster client-side (not the engineer that discovered the IOMMU fix) with Clyso. There was no significant difference when testing between the latest HWE on Ubuntu 20.04 and kernel 6.2 on Ubuntu 22.04. In both cases we ran into the same IOMMU behaviour. Our tooling is all very much catered around Ubuntu so testing newer kernels with other distros just wasn’t feasible in the timescale we had to get this built. The plan was < 2 months from initial design to completion. Awesome to see this on HN, we’re a pretty under-the-radar operation so there’s not much more I can say but proud to have worked on this!
- brobinson 3y agoHey, thanks for the response! There have been a lot of across the board improvements in the kernel in the last four years so I'm surprised there's not a noticeable performance improvements in 6.2 (although I also consider 6.2 old at this point).
- kylegalbraith 3y agoThis is a fascinating read. We run a Ceph storage cluster for persisting Docker layer cache [0]. We went from using EBS to Ceph and saw a massive difference in throughput. Went from a write throughput of 146 MB/s and 3,000 IOPS to 900 MB/s and 30,000 IOPS. The best part is that it pretty much just works. Very little babysitting with the exception of the occasional fs trim or something. It’s been a massive improvement for our caching system. [0] https://depot.dev/blog/cache-v2-faster-builds https://depot.dev/blog/cache-v2-faster-builds
- guywhocodes 3y agoDid something very similar almost 10 years ago, EBS costs were 10x+ the cost for same perfomance CEPH cluster on the node disks. Eventually we switched to our own racks and cut it almost in ten again. We developed the inhouse expertise for how to do it and we were free.
- e12e 3y agoDid you host ebs on bare metal? How are you hosting ceph - your own/rented metal, ec2 - VMs? Wasn't immediately clear to me from the blog.
- kylegalbraith 3y agoWe started with AWS EBS volumes with BuildKit on EC2. We've now moved to BuildKit on EC2 and a Ceph storage cluster on bare metal EC2 instances.
- louwrentius 3y agoI wrote an intro to Ceph[0] for those who are new to Ceph. It featured in a Jeff Geerling video briefly recently :-) [0]: Understanding Ceph: open-source scalable storage https://louwrentius.com/understanding-ceph-open-source-scalable-storage.html https://louwrentius.com/understanding-ceph-open-source-scala...
- justinclift 3y agoHas anything important changed since 2018, when you wrote that? :)
- louwrentius 3y agoConceptually not as far as I know.
- mobilemidget 3y agoCool benchmark, and interesting, however it would have read a lot better if abbreviations are explained at first usage. Not everybody is familiar with all terminology used in the post. Nonetheless congrats with results.
- up2isomorphism 3y agoThis is an insanely expensive cluster built to show a benchmark. 68 node cluster serving only 15TB storage in total.
- PiratesScorn 3y agoThe purpose of the benchmarking was to validate the design of the cluster and to identify any issues before going into production, so it achieved exactly that objective. Without doing this work a lot of performance would have been left on the table before the cluster could even get out the door. As per the blog, the cluster is now in a 6+2 EC configuration for production which gives ~7PiB usable. Expensive yes, but well worth it if this is the scale and performance required.
- up2isomorphism 3y agoYou are talking different thing. I don’t care what “purpose “ you want to achieve. I merely point out this performance number is mediocre at best, because the enormous computing power thrown at it , wether you like it or not. To put it into perspective there are 68 nodes with 98 hard thread each, means only 1000/7000 = 140MB/s per thread or 280MB/s per core, and that’s not that impressive, to be honest.
- mrunkel 3y ago> This is an insanely expensive cluster built to show a benchmark. 68 node cluster serving only 15TB storage in total. This reads to me (and the OP) that you are saying the purpose of this "insanely expensive cluster" was to "show a benchmark." That's what OP is addressing in his response. No where do you mention anything about performance.
- up2isomorphism 3y ago1TB/second is a benchmark number, and obviously it is trying to impress. This is already a purpose without further clarification, on the other hand it does not necessarily mean it can not have other purpose - which again looks an extremely expensive cluster for that purpose. And I do not see the reason to down vote except some one got hurt in the feeling with a fact. With the actually configuration shown, it is just not that performant nor economical, as I said in the reply, if you had read anything about it.
- francoismassot 3y agoDoes someone knows how Ceph compares to other object storage engine like MinIO/Garage/...? I would love to see some benchmarks there.
- matesz 3y agoThis would be great, to have a universal benchmark of all available open source solutions for self-hosting. Links appreciated!
- Proven 3y ago[dead]
- mxroeoek 3y ago[dead]
- kaliszad 3y agoWhat surprises me is, why they went with the harder to cool 1U nodes and 10 SSDs/2x100Gb NICs instead of 2U nodes with 24 SSDs/2x200 or even 400Gb NICs. They could remove the network bottleneck and save on power thanks to larger, lower speed fans and less CPU packages, possibly with more cores per socket though. Also, having a smaller number of nodes increases the blast radius but with even 34 nodes this is probably not such a problem. However, with less nodes they could have a flatter network with 4 switches or so too.
- PiratesScorn 3y agoBlast radius is the primary factor as you say and just generally makes things like patching and HW replacements less stressful. The racks and switches already exist and are heavily utilised for other purposes so the additional physical footprint for ceph is pretty tiny :)