5 ms·
I was the Engineering Manager for the Storage team at DigitalOcean that took the Block Storage project from conception to launch (though I no longer work there)
by nickvanw 9y ago
I was the Engineering Manager for the Storage team at DigitalOcean that took the Block Storage project from conception to launch (though I no longer work there), so I might be able to shed some light.
In general, it's really hard to do at-scale network-backed storage - by the time your applications get access to the file system, there are a myriad of abstractions that aren't always receiptive to the idea of the network "going away", or even a modicum of lag. On top of that, in order for it to be profitable, you need to work at a massively-shared scale. This means expensive SSDs, servers and switches that require a lot of capex and no guaranteed revenue because it's a new product. For us, this meant building entirely new network architecture in some places to allow the massive amount of data being shared in the storage cluster and across VMs to not overwhelm existing traffic, etc.
In order to create the reliability in network and persistency that your normal application desires, you need extremely strong consistency and low latency. Every replication strategy (replication and erasure coding) requires each write to touch more than one SSD/HDD/NVMe device in order to acknowledge the write, and that all needs to happen in a shared system with an immense amount of contention, every time.
It takes a while because you only get one opportunity to get all of this right - it's one thing if the network has a few more blips in a month, or if there's a bit more CPU contention than you'd like, but you absolutely can't lose peoples' data.
I can understand why companies are so hesitant to do this - there may be technical debt in their software/network stack that makes it very difficult, or they may not want to proceed unless they have the right set of experts working on the project.
- petecooper 9y agoAs a DO customer – thanks for the insight, this is really good to know.
- neom 9y agoHow do you like Ceph generally as a technology, rest of the implementation aside, thoughts on Ceph?
- deleted 9y ago[deleted]
- nickvanw 9y agoCeph is good! I'm a bit removed from keeping up on the day-to-day, but I always respected it as being a solid and dependable piece of open source software. There are options that will perform better, but they are almost always considerably more expensive than FOSS, and all have their own weird scaling quirks. With the launch of BlueStore a few months ago as well as improvements in erasure coding, I wouldn't hesitate to take a look at it again if I was starting a new project.
- nik736 9y agoIs DO using Ceph for their block storage solution?
- antongribok 9y agoNot the OP, but I've been running production Ceph clusters for the past 4.5 years at two different fortune-50 companies. We've had very good success with Ceph for block storage and a fairly rough time with it for object storage. We're currently doing our best to try to improve it (both our own upstream contributions and collaborating with RedHat). From a technology standpoint, I think it is very interesting and for the most part has a lot of very good engineering. However, it is fairly complex and even today it's very easy to have a hard time with it when starting out. You really need to pay attention to every detail and your hardware selection is extremely important. It is extremely resilient, and goes to great lengths to preserve your data. Ceph can be performant, however that requires very good hardware and network. My experience is limited up to the Jewel release (we haven't upgraded to Luminous and we are not planning on using BlueStore anytime soon).
- rodgerd 9y ago> However, it is fairly complex and even today it's very easy to have a hard time with it when starting out. Sage's talk at LCA covered the work they're doing here; https://www.youtube.com/watch?v=GrStE7XSKFE https://www.youtube.com/watch?v=GrStE7XSKFE But yes, at small scale Gluster is still a lot easier to deploy and run.