7 ms·
Torus development has been stopped at CoreOS
- aerioux 10y agocould anyone more familiar with the situation give context around the decision + what is happening moving forward? Thanks
- Zelmor 10y agoNothing of value was lost, don't worry about it. People reinventing the wheel, realised it takes a lot more than they are capable of. Same old, same old.
- parenthephobia 10y agoI can't speak to the value of Torus, but in this case the best wheels are proprietary and their design is secret. Reinventing them is reasonable.
- chuhnk 10y ago"But we didn't achieve the development velocity over the 8 months that we had hoped for when we started out, and as such we didn't achieve the depth of community engagement we had hoped for either." Open source is tough, even as a successful VC funded company. Gotta give credit to CoreOS though, rather than beating a dead horse they're acknowledging there's little external interest and their time would be better spent focused elsewhere. Seeing as they've discontinued Fleet as well it's likely they're doubling down on the commercial Tectonic product built on kubernetes. There's also likely pressure from their investors to start making money, especially as they're well into using their series B and might want to look to raise more money. Distributed file storage is a tough market in itself. Developers get excited about technology but how many need this as opposed to a highly available database? I love the building blocks of distributed systems and understand that Google's technology is built on a layering of tech (Colossus, Spanner, etc) but it seems the world is not yet ready. Everyone is already struggling to understand the complexities of this new ecosystem and how the pieces all fit together. Again, good move by CoreOS, wish them luck with their commercial strategy.
- pyb 10y agoAnd, well, it looks like it was mostly just one guy working on it.
- michaelcampbell 10y agoI missed the fleet discontinuation notice. Is there a link somewhere about it?
- parenthephobia 10y agohttps://news.ycombinator.com/item?id=13592864 https://news.ycombinator.com/item?id=13592864 was the thread here
- parenthephobia 10y ago> Everyone is already struggling to understand the complexities of this new ecosystem and how the pieces all fit together I believe Google's (and Amazon's) secret sauce is having large sysops and devops teams filled with subject matter experts - and the software's original developers. Efforts like Torus are trying to build zero-to-low-maintenance turn-key solutions, whilst Google is happy to have tens of full-time distributed storage engineers.
- fh973 10y agoFor Google at least the SRE teams for infrastructure components are not large, neither in total nor relative to the huge infrastructure that they are managing. You could call this operational scalability, it is made possible by decoupling service quality from individual pieces of hardware. The key ingredient are redundancy across relatively large failure domains and auto mated handling of all foreseeable events. Everything is built around this concept, compute with a fault-tolerant container scheduler, storage on a fault-tolerant file system, checksuming everywhere,... Another enabler is probably keeping things simple at the component level.
- vr46 10y agoI just spoke with some of the CoreOS team two days ago and Tectonic development is definitely being accelerated, which is critical given a few complete blockers in their setup (e.g. the VPC CIDR range in AWS must be 10.0.0.0/16 - Whoa there, Tex, that's gonna cause all sorts of conflicts with our internal network and Direct Connect)
- fidget 10y agoWell that seems eminently reasonable
- geku 10y agoDoes anyone have experience with https://rook.io/ https://rook.io/ which is mentioned in the message?
- usgroup 10y agoThose of us in the game for some time ultimately read "beta" to mean "30% chance of survival" so this doesn't come as a surprise but it probably isn't true for our more optimistic colleagues. I think more tempered marketing would have really helped.
- usgroup 10y agoFor reference: https://coreos.com/blog/torus-distributed-storage-by-coreos.html https://coreos.com/blog/torus-distributed-storage-by-coreos.... "Releasing today's initial version of Torus is just the beginning of our effort to build a world-class cloud-native distributed storage system..." 2016-06 "Torus development has stopped on Core OS" 2017-02 Simply pointing out what is the case.
- joshbaptiste 10y agoBcantrill on how hard such a feat would be to accomplish. https://news.ycombinator.com/item?id=11817387 https://news.ycombinator.com/item?id=11817387 https://news.ycombinator.com/item?id=11818081 https://news.ycombinator.com/item?id=11818081
- bassamtabbara 10y agodisclaimer: I work on project Rook. Yes building a whole new data path is complicated and takes many years to get right. Its also somewhat of a moving target with new storage technologies appearing on the scene (like SMR drives, NVME, and persistent memory). It would be much more effective for the community to coalesce around a common data path just like we coalesce around kernels. Something like Ceph is a great start, its battle-tested and has storage vendors (like Intel, Samsung, SanDisk and others) updating/optimizing it for new kinds of storage. The focus on simplicity and integration into cloud-native environments is critical however and Torus' vision was spot on. Kudos to the CoreOS team for raising the bar on this.
- sysexit 10y agoI called it here, right at the initial announcement: https://news.ycombinator.com/item?id=11816951 https://news.ycombinator.com/item?id=11816951 This is just too hard of a problem to solve frivolously. Kudos to CoreOS for trying, and coming to the inevitable conclusion sooner rather than later.
- epowell2017 10y agoAnyone looking at openEBS.io? This is open source scale out block for containers.
- umamukkara 10y agohttps://blog.openebs.io/torus-from-coreos-steps-aside-as-cloud-native-storage-platform-what-now-2375e7f5b145#.fslj3rwsg https://blog.openebs.io/torus-from-coreos-steps-aside-as-clo...
- dankohn1 10y agoIn addition to Rook https://rook.io/ https://rook.io/ , which CoreOS mentions and we need to add, please take a look at the other cloud-native storage options listed on the CNCF cloud native landscape: https://github.com/cncf/landscape https://github.com/cncf/landscape Disclosure: I'm executive director of CNCF, and co-author of the landscape.
- chrissnell 10y agoDo you have any details about running Rook on Kubernetes? The Rook docs link to an outdated document about running Kubernetes on CoreOS.
- josephjacks 10y agoThe folks behind Rook at Quantum have put together an operator (custom K8s controller and TPR) for Rook: https://github.com/rook/rook/tree/master/demo/kubernetes https://github.com/rook/rook/tree/master/demo/kubernetes
- hackuser 10y agoThank you. Would you share with us a quick overview of the differences?
- alrs 10y agoI understand they don't hand out Internet points for this sort of thing anymore: https://news.ycombinator.com/item?id=11816821 https://news.ycombinator.com/item?id=11816821
- wmf 10y agoNote that CoreOS has not given up on the concept of distributed storage; they just gave up on writing their own. So they haven't proved you right. I realize reliable block/file storage isn't "cloud native" but legacy apps require it and they are willing to spend billions to have it.
- alrs 10y agoOf course people want it, but can they have it? The world has yet to see a successful distributed block project.
- zzzcpan 10y ago"Of course people want it, but can they have it?" Why not? Block storage is a weird beast, targeting legacy apps unable to run in multiple data centers. The same apps are also likely to be willing to lose some of the most recent data in case of a data center outage and trade this for performance, so the barrier is already low. The storage might only need consensus somewhere close by, where latency is very good and the network capacity is huge. Nodes in other data centers could receive data asynchronously (otherwise write latency is going to render the whole thing useless anyway). The question is whether there really are billions to be paid for something that ultimately cannot do well compared to proper distributed solutions.
- epowell2017 10y agoWhat does "cloud native" mean here? To me it suggests purpose built much as Torus was /is - though the OpenEBS engineers are asserting there are really only two "container native" storage solutions going, their open source project and PortWorx. https://medium.com/@kiranmova/persistent-storage-for-containers-alternatives-to-torus-2375e7f5b145 https://medium.com/@kiranmova/persistent-storage-for-contain...
- alrs 10y agoThe next question that needs to be answered at CoreOS: "Why, exactly, are we maintaining our own Linux distro when the Go binaries that we're writing can mostly ignore userspace?"
- el_isma 10y agoFor one, CoreOS auto-updates smartly, so you can install and forget.
- alrs 10y agoThere were Linux distros that could upgrade across releases twenty(!) years ago. There is so much opportunity in infrastructure right now, so it strikes me as weird that CoreOS took a bunch of VC money to go off on an extended Linux-From-Scratch-Adventure.
- hagbarddenstore 10y agoAh.... Hahah... Hahhahahahhahahahahahahahah. No. During the 1 year I ran CoreOS in production, updates were turned off, because they caused all sorts of issues. They only reliable way of doing updates in CoreOS is to replace the machine and reconfiguring it. But then you need to automate joining etcd, which itself is a major pain in the ass.
- politician 10y agoYou don't deserve the downvotes. When Docker arbitrarily changes something important and pushes those changes to Docker Hub, CoreOS is dragged along for the ride. We had our auto-updating servers move to Docker 1.10 over a weekend. Of course, this brought down our CI/CD process because that version of Docker changed something important. Our staging environment was totally horked, but our production environment survived due to an unexplained reboot lock. We were lucky. Turn off auto updates.
- smlacy 10y agoWhat's Torus and why should I care?
- marknadal 10y agoDangit, I trust the CoreOS team more/better than a lot of people in the space. Torus would have been so useful. At the other end of the spectrum though, maybe this is reasonable? As a developer, my first thoughts for "I want my own S3" is not etcd (strong consistency) but projects like https://github.com/minio/minio https://github.com/minio/minio , or even using eventually consistent SQLite replication / synchronization tools https://github.com/gundb/sqlite https://github.com/gundb/sqlite . So that makes me ask about rook.io too, what layer of the "stack" is it trying to fit into? Obviously pretty low, but that also seems unnecessary (and part of why I suspect Torus is stopping).
- jacques_chester 10y agoAt a glance, Torus was intended to be a distributed file system. As I understand it, distributed file systems are easier than distributed block systems, but harder than distributed blob systems. A blob system is all-or-nothing. You create or replace the entire blob at once. This makes bookkeeping and replication much easier for the implementer. A filesystem supports much richer semantics, including the ability to seek parts of files and modify small regions of files. You need a lot more mechanics to maintain consistency across a network. A block store is difficult because you're trying to work at very high speed on very small units of state wooshing back and forth willy-nilly. You don't get to rely on any of the higher semantics provided by a filesystem or blobstore, since you're pretending to be a magical harddrive. I am often wrong in these matters, as an interested outsider, so I'd be happy to receive correction.
- umamukkara 10y agoDisclosure: I work for OpenEBS project Torus was intending to write distributed block storge that is container native. Metadata management using key value (KV / etcd) method is increases the complexity and not new. Ceph tried it. OpenEBS uses a novel approch, linux sparse files for managing the blocks of a volume. Fork of Rancher longhorn. The issue of managing the large scale distributed block storage metadata is solved easily throught he management of the files (not blocks). https://blog.openebs.io/torus-from-coreos-steps-aside-as-cloud-native-storage-platform-what-now-2375e7f5b145#.fslj3rwsg https://blog.openebs.io/torus-from-coreos-steps-aside-as-clo...
- KaiserPro 10y agoI have karma to burn on this, so here goes: I worked for several years in VFX/HPC. 30k+ cpus and 15pbs of storage. Firstly with storage its very rare that people want actual block storage (unless you are hosting VMs, but thats so 2007.....) Yes, I know, openstack, but that's just fucking horrific, seriously just use netboot and be done with it. I've seen people do it inside new clustereing systems, but its really not fun to do, especially if you consumer is prone to disappearing without warning. (FSCK is a terrible mechanism for fast recovery) Most apps, unless they have bought into the "shove everything over HTTP and pay the penalty", want a posix file system to store anything of importance. (yes, yes database, but where is that writing the data to?) Now, there are three ways you can do this: o use a clustered file system o Use NFS (with or without a clustered filesystem underneath) o Fuck about with iscsi/SAS/FC and dynamically map block dynamically. Using a clustered filesystem spread over many clients is begging for trouble, mainly because one client can fuck it up for everyone. Some FSs are dynamic and sexy, but they have a habit of fucking up in new and interesting ways that even the authors can't figure out. The common ground is having storage nodes attached directly to a pack of big fat disks(for streaming IO) or NVME/SSDs for random IO. They then serve out NFS traffic. Now, you can either have a clustered file system underneth, or not. (Having stand alone servers can be advantageous, if you can map your filesystem out hierarchically) Now, unless you have a Storage area network, then the last option is just begging for shit performance. You really don't want IO traffic fighting with network traffic. However, if you want raw throughput, this is the way to go, but be warned, you won't get any friendly help if you accidentally disconnect a disk. Basically, kubernetes/HPC and storage is a solved problem ducks no really, just map in NFS shares and be done with it. If its exotic, its probably going to fail hard, and in ungoogleable ways. More importantly only a few people are going to be able to help, and they may or may not still employed at your company.
- stonogo 10y agoI find your offhand dismissal of clustered filesystems, on which literally every supercomputer relies, to be a little strange. They might not have worked well for you, but "googleable" is not the bar that is generally set for the HPC problem space.
- 10y ago
- SEJeff 10y agoI kind of wonder if there will ever be a kubernetes operator built for Ceph (not rook ontop of Ceph). Besides it being a bit of a PITA to maintain, Ceph is about as good as exists regarding OSS distributed object storage currently. If they could kill some of the operational overhead via an operator that did much of it, they might have a serious winner on their hands. Note that I'm just referring to the radosgw bits for the S3 style storage API, not the posix filesystem bits.
- hunter_n 10y agoThere is some discussion on this in the ceph-docker project - https://github.com/ceph/ceph-docker/issues/472 https://github.com/ceph/ceph-docker/issues/472 Interestingly there is an unannounced project by CoreOS for a storage Operator that will handle Ceph, Gluster, etc. I'm sure we'll hear more about that now that Torus has been retired.
- SEJeff 10y agoSource for the unannounced project, or just overheard in person from someone? I don't see anything obvious on their github but it could be private.
- hunter_n 10y agoSome more info here: https://docs.google.com/document/d/1Nm3ZQXtojd7Ruw8gQ-8xNo0v1rReHl7v-KSd19UJO4I/edit https://docs.google.com/document/d/1Nm3ZQXtojd7Ruw8gQ-8xNo0v...
- 123jfeichabc 10y agoThis is good news - CoreOS needs to focus on what's most important to their core business to be successful. Being chock full of bright, relatively young and enthusiastic engineers drunk on the Golang kool-aid, there's a very real risk of getting distracted by reimplementing everything under the sun in their favorite shiny new language. Even if Torus is a good idea, CoreOS has to prioritize, commit, and execute. They can't afford too many diversions. This is a competitive space, their opportunity window and runway are both limited, as usual.