2 ms·
to me the goal of having 'ps' show everything in the cluster is kind of cute at best. but yes, right now you install all kinds of services and support nodes, a
by convolvatron 13d ago
to me the goal of having 'ps' show everything in the cluster is kind of cute at best.
but yes, right now you install all kinds of services and support nodes, and distributed filesystems, and orchestration tools and software management tools and job schedulers. every cluster is a bespoke mess that takes a large staff to maintain and is always broken.
this is kind of the projection of web software onto hpc clusters.
in the 90s it wasn't nearly as hard to run a supercomputer because it was an integrated software platform. they were still a lot dodgier than the needed to be. but you could run a 64k node system with one support person, and there weren't really very many support tickets because things just mostly worked.
but I certainly can't imagine trying to do a company that solves this problem. or I should say I keep trying to and just seeing failure. part of the issue is that the people that you are selling to are personally and monetarily invested in the status quo. I don't think they _want_ to relieved of the burden of messing around with Kubernetes all the time, and having distributed filesystems that need to be nursed all the time, or having provisioning tools that have a 80% success rate and take hours to spin up a node.
- toast0 12d ago> but I certainly can't imagine trying to do a company that solves this problem. or I should say I keep trying to and just seeing failure. Yep, I don't think a system for this is something you can sell by itself, unless you're selling mainframes. You'd need to be selling an application that needs the system to run. I don't really know what kind of application really needs SSI clustering though. I've used clustering in a few different domains, but never really felt like SSI would be a clear win. Maybe it's chicken and egg though? If SSI clustering was more available, would applications that really leveraged it arise? The process migration + use spare capacity on the corporate network case feels potentially compelling, but only with the right kinds of workloads ... and that would be hard to sell.
- convolvatron 12d agofwiw I think it would reduce a lot of friction for the end user and probably increase utilization quite a bit. my new take on this problem is to arrange that shared resources are dynamically associated with your environment, rather than being assembled into one giant shared environment. so that if you run a program on your laptop that uses large scale compute resources, those remote threads see the same OS view that a local thread would. still not a business but I think potentially valuable. its strange how many solutions fall into that space.