11 ms·
Hey, creator here. Thanks for sharing this! Uncloud[0] is a container orchestrator without a control plane. Think multi-machine Docker Compose with automatic W
by psviderski 10mo ago
Hey, creator here. Thanks for sharing this!
Uncloud[0] is a container orchestrator without a control plane. Think multi-machine Docker Compose with automatic WireGuard mesh, service discovery, and HTTPS via Caddy. Each machine just keeps a p2p-synced copy of cluster state (using Fly.io's Corrosion), so there's no quorum to maintain.
I’m building Uncloud after years of managing Kubernetes in small envs and at a unicorn. I keep seeing teams reach for K8s when they really just need to run a bunch of containers across a few machines with decent networking, rollouts, and HTTPS. The operational overhead of k8s is brutal for what they actually need.
A few things that make it unique:
- uses the familiar Docker Compose spec, no new DSL to learn
- builds and pushes your Docker images directly to your machines without an external registry (via my other project unregistry [1])
- imperative CLI (like Docker) rather than declarative reconciliation. Easier mental model and debugging
- works across cloud VMs, bare metal, even a Raspberry Pi at home behind NAT (all connected together)
- minimal resource footprint (<150MB ram)
[0]: https://github.com/psviderski/uncloud https://github.com/psviderski/uncloud
[1]: https://github.com/psviderski/unregistry https://github.com/psviderski/unregistry
- olegp 10mo agoHow's this similar to and different from Kamal? https://kamal-deploy.org/ https://kamal-deploy.org/
- psviderski 10mo agoI took some inspiration from Kamal, e.g. the imperative model but kamal is more a deployment tool. In addition to deployments, uncloud handles clustering - connects machines and containers together. Service containers can discover other services via internal DNS and communicate directly over the secure overlay network without opening any ports on the hosts. As far as I know kamal doesn’t provide an easy way for services to communicate across machines. Services can also be scaled to multiple replicas across machines.
- olegp 10mo agoThanks! I noticed afterwards that you mention Kamal in your readme, but you may want to add a comparison section that you link to where you compare your solution to others. Are you working on this full time and if so, how are you funding it? Are you looking to monetize this somehow?
- psviderski 10mo agoThank you for the suggestion! I’m working full time on this, yes. Funding from my savings at the moment and don’t have plans for any external funding or VC. For monetisation, considering building a self-hosted and managed (SaaS) webUI for managing remote clusters and apps on them with value-added PaaS-like features.
- olegp 10mo agoThat sounds interesting, maybe I could help on the business side of things somehow. I'll email you my calendar link.
- psviderski 10mo agoAwesome, will reach out!
- cpursley 10mo agoThis is neat, regarding clustering - can this work with distributed erlang/elixir?
- psviderski 10mo agoI don't know what the specific requirements for the distributed erlang/elixir but I believe the networking should support it. Containers get unique IPs on a WireGuard mesh with direct connectivity and DNS-based service discovery.
- jabr 10mo ago
- topspin 10mo ago"I keep seeing teams reach for K8s when they really just need to run a bunch of containers across a few machines" Since k8s is very effective at running a bunch of containers across a few machines, it would appear to be exactly the correct thing to reach for. At this point, running a small k8s operation, with k3s or similar, has become so easy that I can't find a rational reason to look elsewhere for container "orchestration".
- nullpoint420 10mo ago100%. I’m really not sure why K8S has become the complexity boogeyman. I’ve seen CDK apps or docker compose files that are way more difficult to understand than the equivalent K8S manifests.
- esseph 10mo agoManaging hundreds or thousands of containers across hundreds or thousands of k8s nodes has a lot of operational challenges. Especially in-house on bare metal.
- sceptic123 10mo agoI don't think that argument matches with they "just need to run a bunch of containers across a few machines"
- lnenad 10mo agoBut that's not what anyone is arguing here, nor what (to me it seems at least) uncloud is about. It's about simpler HA multinode setup with a single/low double digit containers.
- esseph 10mo ago> I’m really not sure why K8S has become the complexity boogeyman. Was what i was responding to. It's not the app management that becomes a pain, it's the cluster management, lifecycle, platform API deprecations, etc.
- mosselman 10mo agoYou have a graph that shows a multi provider setup for a domain. Where would routing to either machine happen? As in which ip would you use on the dns side?
- calgoo 10mo agoNot OP, but you could do "simple" dns load balancing between both endpoints.
- psviderski 10mo agoAs I mentioned in the sibling comment, please note that in this case you only get round-robin, not failover. If one of the addresses is down, the DNS record will continue returning it and users will hit a dead end. A proper load balancer or Cloudflare DNS proxy would handle this.
- psviderski 10mo agoFor the public cluster with multiple ingress (caddy) nodes you'd need a load balancer in front of them to properly handle routing and outage of any of them. You'd use the IP of the load balancer on the DNS side. Note that a DNS A record with multiple IPs doesn't provide failover, only round robin. But you can use the Cloudflare DNS proxy feature as a poor man's LB. Just add 2+ proxied A records (orange cloud) pointing to different machines. If one goes down with a 52x error, Cloudflare automatically fails over to the healthy one.
- snthpy 10mo agoI looked into this yesterday for making Caddy HA on my Proxmox cluster and stumbled upon keepalivd. It will provide you with a virtual IP and failover but not load balancing so you'd need to still point that at something like HAProxy for that. Could be something interesting to integrate though.
- woile 10mo agodoes it support ipv6?
- psviderski 10mo agoThere is an open issue that confirms enabling ipv6 for containers works: https://github.com/psviderski/uncloud/issues/126 https://github.com/psviderski/uncloud/issues/126 But this hasn’t been enabled by default. What specifically do you mean by ipv6 support?
- miyuru 10mo ago> What specifically do you mean by ipv6 support? This question does not make sense. This is equivalent to asking "What specifically do you mean by ipv4 support" These days both protocols must be supported, and if there is a blocker it should be clearly mentioned.
- justincormack 10mo agoHow do you want to allocate ipv6 addresses to containers? Turns out there are lots of answers. Some people even want to do ipv6 NAT.
- lifty 10mo agoA really cool way to do it is how Yggdrasil project does it (https://yggdrasil-network.github.io/implementation.html#how-are-nodes-identified https://yggdrasil-network.github.io/implementation.html#how-...). They basically use public keys as identities and they deterministically create an IPv6 address from the public key. This is beautiful and works for private networks, as well as for their global overlay IPv6 network. What do you think about the general approach in Uncloud? It almost feels like a cousin of Swarm. Would love to get your take on it.
- GoblinSlayer 10mo agoLike docker? --fixed-cidr-v6=2001:db8:1::/64
- zbuttram 10mo agoVery cool! I think I'll have some opportunity soon to give it a shot, I have just the set of projects that have been needing a tool like this. One thing I think I'm missing after perusing the docs however is, how does one onboard other engineers to the cluster after it has been set up? And similarly, how does deployment from a CI/CD runner work? I don't see anything about how to connect to an existing cluster from a new machine, or at least not that I'm recognizing.
- jabr 10mo agoThere isn't a cli function for adding a connection (independently of adding a new machine/node) yet, but they are in a simple config file (`~/.config/uncloud/config.yaml`) that you can copy or easily create manually for now. It looks like this: current_context: default contexts: default: connections: - ssh: admin@192.168.0.10 ssh_key_file: ~/.ssh/uncloud - ssh: admin@192.168.0.11 ssh_key_file: ~/.ssh/uncloud - ssh: administrator@93.x.x.x ssh_key_file: ~/.ssh/uncloud - ssh: sysadmin@65.x.x.x ssh_key_file: ~/.ssh/uncloud And you really just need one entry for typical use. The subsequent entries are only used if the previous node(s) are down.
- psviderski 10mo agoFor CI/CD, check out this GitHub Action: https://github.com/thatskyapplication/uncloud-action https://github.com/thatskyapplication/uncloud-action. You can either specify one of the machine SSH target in the config.yaml or pass it directly to the 'uc' CLI command, e.g. uc --connect user@host deploy
- utopiah 10mo agoNeat, as you include quite a few tool for services to be reachable together (not necessarily to the outside), do you also have tooling to make those services more interoperable?
- unixfox 10mo agoAwesome tool! Does it provide some basic features that you would get from running a control plane. Like rescheduling automatically a container on another server if a server is down? Deploying on the less filled server first if you have set limits in your containers?
- psviderski 10mo agoThank you! That's actually the trade off. There is no automatic rescheduling in uncloud by design. At least for now. We will see how far we can get without it. If you want your service to tolerate a host going down, you should deploy multiple replicas for that service on multiple machines in advance. 'uc scale' command can be used to run more replicas for an already deployed service. Longer term, I'm thinking we can have a concept of primary/standby replicas for services that can only have one running replica, e.g. databases. Something similar to how Fly.io does this: https://fly.io/docs/apps/app-availability/#standby-machines-for-process-groups-without-services https://fly.io/docs/apps/app-availability/#standby-machines-... Regarding deploying on the less filled machine first is doable but not supported right now. By default, it picks the first machine randomly and tries to distributes replicas evenly among all available machines. You can also manually specify what target machine(s) each service should run on in your Compose file. I want to avoid recreating the complexity with placement constraints, (anti-)affinity, etc. that makes K8s hard to reason about. There is a huge class of apps that need more or less static infra, manual placement, and a certain level of redundancy. That's what I'm targeting with Uncloud.
- avan1 10mo agoThanks for the both great tools. just i didn't understand one thing ? the request flow, imaging we have 10 servers where we choose this request goes to server 1 and the other goes to 7 for example. and since its zero down time, how it says server 5 is updating so till it gets up no request should go there.
- psviderski 10mo agoI think there are two different cases here. Not sure which one you’re talking about. 1. External requests, e.g. from the internet via the reverse proxy (Caddy) running in the cluster. The rollout works on the container, not the server level. Each container registers itself in Caddy so it knows which containers to forward and distribute requests to. When doing a rollout, a new version of container is started first, registers in caddy, then the old one is removed. This is repeated for each service container. This way, at any time there are running containers that serve requests. It doesn’t say any server that requests shouldn’t go there. It just updates upstreams in the caddy config to send requests to the containers that are up and healthy. 2. Service to service requests within the cluster. In this case, a service DNS name is resolved to a list of IP addresses (running containers). And the client decides which one to send a request to or whether to distribute requests among them. When the service is updated, the client needs to resolve the name again to get the up-to-date list of IPs. Many http clients handle this automatically so using http://service-name as an endpoint typically just works. But zero downtime should still be handled by the client in this case.
- doctorpangloss 10mo agohaha, uncloud does have a control plane: the mind of the person running "uc" CLI commands > I’m building Uncloud after years of managing Kubernetes did you manage Kubernetes, or did you make the fateful mistake of managing microk8s?
- oulipo2 10mo agoSo it's a kind of better Docker Swarm? It's interesting, but honestly I'd rather have something declarative, so I can use it with Pulumi, would it be complicated to add a declarative engine on top of the tool? Which discovers what services are already up, do a diff with the new declaration, and handles changes?
- psviderski 10mo agoThis is exactly how it works now. The Compose file is the declarative specification of your services you want to run. When you run 'uc deploy' command: - it reads the spec from your compose.yaml - inspects the current state of the services in the cluster - computes the diff and deployment plan to reconcile it - executes the plan after the confirmation Please see the docs and demo: https://uncloud.run/docs/guides/deployments/deploy-app https://uncloud.run/docs/guides/deployments/deploy-app The main difference with Docker Swarm is that the reconciliation process is run on your local/CI machine as part of the 'uc deploy' CLI command execution, not on the control plane nodes in the cluster. And it's not running in the loop automatically. If the command fails, you get an instant feedback with the errors you can address or rerun the command again. It should be pretty straightforward to wrap the CLI logic in a Terraform or Pulumi provider. The design principals are very similar and it's written in Go.
- snthpy 10mo agoThat's really interesting and cool. In that case calling it imperative rather than declarative is underselling it imho. I haven't worked that much with Terraform but in my usage from the cli, that is how it works too and I consider that declarative. I get that putting the declarative spec in the control plane and having the service autoreconcile continuously is another layer but this is great as a start. In fact could you not just cron the cli deployment command on the nodes and get an effective poor man's declarative layer to guard against node failures if your ok with a 1 min or 1 sec recovery objective?
- jabr 10mo ago> In fact could you not just cron the cli deployment command on the nodes and get an effective poor man's declarative layer In the project discord, a user recently experimented with a custom setup that sounds very similar to what you describe. In fact, a big part of uncloud’s appeal to me is that it also provides powerful building blocks for more complex, custom systems like this, not just the streamlined workflow for simpler, standard cases.
- tex0 10mo agoThis is a cool tool, I like the idea. But the way `uc machine init` works under the hood is really scary. Lot's of `curl | bash` run as root. While I would love to test this tool, this is not something I would run on any machine :/
- redrove 10mo ago+1 on this I wanted to try it out but was put off by this[0]. It’s just straight up curl | bash as root from raw.githubusercontent.com. If this is the install process for a server (and not just for the CLI) I don’t want to think about security in general for the product. Sorry, I really wanted to like this, but pass. [0] https://github.com/psviderski/uncloud/blob/ebd4622592bcecedbc51ad653513d91f43ef124e/internal/cli/machine.go#L45 https://github.com/psviderski/uncloud/blob/ebd4622592bcecedb...
- psviderski 10mo agoTotally valid concern. That was a shortcut to iterate quickly in early development. It’s time to do it properly now. Appreciate the feedback. This is exactly the kind of thing I need to hear before more people try it.
- tontony 10mo agoCurious, what would be an ideal (secure) approach for you to install this (or similar) tool?
- rovr138 10mo agoIt's deploying a script, which then downloads uncloud using curl. The alternative is, deploying the script and with it have the uncloud files it needs.
- yabones 10mo agoThe correct way would be to publish packages on a proper registry/repository and install them with a package manager. For example, create a 3rd party Debian repository, and import the config & signing key on install. It's more work, sure, but it's been the best practice for decades and I don't see that changing any time soon.
- Glemkloksdjf 10mo agoSo you build an insecure version of nomad/kubernetes and co? If you do anything professional, you better choose proven software like kubernetes or managed kubernetes or whatever else all the hyperscalers provide. And the complexity you are solving now or have to solve, k8s solved. IaC for example, Cloud Provider Support for provisioning a LB out of the box, cert-manager, all the helm charts for observability, logging, a ecosystem to fall back to (operators), ArgoCD <3, storage provisioning, proper high availability, kind for e2e testing on cicd, etc. I'm also aways lost why people think k8s is so hard to operate. Just take a managed k8s. There are so many options out there and they are all compatible with the whole k8s ecosystem. Look if you don't get kubernetes, its use casees, advantages etc. fine absolutly fine but your solution is not an alternative to k8s. Its another container orchestrator like nomad and k8s and co. with it own advantages and disadvantages.
- mgaunard 10mo agoThose are all sub-par cloud technologies which perform very badly and do not scale at all. Some people would rather build their own solutions to do these things with fine-grain control and the ability to handle workloads more complex that a shopping cart website.
- Glemkloksdjf 10mo ago[dead]
- RyanHamilton 10mo agoI've tried to refrain from commenting but your comment pushed me over the edge. I either want to dismiss your comment as ignorant that amazon is just a shopping cart or ignorant that you even need cloud technologies until you have 1000s of customers. But I must concede there's a chance you fall in that middle area and I'm wrong. It's < 5 percent. But yeah sure.. we have a scale problem and you're right you've identified the nonsense cloud technologies that won't fix it. I'm glad you chimed in to convince us but to build our own for 5000 customers.
- 10mo ago
- 11mariom 10mo ago> - uses the familiar Docker Compose spec, no new DSL to learn But this goes with assumption that one already know docker compose spec. For exact same reason I'm in love for `podman kube play` to just use k8s manifests to quickly test run on local machine - and not bother with some "legacy" compose. (I never liked Docker Inc. so I never learned THEIR tooling, it's not needed to build/run containers)
- TingPing 10mo agopodman-compose works fine. It’s a very simple format.
- sam-cop-vimes 10mo agoI really like what is on offer here - thank you for building it. Re the private network it builds with Wireguard, how are services running within this private network supposed to access AWS services such as RDS securely? Tailscale has this: https://tailscale.com/kb/1141/aws-rds https://tailscale.com/kb/1141/aws-rds
- psviderski 10mo agoThanks! If you're running the ucloud cluster in AWS, service containers should be able to access RDS the same way the underlying EC2 instances can (assuming RDS is in the same VPC or reachable via VPC peering). The private container IPs will get NATed to the underlying EC2 IPs so requests to RDS will appear as coming from those instances. The appropriate Security Group(s) need to be configured as well. The limitation is that you can't segregate access at the service level, only at the EC2 instance level.
- knowitnone3 10mo agobut if they already know how to use k8s, then they should use it. Now they have to know k8s AND know this tool?
- INTPenis 10mo agoWe have similar backgrounds, and I totally agree with your k8s sentiment. But I wonder what this solves? Because I stopped abusing k8s and started using more container hosts with quadlets instead, using Ansible or Terraform depending on what the situation calls for. It works just fine imho. The CI/CD pipeline triggers a podman auto-update command, and just like that all containers are running the latest version. So what does uncloud add to this setup?
- lifty 10mo agoUncloud seems much easier to manage than writing your own ansible or terraform.
- psviderski 10mo agoGreat setup! Where Uncloud helps is when you need containers across multiple machines to talk to each other. Your setup sounds like single-node or nodes that don't need to discover each other. If you ever need multi-node with service-to-service communication, that's where stitching together Ansible + Terraform + quadlets + some networking layer starts to get tedious. Uncloud tries to make that part simple out of the box. You also get the reverse proxy (Caddy) that automatically reconfigures depending on what containers are running on machines. You just deploy containers and it auto-discovers them. If a container crashes, the configuration is auto-updated to remove the faulty container from the list of upstreams. Plus a single CLI you run locally or on CI to manage everything, distribute images, stream logs. A lot of convenience that I'm putting together to make the user experience more enjoyable. But if you don't need that, keep doing what works.
- INTPenis 10mo agoIt's starting to sound a lot like k8s to me. :D Technically I could allow my web proxy to discover my services today already, but I refuse to have Traefik (in my case) running as the same user as my services. I prefer to only let them talk over TCP/IP and configure them dynamically with Ansible instead. It always amazed me that people used that feature in Traefik or Caddy, because it essentially requires your web proxy to have container access to all your other services. It seems a bit intimate to me, but maybe I'm old school.
- snthpy 10mo agoWow, this sounds very cool. I share the same concern as top comments on security but going to check out out in more detail. I wonder if you integrated some decentralized identity layer with DIDs, if this could be turned into some distributed compute platform? Also, what is your thinking on high availability and fail failovers?