4 ms·
What I found building multiplayer editors at scale is that it's very easy to very quickly overcomplicate this. For example, once you get into pub/sub territory,
by blixt 2y ago
What I found building multiplayer editors at scale is that it's very easy to very quickly overcomplicate this. For example, once you get into pub/sub territory, you have a very complex infrastructure to manage, and if you're a smaller team this can slow down your product development a lot.
What I found to work is:
Keep the data you wish multiplayer to operate on atomic. Don't split it out into multiple parallel data blobs that you sometimes want to keep in sync (e.g. if you are doing a multiplayer drawing app that has commenting support, keep comments inline with the drawings, don't add a separate data store). This does increase the size of the blob you have to send to users, but it dramatically decreases complexity. Especially once you inevitably want versioning support.
Start with a simple protocol for updates. This won't be possible for every type of product, but surprisingly often you can do just fine with a JSON patching protocol where each operation patches properties on a giant object which is the atomic data you operate on. There are exceptions to this such as text, where something like CRDTs will help you, but I'd try to avoid the temptation to make your entire data structure a CRDT even though it's theoretically great because this comes with additional complexity and performance cost in practice.
You will inevitably need to deal with getting all clients to agree on the order in which operations are applied. CRDTs solve this perfectly, but again have a high cost. You might actually have an easier time letting a central server increment a number and making sure all clients re-apply all their updates that didn't get assigned the number they expected from the server. Your mileage may vary here.
On that note, just going for a central server instead of trying to go fully distributed is probably the most maintainable way for you to work. This makes it easier to add on things like permissions and honestly most products will end up with a central authority. If you're doing something that is actually local-first, then ignore me.
I found it very useful to deal with large JSON blobs next to a "transaction log", i.e. a list of all operations in the order the server received them (again, I'm assuming a central authority here). Save lines to this log immediately so that if the server crashes you can recover most of the data. This also lets you avoid rebuilding the large JSON blob on the server too often (but clients will need to be able to handle JSON blob + pending updates list on connect, though this follows naturally since other clients may be sending updates while they connect).
The trickiest part is choosing a simple server-side infrastructure. Honestly, if you're not a big company, a single fat server is going to get you very far for a long time. I've asked a lot of people about this, and I've heard many alternatives that are cloud scale, but they have downsides I personally don't like from a product experience perspective (harder to implement features, latency/throughput issues, possibility of data loss, etc.) Durable Objects from Cloudflare do give you the best from both worlds, you get perfect sharding on a per-object (project / whatever unit your users work on) basis.
Anyway, that's my braindump on the subject. The TLDR is: keep it as simple as you can. There are a lot of ways to overcomplicate this. And of course some may claim I am the one overcomplicating things, but I'd love to hear more alternatives that work well at a startup scale.
- jvanderbot 2y agoI've worked in robotics for a long time. In robotics nowadays you always end up with a distributed system, where each robot has to have a view of the world, it's mission, etc, and also of each other robot, and also the command and control dashboards do too, etc etc. Always always always follow parent's advice. Pick one canonical owner for the data, and have everyone query it. Build an estimator at each node that can predict what the robot is doing when you don't have timely data (usually just running a shadow copy of the robot's software), but try to never ever do distributed state. Even something as simple as a map gets arbitrarily complicated when you're sensing multiple locations. Just push everyone's guesses to a central location and periodically batch update and disseminate updates. You'll be much happier.
- mikhmha 2y agoWow, this sounds like how the AI simulation for my multiplayer game works. Each AI agent has a view of the world and can make local steering decisions to avoid other agents and self preservation. Agents carry out low-level goals that are given to them by squad leaders. A squad leader receives high level "world" objectives from a commander. High-level objectives are broken down into low level objectives distributed among squad units based on their attributes and preferences.
- rurban 2y agoWell, at least don't update multiple servers. Distributed read-only state is ok though, just updates not. They must be centralized.
- athrun 2y agoThanks for sharing your experience, and what you have found to work. Sometimes I feel we (fellow HN readers) get caught into overly complex rabbit holes, so it's good to balance it out with some down-to-earth, practical perspectives.
- NathanFlurry 2y agoDurable Objects solves so many problems for realtime & persistence. The biggest problem is they're vendor locked, so we can't use them if we want to keep all of our infrastructure on AWS. So we built out a library that lets you run Durable Object-like backends on any cloud: https://github.com/rivet-gg/actor-core https://github.com/rivet-gg/actor-core