26 ms·
Slack’s migration to a cellular architecture
- skullone 3y agoSo they used a feature built into a load balancer to gracefully drain traffic from specific availability zones? Odd that a feature found in load balancers from the last 25 years is a blog post worthy thing.
- progbits 3y agoThe other bit is separating the service into isolated cells so issues in one don't affect dependent services everywhere like they had experienced before. But yeah any good SRE could point this out years ago.
- skullone 3y agoJust odd a company worth billions and billions of dollars is just now discovering HA models standard since the 90s. Can expand the Clos network architecture to these distributed service applications too. But judging by Slack's client quality, mature concepts such as those must be new to them.
- deleted 3y ago[deleted]
- antoniojtorres 3y agoThe linked AWS article specifically explains that it’s not just the typical single load balancer for cross AZ routing. I frankly don’t know where you’re getting that this means that HA is new to them.
- skullone 3y agoOf course this isn't a typical single load balancer for cross AZ - but the general gist of their "new" architecture is first principles level of design. But sure, we can celebrate their minor achievement I guess
- jameshart 3y agoThat seems like a shallow dismissal. In a distributed system, making sure that sub requests are handled across distributed nodes within the local AZ, and correctly draining traffic from AZs with partial component service outages, is not as trivial as 'using a feature built in to a load balancer'.
- skullone 3y agoIt may be shallow, but architecting for this is not really "advanced, FAANG-only accessible methodology". I'm surprised their services have been as "reliable" as they have been considering such trivial stuff is just now being employed in their architecture.
- robertlagrant 3y ago> architecting for this is not really "advanced, FAANG-only accessible methodology" Sorry - where are you quoting this claim from?
- colmmacc 3y agoClose but I don't think it's quite 25 years! I added graceful draining to Apache httpd's mod_proxy and mod_proxy_balancer either in 2003 or 2004, and at the time I'm nearly certain it was the first software load balancer to have the feature, and it wasn't available on the hardware load balancers of the time that I had access to ... though I later learned that at least BigIP load balancers had the feature. At the time, we had healthy debates about whether the feature was useful enough to justify additional complexity, and whether there could be cases where it would backfire. To this day, it's an underused feature. I still regularly run into customers and configurations that cause unnecessary blips to their end-users, so it's nice to see when people dig in and make sure that the next level of networking is working as well as it can.
- robertlagrant 3y agoWell played, HN.
- skullone 3y agoI migrated some old BigIP load balancers over to Apache in 2004ish, and extended some of mod_proxy to do some "unholy" things at the time. We also did a lot of direct server return stuff when no load balancer you could buy could handle the amount of traffic statefully. Man, how times have changed, and lesson forgotten.
- djbusby 3y agoMicrosoft bought Convoy in 1998[0]. Then incorporated it into NT4sp6a and Win2k as NLB/WLBS. One of its features was to gracefully remove a server from the cluster after all connections were closed - draining. But, cluster not the same as an LB. [0] https://news.microsoft.com/1998/08/24/microsoft-corp-acquires-valence-research-inc/ https://news.microsoft.com/1998/08/24/microsoft-corp-acquire...
- deleted 3y ago[deleted]
- alberth 3y agoIs Slack still written in Hack/PHP?
- aftbit 3y agofrom the article: >Slack does not share a common codebase or even runtime; services in the user-facing request path are written in Hack, Go, Java, and C++.
- skullone 3y agoMan what a mess. Meanwhile, everyone else can extend a library used by their common services in a common language trivially.
- rs_rs_rs_rs_rs 3y agoLet me guess, they should rewrite everything in Javascript?
- skullone 3y agoWoe is us if they actually did.
- tomrod 3y agoNah, Excel. /s
- HumanOstrich 3y ago[flagged]
- wmf 3y agoAlmost everyone embraced polyglotism and microservices together.
- hotnfresh 3y agoMeh. As long as you’ve got a good, typed interface for passing messages between them and for having a common understanding of (and versioning system for) key data structures, that’s fine for this sort of thing where it’s largely processing steams of small messages and events. … but it’s probably JSON and some JSON-Schema-based “now you have two problems” junk instead of what I described. In which case, yeah, ew, gross. Unless they’ve made some unusually good choices.
- anonshadow 3y ago[flagged]
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- heywhatupboys 3y agoIs Slack dead? unironically. Does it have a future? With Teams, etc. coming out, it seems most companies do not want to go the Slack route
- aftbit 3y agoClearly no. Legacy inertia will carry it pretty far, even if literally nobody new tries to sign up for it. Our team is still using Slack and has no plans to migrate away at the moment.
- robertlagrant 3y agoTeams is doing well because it's often an IT department's simplest choice, but I don't find it's great for users.
- zo1 3y agoWhy would I choose Slack for my employees when Teams integrates so nicely with everything else in the "stack". Teams is leaps and bounds ahead already, and Slack really lost the boat many years ago. Speaking of which, I'm going now to buy more Microsoft shares.
- Shared404 3y ago> Why would I choose Slack for my employees when Teams integrates so nicely with everything else in the "stack". Does it really though? In my experience teams has a buggy integration with other things in the stack. And Teams itself ia massively buggy and a resource hog for the whole time I've used it.
- zdragnar 3y agoI don't think I have ever heard someone favorably compare teams chat with slack before. Even when I worked at a company that used teams for video calls and MS for email and calendar and documents and what not, everyone used slack for chat. I don't think anyone was sad that slack didn't integrate with the other MS services "stack".
- 3y ago
- aftbit 3y agoHow can such an architecture function with respect to user data? If the DB instance primary handling your shard is in AZ-1 and AZ-1 gets drained, how can your writes continue to be serviced?
- progbits 3y agoUsually in distributed strongly consistent and durable systems, data is not considered committed until it has been persisted in multiple replicas. So if one goes down nothing is lost, but capacity and durability is degraded.
- skybrian 3y agoThat makes sense on its own, but doesn’t it mean that there are lots of network requests happening between silos all the time? It doesn’t seem very siloed. Or is this some lower-level service that “doesn’t count” somehow?
- progbits 3y agoIt's siloed that if one is down others are not affected as long as enough other replicas are healthy to keep the quorum. You always need cross-AZ traffic, otherwise your data is single homed (which we used to call "your data doesn't exist").
- gibb0n 3y agoI think you are right, probably some service like kinesis of kafka that is keeping them in sync. that 'doesn't count'
- dexwiz 3y agoMultiple tiers of redundancy. There is usually redundancy within the AZ and then a following copy in another AZ. Usually at least four copies exist for a tenant.
- deleted 3y ago[deleted]
- madduci 3y agoCellular architecture? They've just rediscovered the art of redundancy systems
- skullone 3y agoBut if they call it cellular architecture, it sounds much more exotic than a shared-nothing active/active service!
- politelemon 3y agoIt's a common pattern in tech. Everything old will be new again.
- donutshop 3y agoBut kubernetes
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- gumballindie 3y agoOh hey they now have a new buzzword to sell!
- Terretta 3y agoIndeed, for 20+ years of distributed data centers (remember AZs are generally separate DCs near a city but on different grids, regions are geographically disparate cities) we called it "shared nothing" architecture pattern. Here's AWS's 2019 guide for financial services in AWS, where the isolated stack concept is referenced under parallel resiliency section and called "shared nothing": https://d1.awsstatic.com/Financial%20Services/Resilient%20Applications%20on%20AWS%20for%20Financial%20Services.pdf https://d1.awsstatic.com/Financial%20Services/Resilient%20Ap...
- athoscouto 3y agoTheir siloing strategy, which I'll roughly refer as resolving a request from a single AZ, is a good way to keep operations and monitoring simple. A past team of mine managed services in a similar fashion. We had a couple (usually 2-4) single AZ clusters with a thin (Envoy) layer to balance traffic between clusters. We could detect incidents in a single cluster by comparing metrics across clusters. Mitigation was easy, we could drain a cluster in under a minute, redirecting traffic to the other ones. Most traffic was intra AZ, so it was fast and there was no cross-AZ networking fees. The downside is that most services were running in several clusters, so there was redundancy in compute, caches, etc. When we talked to people outside the company, e.g. solution architects from our cloud provider, they would be surprised at our architecture and immediately suggest multi-region clusters. I would joke that our single AZ clusters were a feature, not a bug. Nice to see other folks having success with a similar architecture!
- ComputerGuru 3y agoIt sounds like you didn’t have persistent data, and were only offering compute? If there’s no need for a coherent master view accessible/writeable from all the clusters, there would be no reason to use multi-region cluster whatsoever.
- athoscouto 3y agoWe did. But the persisted data didn't live inside those ephemeral compute clusters though.
- lordofnorn 3y agoYeah, keep stateful stuff and stateless stuff separate; separate clusters, network spaces, cloud accounts, likely a mix of all that. Clearly define boundaries and acceptable behavior within boundaries. Setup up telemetry and observability to monitor for threshold violations. Simple. Right?
- catchnear4321 3y ago
- chrisweekly 3y agoI appreciate the clear explanation of the problem and the solution, which (as is so often the case) seems fairly simple or obvious in retrospect. Semi-related tangent: sometime around mid-2016, I came across a tool that helped visualize requests in near real-time, and showed what it "looks" like (ie, flow slows to trickle in service A during draining, while it ramps up in service B)... there was a really compelling demo, but I never bookmarked it and can't seem to find it. IIRC its name was a single word. Maybe someone reading this will know what I'm talking about... ?
- deleted 3y ago[deleted]
- mrkeen 3y agoVizceral
- xwowsersx 3y agoNeat. > If a graph of nodes and edges with data about traffic volume is provided, it will render a traffic graph animating the connection volume between nodes. How would one go about providing such a graph? :)
- chrisweekly 3y agoYES! Thank you!! :) This kind of shared communal knowledge is one of many reasons I'm very grateful for the HN community.
- ThePhysicist 3y agoSo they run everything in AWS USE1? That doesn't seem very redundant, but then I guess if the whole of USE1 goes down Slack won't be the only service that will be affected.
- radicality 3y agoIsn’t the point of the article that they don’t? And it describes how they implemented region drains to traffic shift between the different regions. edit: Hmm or maybe not? I still sometimes confuse aws terminology. Perhaps it is all in us-east—1, just in different availability zones (buildings?)
- ThePhysicist 3y agoIf I understand it correctly they have an edge network for ingress traffic but host all of their core services in a single AWS region (USE1) in multiple availability zones there.
- jldugger 3y ago>edit: Hmm or maybe not? I still sometimes confuse aws terminology. Perhaps it is all in us-east—1, just in different availability zones (buildings?) Correct, us-east-1 has several AZs, names like us-east-1a, us-east-1b etc. IIRC us-east-1 has six of them now.
- messe 3y agoAWS also uses Slack internally, so add that to the list of shit that can hit the fan if us-east-1/IAD goes down.
- mynameisvlad 3y agoDon’t they also use Chime? It wouldn’t be a single point of failure.
- skullone 3y agoLots of teams use Slack as well. Oddly enough, I didn't mind Chime as an end-user, but 6 years ago their API features were somewhat lacking.
- enduser 3y ago[flagged]
- danielovichdk 3y ago"A single Slack API request from a user (for example, loading messages in a channel) may fan out into hundreds of RPCs to service backends, each of which must complete to return a correct response to the user." Not being a dick here but is this not a fairly obvious flaw? I mean why not keep a structured "message log" of all channels of all time ? For every write the system updates the message log. I am guessing and making assumptions I know.
- skullone 3y agoXMMP was extensible to support all this in the early 2000s. Slack reinvented simple services in the most obtuse way. I have to use Slack and I sideline quarterback all the ways things could have been better every day.
- zx8080 3y agoXMPP Agree with this point of view. Except the Jabber/XMPP Cisco legal thing, there's just no tech answer on why on earth Slack did not use XMPP under the hood.
- alberth 3y agoWhat’s even more interesting is … WhatsApp is XMPP/ejabberd based. Slack would have known about WhatsApp architecture because it was widely talked about pre-FB acquisition (2014). And Slack was founded in 2013.
- zx8080 3y agoProbably, it's the result of schizo-histerical tech decision process happening in some companies "we need fancy tech and NOT THAT XML". Sometimes it's for OKRs and power/politics balance between departments and teams in an enterprise. "If an existing tech like XMPP is used then no serious development could be needed" fear (which is not really true). It can lead (and leads) to a huge waste of resources and overspending. But it's a bit similar to building a luxury house. Not because an owner needs it. But because he can afford it.
- deleted 3y ago[deleted]
- inertially 3y ago[dead]
- purpleturtle22 3y agoCan someone ELI5 the difference between using AWS availability zone affinity and then simply dropping the downed AZ at the top most routing point? Wouldn't that be the same thing, with the obvious caveat you are t using the routing technology Slack is using (We don't - We use vanilla AWS offerings)
- ec109685 3y agoIsn’t that exactly what they are doing? Keeping requests within an AZ and instead of using DNS at the first hop into AZ, they use envoy to control traffic shaping and making that initial decision if traffic needs to be routed away.
- t0mas88 3y agoThey decided to use every routing tool available at least once in their setup, so they can't do this. But there is no explanation in the blog about why they use so many platforms and so many routing tools. Sounds to me like they got themselves into a mess and decided to continue on that path.
- jonathankoren 3y agoSomewhere, an engineering “leader” is going to point to this blog post and then say, “Well, that’s how Slack did it!” and promptly copy this overwrought system
- Terretta 3y ago
- t0mas88 3y agoThey got themselves into a mess here: > This turns out to have a lot of complexity lurking within. Slack does not share a common codebase or even runtime; services in the user-facing request path are written in Hack, Go, Java, and C++. This would necessitate a separate implementation in each language. This sounds crazy. I've seen several products where there is a core stack (e.g. Java) and then surrounding tools, analytics etc in Python, R and others. But why would you create such a mess for your primary user request path? Sure, they're not "just a chat app" they have video, file sharing etc included and a lot of integrations. But still this sounds like a company that had too much money and too little sense while growing rapidly.
- skullone 3y agoIt even pains me to see they're suffering from so many own goals. And it's unfortunately reflected in the poor experience using the Slack client. Not to mention the multiple deprecated bot/integration APIs with such bad feature parity between all the different ways to integrate your own tooling into Slack.
- nostrebored 3y agoWhat do you mean? Slack is one of the most responsive and reliable tools I touch every day.
- skullone 3y agoHow slow are the rest of your tools? The Slack client probably performs worse today than it did a few years ago. It has the laggiest interface of any of my tools, you can watch your CPU spike to 60-80% just switching channels. Just do it right now, open up htop/top/atop/Activity Monitor - whatever you want, and just switch channels. Laugh as the Slack client wastes a universe's worth of time just... rendering a DOM with plain text. It is genuinely pathetic how bad the client is.
- lopkeny12ko 3y agoI hope this is satire. Slack is one of the slowest work tools I've ever used. Every interaction and click visibly lags. It's a sad state of the world that almost every application now is written in Javascript and deployed with Electron, and massive memory usage and slow UIs have become accepted as the norm. Try any IRC client and tell me, with a straight face, that Slack is just as responsive.
- gumballindie 3y ago"cellular architecture" What? Does amazon need to push for new sales points or are they simply making up architectures now?
- mike_hock 3y agoWhy is than an either/or?
- ignoramous 3y agoex-AWS here May be marketing but it is an architecture born out of Amazon's (and AWS's) use of AWS: - Reliable scalability: How Amazon.com scales in the cloud, https://www.youtube.com/watch?v=QeW9wCB36ck&t=993 https://www.youtube.com/watch?v=QeW9wCB36ck&t=993 (2022) - How AWS minimizes the blast radius of failures, https://youtu.be/swQbA4zub20 https://youtu.be/swQbA4zub20 (2018) For massive enterprise products like Slack that need close to 100% uptime across all their services, cells make sense.
- mike_hock 3y agoCells, interlinked.
- gumballindie 3y agoYeah that's what microservices were meant to achieve. Suppose the market is staturated with "microservices", so a new term was needed.
- ignoramous 3y agoMicroservices is one reason you need cells. If you haven't, the second talk I linked to might interest you.
- deleted 3y ago[deleted]
- diarrhea 3y agoA big term for a simple design principle indeed. But their implementation isn’t as grim as what I had initially envisioned when hearing that term. I immediately thought of Smalltalk and the idea of objects sitting next to each other, forming a graph (of no particular structure… just a graph), passing messages to neighbours. Like cells in an organism send hormones and whatnot. That makes for a huge mess that cannot be reasoned about, hence why we instead went with stricter structures like trees for (single) inheritance. That’s much closer to this silo approach, which seems nice and reasonable (although I get the impression considerable complexity was swept under the rug, like global DB consistence; the siloes cannot truly be siloed).
- dr_kiszonka 3y agoNice write-up! If no new requests from users are arriving in a siloed AZ, internal services in that AZ will naturally quiesce as they have no new work to do. Not necessarily because, due to some bug, there may be resource-hungry jobs running indefinitely. (Slack's engineers must have considered this; I am just nitpicking this particular part of the text.)
- ninkendo 3y agoIf you replace "because" with "if", your comment makes more sense. "If" there are such bugs, you are right, but such bugs might not exist.
- fiddlerwoaroof 3y agoThe thing I don't understand about Slack is how the core functionality seems to have continuously degraded since I started using it in ~2015. When I started using it, its core message sending features basically didn't have the issues with delayed messages or failure to send that I had experienced with competitors. Now, I routinely have to reset the app/clear the cache and go through various dances to get files to upload reliably (add the file to a message, wait five or ten seconds, then hit send). It's nice to see these technical write-ups about improving the infrastructure behind Slack, but I'd like to see fewer feature launches and more stability improvements to make the web, desktop and mobile apps feel like reliable software again. (nice to haves would be re-launching the XMPP and IRC bridges)
- tmpX7dMeXU 3y agoNot to “works on my machine” you, but I…genuinely do not have these problems. I’ve never heard it from my team either. So we could at the very least say it’s not a widespread global issue. Even the percentage of nerds that would want IRC or XMPP bridges back would have to be vanishingly small. I’d be annoyed if Slack reimplemented such functionality because it no doubt slows down future development. Slack has a number of mechanics that do not carry across to IRC or XMPP, and they did when they killed the bridges. I’d be annoyed if new features were compromised to increase compatibility with this blatant nerd vanity project.
- fiddlerwoaroof 3y agoSo, it's workspace and user/device specific: two of the workspaces I interact with regularly have these problems and the problems also show up intermittently for some users and not others. (Anecdotally, my experience is that Matrix/Element used to be annoying compared to the Slack experience and now I mostly prefer it to Slack) I would be fine with the understanding that the IRC bridge was missing functionality (and it always was). Although threads might make it impossible to implement in a nice way now. As far as new features go, I don't want any new features in Slack: it worked exactly like I wanted it to seven years ago and the new stuff is nice, but not worth the degradation in user experience.
- UncleOxidant 3y agoInitially read this as: "Slack's Migration to Cellular Automata" and now I'm a little disappointed.
- deleted 3y ago[deleted]
- random3 3y agoThis brings back memories - we speced an open distributed operating system called Metal Cell and built an implementation called Cell-OS. It was inspired by the "Datacenter as a computer" paper, but built with open-source tech. We had it running accross bare metal, AWS and Azure and it one of the key aspects was that it handled persistent workloads for big data, including distributed databases. Kubernetes was just getting built when we started and was supposed to be a Mesos scheduler initially. I assumed Kubernetes would get all the pieces in and make things easier, but I still miss the whole paradigm we had almost 10 years ago. This is retro now :) https://github.com/cell-os/metal-cell https://github.com/cell-os/metal-cell https://github.com/cell-os/cell-os https://github.com/cell-os/cell-os
- memefrog 3y ago"For example slack is an incredibly successful product. But it seems like every week I encounter a new bug that makes it completely unusable for me, from taking seconds per character when typing to being completely unable to render messages. (Discord on the other hand has always been reliable and snappy despite, judging by my highly scientific googling, having 1/3rd as many employees. So it's not like chat apps are just intrinsically hard.) And yet slack's technical advice is popular and if I ran across it without having experienced the results myself it would probably seem compelling." https://www.scattered-thoughts.net/writing/on-bad-advice/ https://www.scattered-thoughts.net/writing/on-bad-advice/
- lordgrenville 3y agoI really wouldn't judge the quality of some company's technical advice based on one person's experience with their UI. For almost any consumer software that gets mentioned here, you will find some people who love it and lots of others with gripes. And for e.g. Slack might have bad product/UI people but very good infra people. Better to look at TFA and judge it on its merits.
- vorpalhex 3y agoThe proof is in the pudding, not the recipe blog post.
- memefrog 3y agoIt isn't one person's experience with the UI. It is everyone's. If you don't think Slack is slow then you have forgotten what "slow" means. It is a chat program. It is incredibly simple. It is not doing anything complicated. We have gigabit internet, CPUs with multi-GHz clocks and high IPC rates, NVMe 4 SSDs that load data from disk almost instantly. It should open in milliseconds, not several seconds. That it ever takes a noticeable amount of time to do anything reveals deep flaws in Slack's engineering culture, because it shows they just don't care about performance at all. If they had "good infra people" then their program wouldn't sit and spin for seconds, ever.
- awinter-py 3y agotheir backend being on 2G explains a lot of other stuff about their software
- asankama 3y agoDelighted to be part of this conversation on cell-based architecture. As the author of the cell-based reference architecture https://github.com/wso2/reference-architecture/blob/master/reference-architecture-cell-based.md https://github.com/wso2/reference-architecture/blob/master/r..., I'm here to share insights on this exciting approach. Cell-based architecture introduces modular 'cells' into software systems, each with distinct APIs. This design fosters loose coupling and scalability – key for today's dynamic software landscape. Particularly, for those intrigued by microservices, cells align seamlessly with the independent, scalable components that power microservices architectures. Curious to dive deeper? If you're keen to explore the nitty-gritty technicalities, I invite you to check out the architecture paper https://github.com/wso2/reference-architecture/blob/master/reference-architecture-cell-based.md https://github.com/wso2/reference-architecture/blob/master/r... for an in-depth understanding. Let's kick-start this dialogue on the potential of cell-based architecture and its impact on modern software design. Feel free to join the conversation!
- DandyDev 3y agoThis seems written by AI? And as such it comes across as not genuine
- dikei 3y agoIt's from WSO2, a company selling enterprise middleware. Of course, the paper would read like something you bring to a sales meeting, not a tech talk.
- asankama 3y agoInteresting feedback, I'll see how I can address this perception of a 'sales pitch' during the next revision. The intention was to define a vendor and technology-neutral reference architecture to the community.
- asankama 3y agoNo, it is not written by AI :). You can look at the Reference Implementations section of the spec to find who is using the spec https://github.com/wso2/reference-architecture/blob/master/reference-architecture-cell-based.md#reference-implementations https://github.com/wso2/reference-architecture/blob/master/r...
- 708145_ 3y agoIf each AZ is siloed, then how can different AZs serve the same user/workspace?
- gibb0n 3y agoSeems to be the same collection of services deployed in different AZs with a load balancer? The trick would be how data is replicated across the instances, which I'm guessing is some sort of event publishing or even backup sources of truth? It says that will come in the next article and surely that's the more interesting part than the load balancing... Also explains to me why new features would take a while to roll out of you are cautiously updating instances/AZs one by one
- Cort3z 3y agoIs “cellular architecture“ hipsterspeak for micro services?
- qlkjwenf 3y agoThanks to EU, Microsoft Teams replaced Slack. GDPR makes it way too difficult to work with multiple software vendors, so companies usually only choose products from the absolute minimum number of vendors (even if there are better options). Also Slack asks too much money for what it does.