11 ms·
Silos are fine, as long as there is an API between them
- deleted 3y ago[deleted]
- internetter 3y agoI thought this was going to be about siloed products. Like AWS is a silo, because once you use AWS you are effectively locked in, whereas WinterCG and Unix standards make implementors non-siloed. Oh well, this take would be a lot more spicy if so
- bazil376 3y agoI prefer my takes lukewarm
- alejoar 3y agoNew management in the last company I worked for went on a crusade against silos. Their strategy was to randomly shuffle people across teams. Literally random. You could tell how productivity bottomed and soon after a lot of senior engineers left the company. This is a multimillion american media company.
- cdavid 3y agosilos is one of this keyword that is org health dependent. When things go well, people highlight team autonomy. For the exact same setup, when things go bad, people talk about silos.
- bigbillheck 3y agoMandating that is a terrible idea but I think giving senior people (and go-getting junior ones) in team A an opportunity to do a 6-12 month rotation in team B, and so on, is a great idea.
- braza 3y agoI can resonate with the author, giving the context working on Scaleups and Big companies, where the coordination and communication plays a very heavy toil on the actual work. I do not believe in this model of “big collaboration” with folks swarming around a problem and each one with partial context trying to fix something that structurally has a huge communication, coordination and technical coupling. The best work experiences that I had was in teams that we establish the APIs, S3 buckets where the processed batch files will stay an maybe a e-mail list where someone will reply if something change.
- blastro 3y agoFirst two sentences describe my current situation. How to improve it?
- braza 3y agoI do not have a straight answer, but one thing that I saw working is the concept of “Good fences” where folks still act as a team but has clear touch points related with problems and when things goes south everyone knows what to do.
- blastro 3y agoThanks for your input!
- platformeng 3y agoNot saying these will solve your problems, but it may help give it a little perspective at least: https://www.opslevel.com/resources/optimizing-engineering-communication-for-better-developer-experience https://www.opslevel.com/resources/optimizing-engineering-co... https://fernandovillalba.substack.com/p/improving-engineering-communication https://fernandovillalba.substack.com/p/improving-engineerin...
- blastro 3y agoThank you for this
- candiddevmike 3y agoThis is one of those utopian engineer fallacies that is on the same level as reliable networking and fsync. People don't work like computers and don't respond like API endpoints. There will always be backchannels, backburners, pidgeonholes, and all the other political games humans play happening even with an "API between them". The problem is always management, and the solution is never "more management", IMO.
- dacryn 3y agoan API does not mean there are no hidden dependencies, and that is often the biggest failure. If an API call is not stateless, or requires to be chained with other calls, you're in for a world of pain in the long run
- nostrebored 3y agoBut this is equivalent to saying “a monolith can never work because it’s highly coupled”. In both cases you need to follow best practices to make things work. API design and alignment with consumers of the API is table stakes.
- DanielHB 3y agoThe whole point of the article is that silos are not intrinsically bad, partially because silos reduces communication (and therefor managers) required Two teams agreeing on an API between themselves instead of one mega team fulfilling the needs of several client teams
- throw04323 3y ago> Two teams agreeing on an API between themselves I think that depends on the type of service that the team provides. If you have a central team that many other teams interact with, they risk becoming a bottleneck. They may not be interested in maintaining custom APIs for each team interaction and you will need to agree on a contract that all can live with. Another risk is that the team providing the service also have their own backlog, including work they want to do themselves and requests from other teams. This can cause unwanted dependencies and delays where managers try to fight to be prioritized on the expense of others.
- nine_zeros 3y agoI have seen managers create task funnels for requests coming from other teams. They assume this has fixed all coordination issues. Later I see engineers having to engage week-after-week out to other teams for coordination. Things are not in place. People don't respond, coordination "APIs" don't work as documented. I have seen individuals blamed for dependency failures that they had no part in. I rarely see the obvious solution: Why don't managers spend day-in-day-out coordinating work across teams/orgs/corporate barriers? Instead of doing coordination, managers seem to be spending more time doing HR stuff like vacations, performance calibrations, documenting ways to blame engineers. They create such APIs and think their job is done. Why? How? What a severe deadweight in the company.
- jkoberg 3y agothere is a difference between silos and domains. It’s great for domains to be separated, and independently well-defined. The silos I have seen cause problems are between organizational functions, like “Product” and engineering or “The Business” and implementation teams.
- acjohnson55 3y agoMy overall reaction is that this is a great piece for products/teams that have reached significant scale, once the job to be done is too big and complex for one team to own, end-to-end, or there are truly reusable concerns that can be separated from the core product (e.g. auth, observability). Because interfacing via API is expensive. Writing APIs for others to use productively isn't easy and change management also adds a lot of overhead. And if we're talking about network APIs, there's a ton of distributed systems complexity to account for. > The problem with the DevOps movement is that it ended up taking “shifting left” to the extreme. In this sense, development teams weren’t so much empowered to deliver software faster; rather, they were over-encumbered with infrastructure tasks that were outside of their expertise. This. In truth, I think this is a major misinterpretation of DevOps, which is meant to empower devs without loading them down with incidental complexity. But I experienced exactly this misinterpretation at the first place I worked that had embraced DevOps culture.
- braza 3y ago> Because interfacing via API is expensive. Writing APIs for others to use productively isn't easy and change management also adds a lot of overhead. I agree in principle, but there is a lot of “unseen coordination/communication” costs that it’s easy taken for granted. When I was working on telecom doing interfacing with carriers (e.g T-Mobile, Verizon, etc.) on thing that I noticed was how simple was to work with those folks: Ok, this is the standard XML, those are the endpoints, that’s the list of error codes, the rate limit is X requests per second, a bunch of files will be on this FTP at 5AM daily basis, and if you face more than 100ms latency from our side just call this number. Working with “product” companies without silos most of the time it’s design by committee, folks that won’t keep the service running wanting to have a say in our payload wanting us to change our overly reliable RabbitMQ to use their Kafka.
- acjohnson55 3y ago> When I was working on telecom doing interfacing with carriers (e.g T-Mobile, Verizon, etc.) on thing that I noticed was how simple was to work with those folks: Ok, this is the standard XML, those are the endpoints, that’s the list of error codes, the rate limit is X requests per second, a bunch of files will be on this FTP at 5AM daily basis, and if you face more than 100ms latency from our side just call this number. To me, that probably reflects the maturity of the services the carriers provide. And presumably that there's an explicit customer-producer relationship? These things justify the complexity of maintaining a well curated and operated API. > Working with “product” companies without silos most of the time it’s design by committee, folks that won’t keep the service running wanting to have a say in our payload wanting us to change our overly reliable RabbitMQ to use their Kafka. If I understand what you're saying, you've experienced platform people telling you what tech to use, without having real skin in the game for operating your services? If so, that sounds very irritating. To me, a truly silo-less approach would not have that. To the extent that there are platform teams with a say in architecture, I think they should develop requirements around the external characteristics of the deliverable (performance, cost, observability, contract with other teams, etc) and largely leave the implementation concerns to the people developing and running the service.
- deleted 3y ago[deleted]
- agentultra 3y agoI have heard similar perspectives on this. That folks are moving away from devops and towards platform engineering. The idea being that a platform team reduces the friction to deploy code by building self-serve APIs, libraries, and infrastructure to be used by dev teams. Even when in teams practicing devops well I have always known at least one or two people who are great developers but they don’t want to know anything about operating systems, sockets, file descriptors, service level objectives and all that. I always found working with them to be challenging. I’m very much in the, “you wrote it, you run it,” camp. While platform engineering sounds great I’ve also worked with teams trying this and it has its own trade-offs: as demands on the platform team grow it can take longer to wait for your change requests to be deployed and depending on how ownership at the company works.. it can be frustrating: you could have fixed it yourself and shipped sooner but now you have to live within the constraints set for you by the platform team. You also end up with a development culture that has a hard time understanding service performance objectives. This can be a great thing for some companies for sure. But I haven’t seen a cure-all for siloing teams. Conway’s Law and all.
- 8organicbits 3y ago> as demands on the platform team grow it can take longer to wait for your change requests to be deployed That sounds like an ops or deployment team, not a platform team. A key feature of a platform should be that the developers choose when a deployment happens.
- Dioxide2119 3y ago> A key feature of a platform should be that the developers choose when a deployment happens. Agreed. When I was on a platform team we wrote tools to take a process that used to be done by a deployment team (change a DNS record was a helpdesk ticket) and move it into a self-serve system (PR your desired DNS changes in, upon merge, the system deploys the changes), which kept audit happy because 'dev' wasn't touching 'prod' in the unfettered way SOC2 people stay up at night worrying about (even though Enron happened because of bad managment not Office Space but anyways), while still giving Devs effective control of when and where they wanted to make production changes, whether relatively ad-hoc or as part of a CI/CD pipeline. Humans could approve the self-service PRs, or if a list of in-code rules had been fulfilled, the PR would be auto approved (and potentially even merged but everyone but us was too afraid to set that part up).
- deleted 3y ago[deleted]
- PaulHoule 3y ago... and some system for coining unique ids across all the silos, whether that is UUIDs or something like RDF-style namespaces.
- jon_richards 3y agoI was really hesitant to use uuids as primary keys because it’s basically worst case clustering performance. Ended up just using 32 bit Unix time followed by 12 random bytes. Not sure why that isn’t an official version.
- PaulHoule 3y agoIf you insist on using a B-Tree index and random UUIDs that's certainly the case. Many of the formulas used to generate UUIDs in the past had terrible privacy implications: Office 95 would fill documents with UUIDs generated based on MAC addresses and timestamps so Office documents could be tracked to particular machines until Microsoft changed this with little fanfare.
- pphysch 3y agoNothing wrong with having clear boundaries and different owners across high-quality data sources, but IME "silo" usually comes up in the context of egregious data duplication and ambiguous sources of truth.
- 0xbadcafebee 3y agoAn API is actually not good enough. It's the minimum you could possibly have to allow communication. So the silos can "work together" though that API, but problems then occur due to lack of understanding of how each silo actually works under the hood. Imagine a bunch of microservices built by different teams. They just send each other their OpenAPI specs and some URIs. So they start calling each other's APIs. Everything seems to work fine. But wait. Are there limits on this API? How many calls, or how much data can I send? What's the SLA on these transactions? What happens to data, how it's stored, processed, backed up? If I send X data to service Y, do I know service Z is going to get the same data? If one of these services goes down, is everything going to go down, and is that team staffed for 24/7 support? Do they even know what to do when things are down? When everything does down, how do these silos know which thing was the cause and who to alert to fix it? Does it require multiple silos to fix? All of that and much, much more, is deeper knowledge related to the entire system, which is the inter-relation of all these silos from top to bottom and sideways. The API doesn't solve these problems or answer the questions. The API only tells you how to do one thing. The premise of silos is the idea that you don't need to know anything else about the rest of the world but some tiny bit of information. Well, reality says different. DevOps is not about "merging teams", but communication and collaboration between teams. They should understand each other well, or at least make it much easier to discover the right information in order to improve outcomes. You absolutely have to have specialized teams where people have domain knowledge. But you also need to provide the tools and practices that enable very different teams to work together to build things right and solve problems quickly. Say you're building cars. A dealership's mechanics notice a belt keeps rubbing on a cable or hose. That needs to be quickly notified to the assembly people to see if it's an assembly problem, and if not, it needs to be sent to the mechanical engineers to address a potential design flaw. That all needs to be done quickly, as soon as the problem is noticed, because cars are shipping every day. The longer it takes for that whole loop to complete, the more bad cars are shipped. By focusing on improving the loop between all these different groups, you improve business outcomes. But "an API" isn't going to do that. That's why the idea of a "DevOps Engineer" is wrong. This isn't an engineering problem. This is a business process problem. The communication between teams is not an API, it is really organizational structure and practice. Engineers noticed the problem, and wanted to fix it, but they failed to use the language of management. So instead people slapped "Engineer" on the concept and everyone got confused.
- tanseydavid 3y agoWith two silos, wouldn't you need to have at least two APIs?
- josh-sematic 3y agoNo; one solo could provide an API that the other consumes, but the producer consumes no APIs from the consumer.
- whynotmaybe 3y agoUnless you work at a place where ops don't write code because "it's the dev's job to write code", so everything is built and done manually; and devs don't have access to any infrastructure and don't bother about it, because "it's the ops' job". And management of both sides agree with this vision. "works on my machine" and "the issue must be with the code" are the most used excuses by both sides when something fails. The ops "api" works by email, but the response delay is usually expressed in weeks. Most memorable quote from that place : "Yes, a 4 month delay for an answer about your new server might seem long"
- slaymaker1907 3y ago"We can have that new release deployed in about 6-8 weeks."
- jahsome 3y agoI find myself these days on one of these ops teams. The lead time is the same for deploying either a code or infrastructure change, and probably closer to 3-4 months here. It's not an operational capacity or competency issue for this org, it's the result of hours worth of "sync" or "review" meetings with no discernible agenda, negotiating maintenance windows, facilitating approvals from a dozen or more parties who don't even comprehend what they're approving, and weeks of manual acceptance testing. On the other extreme, in past roles at different orgs, I've been on teams doing multiple deployments to production every day, both on the dev and ops sides. I find it exhausting and soul crushing being completely untrusted because of the mistakes made by people who left the org years before I started.
- everdrive 3y agoI have no comment on this article's specific claims, and I have every reason to believe the author is both intelligent and insightful. That said, so many times in my professional life I run into the same problem: people think some idea is a truism which broadly applies to all or most situations. The problems faced are seldom deeply understood, and so that truism is misapplied by people attempting to follow the latest best practices. This seems like a pernicious meta-problem related to profitability and work resources. ie, people cannot deeply understand all problems, and so they are content to make larger errors in a number of places so long as "more work" is getting done. I don't think this problem will ever disappear.
- esafak 3y agoThis is true for some definition of a silo, but to me the term inherently implies difficulty in interfacing. If you can easily pull information out or push it in, it's not a silo in the sense people complain about. So to me this not so much a correction of the definition, but a technical resolution. There are several problems, though. The owner of the silo may benefit from it. This is called "lock in". If the company has any say, the silo owner should be properly incentivized to ensure "good citizenship".
- throwup238 3y agoThis is how it starts. First ChatGPT starts suggesting all these bloggers "Silos are fine as long as they have APIs" articles and "silo this, silo that". Soon enough the overton window has shifted and we're talking about military silos and silos having APIs in the same sentence. Before you know it, GPTskyNet5 is firing nukes off using poorly secured and thought-out webhooks setup by some random defense contractor at missile silos because some DoE scrum master needed a promotion. This is the beginning of the end!
- osigurdson 3y agoI'd suggest that silos are necessary as human communication bandwidth is limited. Everyone talking to everyone, all the time doesn't scale. The key is to create silos along natural "fault lines" such that less communication is required.
- solatic 3y agoAPIs don't live in a vacuum. Sorry, but you can't just ship a service that exposes an API and expect people to use it. It MUST be documented. And no, your auto-generated Swagger/OpenAPI docs don't cut the mustard. If, as an executive, you expect the teams you set up to actually be independent, then treat them like Product teams in their own right. Set expectations to write Product-quality documentation, at least as good as what you ship to customers. Hire internal Product managers for those teams, who will learn how internal customers use those APIs and what else they need to solve their problems. Hire internal Marketing to ensure everybody else knows that the APIs exist. Sound ridiculous? What, do you expect your Engineering teams to have these skillsets already? If you don't expect your company's product to succeed without Product and Marketing then why, pray tell, would you expect your internal products to be any different?
- nox101 3y agoWorst docs I've used recently are the chrome extension docs. They're auto generated and practically incomprehensible, to me.
- hayst4ck 3y agoSilos are not a problem. Leadership quality is a problem that silos can greatly exacerbate. Silos exacerbate two problems. The first problem is that there is an area between two silo's which can be poorly owned or lack stewards. The second problem is that the first manager above two different silos is the de facto resolver of disputes and resource application for and around those silos. This area often has no advocate, because to advocate for resource application to that area is to volunteer yourself. This area becomes a blind spot to leadership through systematic neglect. In many cases this leadership can be a checked out ex-google CTO who has never run an org under resource constraints who isn't hungry because they are already well off. Checked out Rest and Vest early employees who, through the peter principle, end up in CTO positions can also be extremely dangerous to organizations. If you have poor leadership, you end up with two silos that don't want extra responsibility, because more responsibility without more resources is a losing prospect. Under bad leadership, anything that is not feature production is not rewarded. The end result is that each silo becomes aligned against other silos, rather than aligned to a business goal. Defining an API seems like the primary goal is to completely remove the area between two silos, thus alleviating the problem of opaque ownership. I agree. Explicit ownership for all artifacts helps an organization run much more effectively. The real problem is a culture where every individual employee does not feel responsible for business outcomes with proportionate recognition by leadership of taken responsibility for those outcomes. Here is Admiral Rickover's take on culture: > Professionalism occurs when individuals act in the best interest of those being served according to objective values and ethical norms, even when an action is perceived to not be in the best interest of the individual or their organization. That is, there are times when professionals must sacrifice their own interest (or that of their organization) to meet the objective values and ethical norms of the profession. Professionals, in this sense, are serving something greater than the bureaucratic organization that employs them. > If Admiral Rickover had a mantra to shape a professional culture, it would have been, "I am personally responsible." When leadership does not practice personal responsibility or engages in blame, that ripples through the entire organization. Silos that don't function well are symptoms of leadership failing to take responsibility and cultural failure, not necessarily structural failure.
- moribvndvs 3y agoComing from a job where we transitioned to cross-functional teams, all we did was reorient the silo. The business was unhappy with teams oriented around specific projects, and felt we were struggling to produce quality releases according to the roadmap because teams had lack of understanding of the big picture along with competing priorities. With a CFT, we would focus on a vertical slice of whatever systems it took to deliver a feature, and orient our team structure around the right expertise in the systems we’d need to touch. We would also take responsibility for managing our own releases rather than relying on bottleneck-prone maintainers to do it. This seemed sensible at first, it even gave us a sudden burst of productivity. It soon fell apart, however, as now it became incredibly difficult to manage releases across products where multiple teams have competing priorities. Quality quickly plummeted and we had an even worse time releasing software, and every end of sprint the problem compounded. We still ended up with a bottleneck of a few experts to sort through the firehose of changes each team was trying to push out and get it handled. Our workload actually increased, along with stress and frustration. The issue in this case was a perfect storm of tech debt, poor planning, poor architecture, and simply too many priorities. The business thought that rotating the team structure 180 degrees magically multiplies our productivity, and somehow convinced themselves that the work these teams do are completely independent, all while building out a roadmap with very little cooperation from the engineers. On top of investing in fixing the technical and organizational issues just producing software, I would rather leadership had focused the entire business on a small set of complimentary or at least non-competing (limited by available resources, of course) priorities so we could each focus on an piece of the puzzle and bring our particular expertise to bear on producing a coordinated and high quality solution. I certainly agree that teams oriented around their microservice or whatever was bad, but orienting around a project/feature really any different; we should have been oriented around a holistic product and the customers’ well-being.