8 ms·
Modular Monoliths Are a Good Idea
- IshKebab 2y agoTo me it seems like the main advantages of microservices are a) you can use different languages b) you can run different parts of your system on different servers I feel like you can solve both without giving up the niceties of a monolith just with a good RPC framework. A really good one would even give you the flexibility to run "microservices" as separate local threads for easy development. I've never seen anyone actually do that though.
- mattnewton 2y agoAs an ex-fang engineer myself, I have never advocated for more services, usually have pushed for unifying repos and multiple build targets on the same codebase. I am forever chasing the zen of google3 the way I remember it. If anything my sin has been forgetting how much engineering went into supporting the monorepo at Google and duo-repo at Facebook when advocating for it.
- paperplatter 2y agoDo FAANG engineers normally advocate for more services instead of fewer? I haven't gotten that impression.
- aleksiy123 2y agoSmaller services but not necessarily more binaries. The current direction I think is to build composable services that could be run together or separately. Where a service is a logical grouping of an RPC interface. Here is some public work in this direction from Google https://serviceweaver.dev/ https://serviceweaver.dev/
- paperplatter 2y agoMakes sense, since Google uses a monorepo (google3).
- paulddraper 2y agoNormally more services and fewer repos. From my experience.
- recursivecaveat 2y agoWhat were the two facebook repos? I can't find any reference to them.
- simscitizen 2y agoThe main ones were www which contained most of the PHP code and fbcode which contained most of the other backend services. There were actually separate repos for the mobile apps also.
- disgruntledphd2 2y agoFbcode then became fbsource when it acquired the android and some data code.
- ecshafer 2y agoThere seems to be more interest in building monorepo support now. Some tools, start ups, etc. I would bet Github is working on increasing support as well for large repos. So I think Google was ahead of the curve there.
- paperplatter 2y ago"To get similar characteristics from a monolith, developers need: Incremental build systems Incremental testing frameworks Branch management tooling Code isolation enforcement Database isolation enforcement" This sounds a lot like microservices, most of all the last point. Is the only difference that you don't use RPCs?
- nine_k 2y ago> the only difference that you don't use RPCs But it's a huge difference. No RPC overhead. No lost / duplicate PRC messages. All logs can literally go to the same file (via e.g. simple syslog). Local deployment is dead simple, and you can't forget to start any service. Prod deployment never needs to handle a mix of versions among deployed services. Beside that, the build step is much simpler. Common libraries' versions can never diverge, because there's one copy per the whole binary (can be a disadvantage sometimes, too). You can attach a debugger and follow the entire chain, even if it crosses the boundaries of the modules. With that, you can make self-contained modules as small is it makes logical sense. You can pretty cheaply move pieces of functionality from one module to another, if it makes better sense. It's trivially easy to factor out common parts into another self-contained module. Still you have all the advantages of fast incremental / partial builds, contained dependencies, and some of the advantages of isolated / parallel testing. But most importantly, it preserves your sanity by limiting the scope of most changes to a single module.
- paperplatter 2y agoThere would be a mix of versions, managed via branches. The part about debuggability sounded appealing at first, but if the multiple services you want to run are truly that hard to spin up locally, it won't be any easier as a monorepo. First thing you'll do is pass in 30 flags for the different databases to use. If these were RPCs, you could use some common prod or staging instance for things you don't want to bother running locally.
- nine_k 2y ago> There would be a mix of versions, managed via branches "We build the image slated for deployment from the release branch which is cut from master daily / weekly at noon." Works for MPOW, and some previous places. It's a monolith, there are rules! > but if the multiple services you want to run are truly that hard to spin up locally, it won't be any easier as a monorepo. First thing you'll do is pass in 30 flags for the different databases to use. Agreed! Maye that would be a reason to split the thing finally into separate services. Not necessarily micro-services, just into parts that are self-contained enough. But most code bases are not nearly as heavyweight. They can work pretty well as a monolith, run as a whole on a decent laptop, along with a database or two, or even three (typically Postgres, Redis, and Elastic). I know because I did it many times, and a ton of other people did. Worse yet, I ran the whole bunch of microservices, much like production, locally, again, as many a developer here did, too. At the scale this small, the complexity just slows you down. > use some common prod or staging instance for things you don't want to bother running locally. I've seen this in much bigger projects, and there it made complete sense. When you have to move literally a ton, it makes sense to use a forklift. But if the thing is a stack of papers that fits into a backpack, the forklift is an unnecessary bulk and expense. It could sill be a neatly organized stack of papers, not a shapeless wad.
- throwaway984393 2y agoI'm nearing greybeard status, so I have to chime in on the "get off my lawn" aspect. There is no one general "good engineering". Everything is different. Labels suck because even if you called one thing "microservices", or even "monolith of microservices", I can show you 10 different ways that can end up. So "modular monolith" is just as useless a descriptor; it's too vague. Outside of the HN echo chamber, good engineering practice has been happening for decades. Take open source for example. Many different projects exist with many different designs. The common thread is that if a project creates some valuable functionality, they tend to expose it both at the application layer and library layer. They know some external app will want to integrate with it, but also they know somebody might want to extend the core functionality. I personally haven't seen that method used at corporations. If there are libraries, they're almost always completely independent from an application. And because of that, they then become shared across many applications. And then they suddenly discover the thing open source has been dealing with for decades: dependency. If you aren't aware, there is an entire universe out there of people working solely on managing dependencies so that you, a developer or user, can "just" install software into your computer and have it magically work. It is fucking hard and complicated and necessary. If you've never done packaging for a distro or a language (and I mean 250+ hours of it), you won't understand how much work it is or how it will affect your own projects. So yes, there are modular moniliths, and unmodular monoliths, and microservices, and libraries, and a whole lot of varied designs and use cases. Don't just learn about these by reading trendy blog posts on HN. Go find some open source code and examine it. Package some annoying ass complex software. Patch a bug and release an update. These are practical lessons you can take with you when you design for a corporation.
- ryze20245 2y agoI read this yesterday and then came back today to upvote and comment because I thought it was so beautifully said
- rramadass 2y agoWell said! I am already in the latter half of my fifties and find articles/discussions like these irksome and a sad reflection on the state of knowledge of the Programmers today. Everything is merely cookie cutter recipes, patterns, cute jargons/acronyms, a general lack of understanding of computation models/paradigms, an inability to disambiguate actual concepts from language constructs, a lack of knowledge of fundamentals/important nuances all of which leads to simplistic cargo-culting. There seems to be no emphasis on thinking through the problem and a solution but only an eagerness to reach for the latest faddish framework/library/pattern to put together something and "make it work". Reminds me of Tesla's observation on Edison; “His [Thomas Edison] method was inefficient in the extreme, for an immense ground had to be covered to get anything at all unless blind chance intervened and, at first, I was almost a sorry witness of his doings, knowing that just a little theory and calculation would have saved him 90 per cent of the labor. But he had a veritable contempt for book learning and mathematical knowledge, trusting himself entirely to his inventor's instinct and practical American sense. In view of this, the truly prodigious amount of his actual accomplishments is little short of a miracle.”
- ljm 2y agoI can't help but feel like the author has taken some fairly specific experiences with microservice architecture and drawn a set of conclusions that still results in microservices, but in a monorepo. There's nothing about microservices that suggests you have to go to the trouble of setting up K8s, service meshes, individual databases per service, RPC frameworks, and so on. It's all cargo culting and all this...infra... simply lines the pockets of your cloud provider of choice. The end result in the context of a monolith reads more like domain driven design with a service-oriented approach and for most people working in a monolithic service, the amount of abstraction you have to layer in to make that make sense is liable to cause more trouble than it's worth. For a small, pizza-sized team it's probably going to be overkill where more time is spent managing the abstraction instead of shipping functionality that is easy to remove. If you're going to pull in something like Bazel or even an epic Makefile, and the end result is that you are publishing multiple build artifacts as part of your deploy, it's not really a monolith any more, it's just a monorepo. Nothing wrong with that either; certainly a lot easier to work with compared to bouncing around multiple separate repos. Fundamentally I think that you're just choosing if you want a wide codebase or a deep one. If somehow you end up with both at the same time then you end up with experiences similar to OP.
- paperplatter 2y agoI think the assumption here is that "microservices" means each team is dealing with lots of services. Sometimes it's like that. But if you go by the "one service <=> one database" rule of thumb, there will probably be 1-3 services per team. And when you want to use other teams' stuff, you'll be thankful it's across an RPC. First basic reason is if you don't agree with that other team on what language to write in. It'd really help to see a concrete example of a modular monolith compared to the microservice equivalent.
- mushufasa 2y agoWould Django's concept of an 'app' fit your definition of modular monoliths? https://docs.djangoproject.com/en/5.1/ref/applications/ https://docs.djangoproject.com/en/5.1/ref/applications/ In a nutshell, each django project is an 'app' and you can 'install' multiple apps together. They can come with their own database tables + migrations. But all live under the same gunicorn and on the same infra, within the same codebase. Many Django plugins are setup as an 'app'.
- halfcat 2y agoDjango apps can be the modular part of a modular monolith, but it requires some discipline. Django apps do not have strong boundaries within a Django project. Often there will be foreign keys crossing app boundaries which makes one wonder why there are multiple apps at all. In fact some people opt for putting everything into a single app [0]. Others opt for no app [1]. Django apps are good for installing Django packages into Django projects. But there’s no firm mechanism that enforced any real separation. It’s just other Python modules in a different folder (that you can just import into your other app). The rule would be something like, if you can’t pip install your Django app into a project, it’s probably too weak of a boundary (that might be a bit too extreme, but if it is, it’s not too far off). [0] https://careers.doordash.com/blog/tips-for-building-high-quality-django-apps-at-scale/ https://careers.doordash.com/blog/tips-for-building-high-qua... [1] https://noumenal.es/notes/django/single-folder-layout/ https://noumenal.es/notes/django/single-folder-layout/
- alganet 2y agoA good analogy is lacking though. "Modular Monolith" sounds like a contradiction. It doesn't help the idea. It inherits culture from OOP stuff, that abstraction was leaked to repositories, then it was leaked to packages, and it's being roughly patched together into meaningless buzzwords. It's no surprise no one understands all of this. I see the react folks trying to come up with a chemical analogy (atoms, molecules and so on), and the functional guys borrowed from a pretty solid mathematical frame of mind. What is the OOP point of view missing here? Maybe it was a doomed analogy from the beginning. Let's not go into biology though, that can't do any good. Spare parts, connectors, moving parts versus passive mechanisms, subsystems. Hard separation and soft separation. It's all about that when doing component stuff. And it has been figured all out, we just keep messing how we frame it for no reason.
- jerf 2y ago"What is the OOP point of view missing here?" The problem isn't that it is missing something but that it has extraneous parts. Inheritance as your default composition method couples together subtyping and interfaces. You basically can't build a modular monolith out of a large class hierarchy because the very act of being a large class hierarchy means you have more coupling between your classes than a modular monolith permits and already have a just-plain-monolith. From there standard code entropy will only grind it in harder. You really need to eschew inheritance for a good modular monolith and base it around interfaces only. Every major bit of functionality in your system needs a cleanly specified set of services it depends on, where that specification is just an interface and not hard-coded to some specific type, and certainly not some sort of class where subclasses must not only fulfill the interface but conform to the Liskov Substitution Principle, which is much, much harsher than most people realize. Then you have a true module, where if one day you need to yank out a particular service and make it a true microservice, you "just" take the interfaces it needs and either carry along the services locally or run them over the network. I scare-quote "just" because it's not infrequent to need to tweak the interface to permit it to work over the network (more things that can fail), but it's still a relatively mechanical and feasible process versus taking a whole bunch of super-hard-coded specific types, accessed randomly through whatever language scoping mechanism made sense at the time, and trying to convert that to run over the network. This is just a summary, of course, because it's an HN post and I can hardly lay out a complete design philosophy to the n'th degree in a comment. But there definitely is a true distinction between a monolith with all sorts of hard wiring internally and a way you can design a modular monolith that is still a monolith, but the coupling between the modules has been slimmed down to a minimum and controlled through some consistent gate rather than just willy-nilly wired together. Or, to put it another way, it's the difference between the module directly loading another module to directly resolve DNS, directly using the filesystem, directly wiring in a specific type for user authentication, and using a global instance of a logging system instance as initialized by some other module, you know, basically how most people write code all the time, and a module that accepts "a thing to resolve DNS addresses", "a filesystem interface", "a thing that authenticates users", and "a logger" where each of those things are interfaces, minimized down to just what the local module actually needs. One is a coupled nightmare no one want to touch. The other is easily swapped out to work with S3 to store files instead, or inject a different authentication system that ends up using a network service rather than directly hitting the DB, or swap in a hard-coded DNS resolver for testing purposes to avoid dependencies on the real network, etc. It isn't even really anything that surprising. A lot of people will at least claim this is just "good programming". But you don't get the benefits if you don't actually do it.
- andy_ppp 2y agoElixir + Phoenix is so great at this with contexts and eventually umbrella apps. So easy to make things into apps that receive messages and services with a structure. I’m amazed it isn’t more popular really given it’s great at everything from runtime analysis to RPC/message passing to things like sockets/channels/presence and Live View.
- sethammons 2y agothe elixir shop I was at, folks just repl'd into prod to do work. Batshit insanity to me. Is that the elixir way? Are you able to easily lock down all writes and all side effects and be purely read only? If so, they never embraced that.
- andrewmutz 2y ago> repl'd into prod to do work Like for debugging production problems and fixing customer data? Or for normal development? If its the former that's a great use of technology, and if its the latter it sound insane.
- psd1 2y agoIt's an order of magnitude harder to debug when you don't have access to prod, but there's a reason to block that access. I think you need to put controls on that fairly early in your project's evolution. Any good strategies to reduce the pain? My previous employer never solved this. I always wanted to explore contextual logging - by which I mean, logging is terse by default, but in an error state, the stack is walked for contextual info to make the log entry richer; and also, ideally, previous debug log entries that are suppressed by default are instead written. I guess that implies buffering log entries and writing only a subset at the end of the happy path. To illustrate what I mean: happy path log: 10:21:04 Authenticated 10:21:05 Scumbulated Error condition log: 10:21:04 Authenticating id 49234 request DEADBEEF IDP response OK for 49234 request DEADBEEF https://idp.dundermifflin.com https://idp.dundermifflin.com Session cookie OK request DEADBEEF Authenticated 10:21:05 Scumbulating flange 7671529 user 49234 request DEADBEEF NullFlangeError flange 7671529 at scum.py:265 Frame vars a=42, password=redacted, flags=0x05 I'm reacting to hard-to-repro bugs at $employer where we chucked logging statements at a dartboard, deployed, waited, didn't capture the issue, repeat several times. At a cadence of 5-10 deploys a week, this is below what I consider acceptable velocity. We often took days to fix major bugs, we'd run degraded for weeks at a time.
- gwbas1c 2y ago> In practice microservices can be just as tough to wrangle as monoliths. What's worse: Premature scalability. I joined one project that failed because the developers spent so much time on scalability, without realizing that some basic optimization of their ORM would be enough for a single instance to scale to handle any predictable load. Now I'm wrangling a product that has premature scalability. It was designed with a lot of loosely coupled services and high degrees of flexibility, but it's impossible to understand and maintain with a small team. A lot of "cleanup" often results in merging modules or cutting out abstraction.
- notjoemama 2y agoI’ve started taking 5 layers out of a Rails app, back to MVC. It’s so much faster now I actually feel bad, and I’m not the one that built the app in the first place. The premise during its construction was that it would scale to millions of active users. It…is not doing that in the wild…
- princevegeta89 2y agoThe company I'm at is a well funded startup that doesn't receive a humongous traffic at all. Yet, the so-called engineers in the early days ended up deciding to split every little functionality into a microservice. Now we have 20+ microservices that are setup together in a fucked up way. Today, every engineer out of our 150+ engineering team struggles with implementing and getting trivial stuff over the finish line. Many tasks require making code changes in multiple codebases and there are way too many moving parts. The knowledge overhead required to setup and test shit locally is too high as well. And the documentation gets so obsolete so quickly and people spend an obscene amount of time reaching out to other teams and running in circles to get unblocked on things. Our productivity would literally 5x if we just had 3 or 4 services overall. Even 1 giant service with clear abstraction between teams would have worked well, actually. Yet, for the flashiness and to keep sounding cool, the folks at our company still keep living with the pain. As an IC, I just fucking do my work 5 hours each day and just keep reminding myself to ignore whatever horrors I keep seeing. Seems to keep going well.
- 2y ago
- stephen 2y agoI mean, of course they are a good idea, what we need is more examples of actually doing them in practice. :-) I.e. quoting from the post: - monolithic databases need to be broken up - Tables must be grouped by module and isolated from other modules - Tables must then be migrated to separate schemas - I am not aware of any tools that help detect such boundaries Exactly. For as much press as "modular monoliths" have gotten, breaking up a large codebase is cool/fine/whatever--breaking up a large domain model is imo the "killer app" of modular monoliths, and what we're missing (basically the Rails of modular monoliths).
- bluGill 2y agoThe thing microservices give is an enforced api boundry. OOP classes tried to do that with public/private but fail because something public for this module is private outside. I've written many classes thinking they were for my module only and then someone discovered and abused it elsewhere. Now their code is tightly coupled to mine in a place I didn't intend to be coupled. i don't know the answer to this it is just a problem I'm fighting.
- Twisol 2y agoDifferent languages handle this in different ways, but the most common seems to be adding access controls to the class itself, rather than just its members. For instance, Java lets you say "public class" for a class visible outside its package, and just "class" otherwise. And if you're using Java 11 modules (nobody is though :( ), you can choose which packages are exported to consumers of your module. In a similar vein, Rust has a `pub` access control that can be applied to modules, types, functions, and so on. A `pub` symbol is accessible outside the current crate; non-pub symbols are only accessible within one crate. Of course, lots of languages don't have anything like this. The biggest offender is probably C++, although once its own version of modules is widely supported, we'll be able to control access somewhat like Java modules and Rust crates, with "partitions" serving the role of a (flattened) internal package hierarchy. Right now, if you do shared libraries, you can tightly control what the linker exports as a global symbol, and therefore control what users of your shared library can depend on -- `-fvisibility=hidden` will be your best friend!
- steveklabnik 2y ago> A `pub` symbol is accessible outside the current crate; This is not universally true; it's more that pub makes it accessible to the enclosing scope. Wrapping in an extra "mod" so that this works in one file: mod foo { mod bar { pub fn baz() { } } pub fn foo() { bar::baz(); } } fn main() { // this is okay, because foo can call baz foo::foo(); // this is not okay, because bar is private, and so even though baz is marked pub, its parent module isn't foo::bar::baz(); }
- eichi 2y agoIt doesn't matter if all of the team welcome the idea toward better productivity and enhance architecture iteratively. Culture and talent matters.
- devit 2y agoWell, that's just the normal way to write software, no? Aside from some websites and small scripts, all software is written like that. You simply create a hierarchical directory structure where the directories correspond to modules and submodules and try to make sure that the code is well split and public interfaces are minimal.
- stavros 2y agoNo, the fashion so far has been to put a network under all that.
- bluGill 2y agoWell you try that but in general someone in a different module discovers this thing you have over here is useful and starts using it and before you know it you have everything tightly coupled to everything else. not getting the above is hard.
- devit 2y agoAll non-toy programming languages support encapsulation, usually implemented with "private" or "public"/"export" keywords (well-designed languages make private the default), which means that unless the "thing" was marked as public/exported, in which case it's designed to be reused and stable and thus it's OK to depend on it, that will trigger a compiler or runtime error (in well-designed languages, a compiler error). Obviously, in that case it's perfectly normal and acceptable to either export or make public the thing, if it is a good idea for it to be part of the module interface, or if that's not a good idea factor out the useful thing and make it a 3rd module that both the original and new modules depend one; this should come with some documentation about the interface if it's not obvious or fully specified by the types.
- bluGill 2y agoThey try but there is puplic to this submodule but not the full module.
- zem 2y agomodular monolith + monorepo, so you get the benefit of continuous integration and automated code maintenance across the codebase.
- jillesvangurp 2y agoModules are almost as old as compiler technology. A good module structure is a time proven way to deal with growing code bases. If you know your SOLID principles, they apply to most module systems at any granularity. It doesn't matter if they are C header files, functions, Java classes or packages, libraries, python modules, micro services, etc. I like to think of this in terms of cohesiveness and coupling rather than the SOLID principles. Much easier to reason about and it boils down to the same kind of outcomes. You don't want a lot of dependencies on other modules (tight coupling) and you don't want to have one module do too many things (lack of cohesiveness). And circular dependencies between modules are generally a bad idea (and sadly quite common in a lot of code bases). You can trivially break dependency cycles by introducing new modules. This is both good and bad. As soon as you have two modules, you will soon find reasons to have three, four, etc. This seems to be true with any kind of module technology. Modules lead to more modules. That's good when modules are cheap and easy. E.g. most compilers can deal with inlining and things like functions don't have a high cost. Small functions, classes, etc. are easy to test and easy to reason about. Being able to isolate modules from everything else is a nice property. If you stick to the SOLID principles, you get to have that. But lots of modules is a problem with micro services because they are kind of expensive as a module relative to alternatives. Having a lot of them isn't necessarily a great idea. You get overhead in the form of build scripts, separate deployments, network traffic, etc. That means increased cost, performance issues, increased complexity, long build times, etc. Add circular dependencies to the mix and you now get extra headaches resulting from that as well (which one do you deploy first?). Things like graphql (aka. doing database joins outside the database) are making this worse (coupling). And of course many companies confuse their org chart with their internal architecture and run into all sorts of issues when those no longer align. If you have 1 team per service, that's probably going to be an issue. It's called Conway's law. If you have more services than teams you are over engineering. If you struggle to have teams collaborate on a large code base, you definitely have modularization issues. Micro services aren't the solution.
- pictur 2y ago[dead]