19 ms·
Good system design
- mgaunard 1y agoSeems biased towards websites, which are mostly easy CRUD.
- kenny239 1y agowise article. thanks op.
- com 1y agoThe advice about logging and metrics was good. I had been nodding away about state and push/pull, but this section grabbed my attention, since I’ve never seen it do clearly articulated before.
- dondraper36 1y agoThe logging part is spot on. It has happened so many times when I thought, "Oh, I wish I had logged this.", and then you face an issue or even an incident and introduce these logs anyways.
- bravesoul2 1y agoIt is a balance. Too many logs cost money and slow down log searches both for the search and the human seeing 100 things on the same trace.
- dondraper36 1y agoYeah, absolutely. But the author's idea of logging all major business logic decisions (that users might question later) sounds reasonable.
- bravesoul2 1y agoYes. I like the idea of assertions too. Log when an assertion fails. Then get notified to investigate.
- jillesvangurp 1y agoThe trick here is to log aggressively and then filter aggressively. Logs only get costly if you keep them endlessly. Receiving them isn't that expensive. And keeping them for a short while won't break the bank either. But having logs pile up by the tens of GB every day gets costly pretty quickly. Having aggressive filtering means you don't have that problem. And when you need the logs, temporarily changing the filters is a lot easier than adding a lot of ad hoc logging back into the system and deploying that. Same with metrics. Mostly they don't matter. But when they do, it's nice if it's there. Basically, logging is the easy and cheap part of observability, it's the ability to filter and search that makes it useful. A lot of systems get that wrong.
- bravesoul2 1y agoNice. I'm going to read up more about filtering.
- bravesoul2 1y agoYes. Everyone should spent the small amount of time getting some logging/metrics going. It's like tests, getting from 0-1 test is psychologically hard in a org but 1-1000 then becomes "how did I live without this". Grafana has a decent free tier or you can self host.
- bravesoul2 1y agoHe doesnt seem to mention Conway or team topology which is an important part of system design too.
- dondraper36 1y agoWell, as sad as it is, such advice is often applicable to new projects when you still have runway for your own decisions. For mostly political reasons, if you are onboarded to a team with a billion microservices and a lot of fanciness, it's unlikely that you will ever get approval or time to introduce simplicity. Or maybe I just got corrupted myself by the reality where I have to work now.
- bravesoul2 1y agoThere is definitely a wood for the trees issue at bigger companies. I doubt there is an architect who understands the full system to see how to simplify it. Hard to even know what "simpler" looks like.
- jillesvangurp 1y agoYou should adapt your team to the architecture, not the other way around. My former Ph.D. supervisor who moonlights as a consultant on this topic uses a nice acronym to capture this: BAPO. Business, Architecture, Process, and Organization. The idea is to end up with optimal business, an optimal architecture & design for that business, the minimum of manual processes that are necessitated by that architecture, and an organization that is efficiently executing those processes. So, you should design and engineer in that order. Most companies do this in reverse and then end up limiting their business with an architecture that matches whatever processes that their org chart necessitated years ago in a way that doesn't makes any logical sense whatsoever except in the historical context of the org chart. If you come in as a consultant to fix such a situation, it helps understanding that whatever you are going to find is probably wrong because of this reason. I've been in the situation where I come in to fix a technical issue and immediately see that the only reason the problem exists is the org chart is bullshit. That can be a bit awkward but lucrative if you deal with it correctly. It helps asking the right questions before you get started. Turning that around means you start from the business end (where's the money coming from?, what value can we create?, etc.), finding a solution that delivers that and then figure out processes and organizational needs. Many companies start out fairly optimal and then stuff around them changes and they forget to adapt to that. Having micro services because you have a certain team structure is a classic mistake here. You just codified your organizational inefficiency. Before you even delivered any business value. And now your organizational latency has network latency to match that. Probably for no good reason other than that team A can't be trusted to work with team B. And even if it's optimal now, is it going to stay optimal? If you are going to break stuff into (micro) services, do so for valid business/technical reasons. E.g. processing close to your data is cheaper, caching for efficiency means stuff is faster and cheaper, physically locating chunks of your system close to the customer means less latency, etc. But introducing network latency just because team A can't work with team B, is fundamentally stupid. Why do you even have those teams? What are those people doing? Why?
- magnio 1y agoI think it's a very good article. Even if you disagree with some of the individual points in it, the advice given are very concrete, pragmatic, and IMO tunable to the specifics of each project. On state, in my current project, it is not statefulness that causes trouble, but when you need to synchronize two stateful systems. Every time there's bidirectional information flow, it's gonna be a headache. The solution is of course to maintain a single source of truth, but with UI application this is sometimes quite tricky.
- klabb3 1y agoYes it’s total madness to synchronize and replicate complex state that logically belongs together. This is why microservices are such hot garbage. Well, how people tend to use them anyway. Monotonic state is also better than mutable state. If you must distribute state, think ownership. Who owns it? Eg Theres nothing necessarily wrong with having state owned by eg a mobile client that can be adjusted by the user. Then you can sync it to the backend if you want, but they are only a reader/listener, and should never try to control it directly.
- ZYbCRq22HbJ2y7 1y ago> You’re supposed to store timestamps instead, and treat the presence of a timestamp as true. I do this sometimes but not always - in my view there’s some value in keeping a database schema immediately-readable. Seems overly negative of broad advice on a good pattern? is_on => true on_at => 1023030 Sure, that makes sense. is_a_bear => true a_bear_at => 12312231231 Not so much, as most bears do not become bears at some point after not being a bear.
- setr 1y agoIf you take the statement at face value — essentially storing booleans in the db ever is a bad smell - then he’s correct. Although I’m not even sure it’s broadly a good principle, even in the on_at case; if you actually care about this kind of thing, you should be storing it properly in some kind of audit table. Switching bool to timestamp is more of a weird lazy hack that probably won’t be all that useful in practice because only a random subset of data is being tracked like that (Boolean data type definitely isn’t the deciding factor on whether it’s important enough to track update time on). The main reason it’s even suggested is probably just that it’s “free” — you can smuggle the timestamp into your bool without an extra column — and it probably saved some effort accidentally; but not because it’s a broadly complete solution to the set of problems it tries to solve for I’ve got the same suspicion with soft-deletes — I’m fairly positive it’s useless in practice, and is just a mentally lazy solution to avoid proper auditing. Like you definitely can’t just undelete it, and it doesn’t solve for update history, so all you’re really protecting against is accidental bulk delete caught immediately? Which is half the point of your backup
- maxbond 1y agoAudit tables are a big ask both in terms of programming effort to design and support them, and in terms of performance hit due to write amplification (all inserts and updates cause an additional write to an audit table). Whereas making a bool into a timestamp is free. Including timestamps on rows (including created_at and updated_at) are real bacon savers when you've deployed a bug and corrupted some rows and need to eg refund orders created in a certain window.
- bambax 1y ago> When querying the database, query the database. It’s almost always more efficient to get the database to do the work than to do it yourself. For instance, if you need data from multiple tables, JOIN them instead of making separate queries and stitching them together in-memory. Oh yes! Never do a join in the application code! But also: use views! (and stored procedures if you can). A view is an abstraction about the underlying data, it's functional by nature, unlikely to break for random reasons in the future, and if done well the underlying SQL code is surprisingly readable and easy to reason about.
- DanielHB 1y agoMicroservice achitecture promotes splitting data cross multiple databases making it impossible to do proper DB JOINs from application code. Then companies buy a solution to aggregate all the different databases in a single "data-lake" (or whatever buzzword is hot right now) so you can do OLAP queries. Without consistency guarantees of course. And I am not saying this is never the _right_ solution, but it should almost never be the _first_ solution
- tossandthrow 1y agoViews make good sense when you can check them in - and DB migrations are a poor way of doing it due to their immutable nature. Depending on the ecosystem the code base adopts a good orm might be a better choice to do joins.
- CafeRacer 1y agoI came here to say an exactly opposite things. There were a few instances where a relatively heavy join would not perform well, no matter what I tried. And it was faster to load/stitch data together with goroutines. So I just opted to doing it that way. Also SQL is easy, but figuring out what's up with indexes and planner is not.
- deleted 1y ago[deleted]
- quietbritishjim 1y agoI think it's ok to have this rule as a first approximation, but like all design rules you should understand it well enough to know when to break it. I worked on an application which joined across lots of tables, which made a few dozen records balloon to many thousands of result rows, with huge redundancy in the results. Think of something like a single conceptual result having details A, B, C from one table, X, Y from another table, and 1, 2, 3 from another table. Instead of having 8 result rows (or 9 if you include the top level one from the main table) you have 18 (AX1, AX2, AX3, AY1, ...). It gets exponentially worse with more tables. We moved to separate queries for the different tables. Importantly, we were able to filter them all on the same condition, so we were not making multiple queries to child tables when there were lots of top-level results. The result was much faster because the extra network overhead was overshadowed by the saving in query processing and quantity of data returned. And the application code was actually simpler, because it was a pain to pick out unique child results from the big JOIN. It was literally a win in every respect with no downsides. (Later, we just stuffed all the data into a single JSONB in a single table, which was even better. But even that is an example of breaking the old normalisation rule.)
- tetha 1y agoThe distinction of stateful and stateless is one of the main criteria how we're dividing responsibilities between platform-infra and development. I know it's a bit untrue, but you can't do that many things wrong with a stateless application running in a container. And often the answer is "kill it and deploy it again". As long as you don't shred your dataset with a bad migration or some bad database code, most bad things at this level can be fixed in a few minutes with a few redeployments. I'm fine having a larger amount of people with a varying degree of experience, time for this, care and diligence working here. With a persistence like a database or a file store, you need some degree of experience of what you have to do around the system so it doesn't become a business risk. Put plainly, a database could be a massive business risk even if it is working perfectly... because no one set backups up. That's why our storages are run by dedicated people who have been doing this for years and years. A bad database loss easily sinks ships.
- mrkeen 1y ago> but you can't do that many things wrong with a stateless application running in a container > As long as you don't shred your dataset with a bad migration or some bad database code, most bad things at this level can be fixed in a few minutes with a few redeployments. At some point between these statements you switched from stateless to stateful and I can't follow the rest of the argument.
- tetha 1y agoIf you mess up your application code in a stateless container, that's boring. Roll code back, and you're back where you want to be. This is stateless and easy. If you introduce a migration like "UPDATE billing SET prices = 0 ; WHERE something < 5", that's an entirely valid migration, but you mess up your state and then everyone is in a world of pain. This could, however, still be caught by various code review strategies, incremental rollouts and a large number of good development practices. This is still easy, you can catch it before it hits prod so you don't have to fix prod. And prod could still be fixed if your database layer manages backups, just with a day or two of downtime. If you don't have backups, you may have permanently lost information, which could kill the company.
- KronisLV 1y ago> Schema design should be flexible, because once you have thousands or millions of records, it can be an enormous pain to change the schema. However, if you make it too flexible (e.g. by sticking everything in a “value” JSON column, or using “keys” and “values” tables to track arbitrary data) you load a ton of complexity into the application code (and likely buy some very awkward performance constraints). Drawing the line here is a judgment call and depends on specifics, but in general I aim to have my tables be human-readable: you should be able to go through the database schema and get a rough idea of what the application is storing and why. I’m surprised that the drawbacks of EAV or just using JSON in your relational database don’t get called out more. I’d very much rather have like 20 tables with clear purpose than seeing that colleagues have once more created a “classifier” mechanism and are using polymorphic links (without actual foreign keys, columns like “section” and “entity_id”) and are treating it as a grab bag of stuff. One that you also need to read the application code a bunch to even hope to understand. Whenever I see that, I want to change careers. I get that EAV has its use cases, but in most other cases fuck EAV. It’s right up there with N+1 issues, complex dynamically generated SQL when views would suffice and also storing audit data in the same DB and it inevitably having functionality written against it, your audit data becoming a part of the business logic. Oh and also shared database instances and not having the ability to easily bootstrap your own, oh and also working with Oracle in general. And also putting things that’d be better off in the app inside of the DB and vice versa. There are so many ways to decrease your quality of life when it comes to storing and accessing data.
- dondraper36 1y agoThere's a great book SQL Antipatterns, by Bill Karwin where this specific antipattern is discussed and criticized. That said, sometimes when I realize there's no way for me to come up even with a rough schema (say, some settings object that is returned to the frontend), I use JSONB columns in Postgres. As a rule of thumb, however, if something can be normalized, it should be, since, after all, that's still a relational database despite all the JSON(B) conveniences and optimizations in Postgres.
- quibono 1y ago> storing audit data in the same DB and it inevitably having functionality written against it, your audit data becoming a part of the business logic What's the "proper" way to do this? Separate DB? Separate data store?
- dennisy 1y agoThis post has some good concepts, but I do not feel it helps you design good systems. It iterates options and primitives, but good design is when and how you apply them, which the post does not provide.
- dondraper36 1y agoBut isn't that type of advice the best we can have? Having read Designing Data Intensive Applications (DDIA) and some system design interview-focused books (like those from Alex Xu), I have noticed two types of resources: * Fundamental books/courses on distributed systems that will help you understand the internals of most distributed systems and algorithms (DDIA is here, even though it's not even the most theoretical treatment) * Hand-wavy cookbooks that tend to oversimplify things, and (I am intentionally exaggerating here) teach to reason like "I have assumed a billion users, let's use Cassandra" I liked the article for its focus on real systems and the sensible rules of thumb instead of another reformulation of the gossip protocol that very few engineers will ever need to apply in practice themselves.
- usernamed7 1y agoOne thing i would add, is that a well designed system is often one that is optimized for change. It is rare that a service remains static and unchanging; browsers and libraries are regularly updated, after all. Thus if/when a developer takes on a feature ticket to add or change XYZ, it should be easy to reason about and have predictable side-effects of how that change will impact the system, and ideally be easy to change as well.
- paintistjksgf 1y agoWhile I think one shouldn't paint themselves into a complete corner, but optimizing for changeability signals a lot of abstract complexity to me, which takes time, which in turn takes money. We can usually have many more cheaper dedicated services for doing a thing that accounts for more good than a single service that grows to become more and more omnipotent. It also means you're much likely to win contracts because you can price yourself competitively
- mrkeen 1y agoAnd you get this part somewhat for free if you're actually testing as you're going. When my service wants to store and retrieve as part of its behaviour, of course I'm going to back it with a hashmap first. Once I know it fulfills its business logic I'll start fiddling with hard-to-change stuff like DB schemas and migrations. And having finished and tested the logic, I'll have a much better idea of the actual access patterns so I can design good tables & indexes.
- setnone 1y agoSome systems are designed to last vs. designed to adapt
- bbkane 1y ago"optimized for change" really only works well if you can predict the incoming changes. Common tools used for this "optimization" often raise the complexity and lower the performance of the system. For example, a db with a single table with just a key and a value is very flexible and "optimized for change" but it offers lower performance (in most cases) and is harder to reason about. I also frequently see people (me too) prematurely make abstractions (interfaces, extra tables, etc) because they're "optimizing for change". Then that part never changes OR it changes in a way that their abstraction doesn't abstract over OR they figure out a better abstraction later on when the app has matured a bit. Then that part of the code is at best wasted space (usually it needs to be rewritten yet no one gets time to do that). Of course, it's also foolish to say "never abstract". I almost always find it worth it to abstract over I/O, just so I can easily add logging, dual writes, or mock it. And when a change is obviously coming down the line it makes sense to plan for it. But usually I'm served best by trying to keep most of my computation pure functions (easy to test), doing as little as possible in the I/O path (it should just persist or print stuff so I can mock it) and otherwise write obvious "deletable" code that does one thing so I can debug it and, only if necessary, replace with a better abstraction if I need to.
- pelagicAustral 1y agoI can definitely feel the "underwhelming" factor. I've been working for +10 years on government software and I really know what an underwhelming codebase looks like, first off, it has my fucking name on it.
- StevenWaterman 1y agoWhat do you call system design, when it's referring to the design of systems in general, and not just computer services? As in: - writing a constitution - designing API for good DX - improving corporate culture I intuitively want to call all of those system design, because they're all systems in the literal sense. But it seems like everyone else uses "system design" to mean distributed computer service design. Any ideas what word or phrase I could use to mean "applying systems thinking to systems that include humans"
- gethly 1y agoActually event-sourcing solves most of the pains - events, schema, push/pull, caching, distribution... whatever. The downside is that it is definitely not suitable for small projects and the overhead is substantial(especially during the development stage when you want to ship the product as soon as possible). On the other hand, once you get it going, it's an unstoppable beast.
- mexicocitinluez 1y agoThere are some tools that try to solve this (MartenDb, for instance), but I wish there was an easier to way to integrate a system where some parts us ES and some parts don't. Almost all the tools I've seen are either fully event-sourced or have nothing to do with event-sourcing. There aren't a ton of in-betweens.
- gethly 1y agoES is at the core of the system where it is being used, so there really is no in-between option here. I've built two production ES systems and I was toying with few ideas to make it more DX friendly by using json as core of any entity/object and using json patch for events but in the end it made no sense because ES must be absolutely strict on schema, which evolves over time, and data types. You must be able to process events that might be a decade old and for objects that no longer exist. There is no wiggle room. Hence the aforementioned overhead. I do not know MartenDb but in essence every db today uses ES as that is how transactions work, except the event log is discarded after the commit. But either way, ES on db level is meaningless, except maybe being able to use it as actual log for auditing purposes but you won't be able to process it by any means as schema changes over time and db schema has little to do with the application itself anyway.
- mattlondon 1y agoThere was an article here recently about how to write good design docs: the TL;DR for that was basically your design doc should make your design seem obvious. I think that is the same conclusion here - good design is simple, straightforward design with no real surprises. Wholly agree.
- sama004 1y agohttps://news.ycombinator.com/item?id=44779428 https://news.ycombinator.com/item?id=44779428
- nvarsj 1y ago> Paradoxically, good design is self-effacing: bad design is often more impressive than good. Rings very true. Engineers are rated based on the "complexity" of the work they do. This system seems to encourage over-engineered solutions to all problems. I don't think there is enough appreciation for KISS - which I first learned about as an undergrad 20 years ago.
- anal_reactor 1y agoThis is unfortunately true. People love complex solutions, and suggesting a simple one usually comes across as incompetent, while the reality is, simple solutions are easy to manage, which ensures the success of the project as a whole. Sure, there are problems that are inherently complex and require complex solutions. But most likely yours isn't one of them, most likely you have a basic web app.
- chrisweekly 1y agoOne of the smartest engineers I've encountered in my 27 year career advised me to strive to do "the simplest thing that could possibly work" - not just to get unblocked on something new, but as a guiding principle. It resonated (and goes beyond "KISS", for me), and IME is real wisdom.
- jdlshore 1y agoThat’s a slogan from Extreme Programming! Coined by Ron Jeffries, I think, along with YAGNI (You Aren’t Gonna Need It), as a way of reminding people not to overengineer for features that in the plan.
- SatvikBeri 1y agoEvery now and then, I try to go through our codebase and write up the parts that we rarely think about – these are usually the cases where we made good decisions early on.
- thisbeensaid 1y agoSince the author praises proper use of databases and talks about event bus, background jobs and caching, I highly recommend to check out https://dbos.dev https://dbos.dev if you have Python or TypeScript backends. DBOS nicely solves common challenges in simple and complex systems and can eliminate the need for running separate services such as Kafka, Redis or Celery. The best: DBOS can be used as a dependency and doesn't require deploying a separate service. Very recently discussed here a week ago: https://news.ycombinator.com/item?id=44840693 https://news.ycombinator.com/item?id=44840693
- bubblebeard 1y agoVery good article, right on point! I do wonder about why the author left out testing, documentation and qa tool design though. To my mind, writing a proper phpcs or whatever to ensure everyone on the team writes code in a consistent way is crucial. Without documentation we end up forgetting why we did certain things. And without tests refactors are a nightmare.
- dondraper36 1y agoEspecially given that generating documentation and tests (of course, with manual revision) is so much faster with, say, Claude Code.
- whodidntante 1y agoNever write an article about good system design. In all seriousness, this is an extraordinary subtle and complex area, and there are few rules. For example, "if you need data from multiple tables, JOIN them instead of making separate queries and stitching them together in-memory" may be useful in certain circumstances. For highly scalable consumer systems, the rule of "avoid joins as much as possible" can work a lot better. There is also no mention of how important it is to understand the business - usage patterns, the customers, the data, the scale of data, the scale of usage, security, uptime and reliability requirements, reporting requirements, etc.
- graphviz 1y agoAnd by "system" we mainly meant "transactional website."
- whodidntante 1y agoAnd this is the point. You need to narrow the scope to make something like this useful. Writing a paper on "Good transportation design" is kind of meaningless. Do you mean cars, trucks, boats, planes, spacecraft, scooters, fighters, tanks ? Do you mean roadways that can accommodate some subset ? If you mean "transactional websites", and assuming you mean something like product catalogs and being able to purchase, that narrows it down quite a lot. Or does it ? For the majority of use cases, Craigs list, ebay, Amazon are the best fit. Next in number of use cases are Wix/Square/etc where you design your UI. Then comes all in one systems with UI/ORM based on Python/Ruby/etc where you need to design your own DB schema and UI, but the "design" is already done for you. The next step is custom designed systems like the one the article talks about, where complete off the shelf is not suitable And then there are the highly scalable systems The article is perfectly fine if we are discussing custom designed but not necessarily the highest in scalability.
- motorest 1y agoWhat a great article. It's always a treat to read this sort of take. I have some remarks though. Taken from the article: > Avoid having five different services all write to the same table. Instead, have four of them send API requests (or emit events) to the first service, and keep the writing logic in that one service. This is not so cut-and-dry. The trade offs are far from obvious or acceptable. If the five services access the database then you are designing a distributed system where the interface being consumed is the database, which you do not need to design or implement, and already supports authorization and access controls out of the box, and you have out-of-the-box support for transactions and custom queries. On the other hand, if you design one service as a high-level interface over a database then you need to implement and manage your own custom interface with your own custom access controls and constrains, and you need to design and implement yourself how to handle transactions and compensation strategies. And what exactly do you buy yourself? More failure modes and a higher micro services tax? Additionally, having five services accessing the same database is a code smell. Odds are that database fused together two or three separate databases. This happens a lot, as most services grow by accretion and adding one more table to a database gets far less resistance than proposing creating an entire new persistence service. And is it possible that those five separate services are actually just one or two services?
- bubblebeard 1y agoI think the author meant, in a general way, it’s better to avoid simultaneous writes from different services, because this is an easy way to introduce race conditions.
- paffdragon 1y ago> the interface being consumed is the database, which you do not need to design or implement You absolutely should design and implement it, exactly because it is now your interface. In fact, it will add more constraints to your design, because now you have different consumers and potentially writers all competing for the same resource with potentially different access patterns. Plus the maintenance overhead that migrations of such shared tables come with. And eventually you might have data in this table that are only needed for some of the services, so you now need to implement views and access controls at the DB level. Ideally, if you have a chance to implement it, an API is cleaner and more flexible. The problem in most cases is simply business pushing for faster features which often leads to quick hacks including just giving direct access to some DB table from another service, because the alternative would take more time, and we don't have time, we want features, now. But I agree with your thoughts in the last paragraph. It happens very often that people don't want to undertake the effort of a whole new design or redesign to match the evolving requirements and just patch it by adding a new table to an existing DB, then another,...
- nasretdinov 1y agoI agree with most of the stuff written in the article (quite a rare thing I must admit :)). But one thing I'd say is a bit outdated: in general whether or not to read from replica is the same decision as whether or not to use caching: it's a (pretty significant) tradeoff. Previously you didn't have much of a choice due to hardware being quite limited. Now, however, you can have literally hundreds of CPU cores, so all those CPUs can very much be busy at work doing reads. Writes obviously do have an overhead, _but_ note that all writes are eventually serialised, _and_ replica needs to handle them as well anyway
- codr7 1y agoReplacing booleans with timestamps might be a good idea sometimes, presenting it as The Solution isn't very constructive imo. Adding a separate table where the presence of a record means 'true' allows recording related state without complicating the main table. And sometimes a boolean is exactly what you want.
- wavemode 1y agothe article presents that as an example of bad advice
- msiyer 1y ago> Avoid having five different services all write to the same table. Instead, have four of them send API requests (or emit events) to the first service, and keep the writing logic in that one service. The ideal solution: Avoid having five different services all write to the same table. If five different services have to write to the same table, there is a major overlap of logic too. Are the five services really different or one would suffice? Taking practical realities into consideration, we can do what the author says. However, we risk implementing a lot of orchestration logic. We introduce a whole new layer of problems. Is that time not better spent refactoring the services: either give them their own DB tables or merge them into one servic?
- alixanderwang 1y ago> I’m often alone on this. Engineers look at complex systems with many interesting parts and think “wow, a lot of system design is happening here!” In fact, a complex system usually reflects an absence of good design. For any job-hunters, it's important you forget this during interviews. In the past I've made the mistake of trying to convey this in system design interviews. Some hypothetical startup app > Interviewer: "Well what about backpressure?" >"That's not really worth considering for this amount of QPS" > Interviewer: "Why wouldn't you use a queue here instead of a cron job?" > "I don't think it's necessary for what this app is, but here's the tradeoffs." > Interviewer: "How would you choose between sql and nosql db?" > "Doesn't matter much. Whatever the team has most expertise in" These are not the answers they're looking for. You want to fill the whiteboard with boxes and arrows until it looks like you've got Kubernetes managing your Kubernetes.
- ajuc 1y agoAnswer what they want and finish with "but in practice it doesn't matter for this much traffic and would be wasted effort". People ask for fizzbuzz in parallel not because it's practical.
- kenny239 1y agolol that's sad and real. modern software engineering has a lot of bloatware, costing security, etc.
- dondraper36 1y agoYes, and this is exactly why LinkedIn-driven development exists in the first place. Listing a million technologies looks much more impressive on paper to recruiters than describing how you managed to only use a modular monolith and a single Postgres instance to make everything work.
- ramraj07 1y agoDo you _want_ to work in these places? In my experience, if they expect you to run kube using kube in the interview, thats exactly what they do in their ststems as well.
- hks0 1y agoThe article starts by criticizing generic rules that come without any context: > Even good system design advice can be kind of bad. I love Designing Data-Intensive Applications, but I don’t think it’s particularly useful for most system design problems engineers will run into. But continues to do the same throughout the rest of its advices. It also says: > ... Drawing the line here is a judgment call and depends on specifics, And immediately mentions: > but in general I aim to have my tables be human-readable ... Which to me reads as "I'm going to ignore the difference of the context everywhere and instead apply mine for everyone, and I'm going to assume most of the wolrd face the same problems as me". It's even worse than the book being criticized in the beginning, as the book at least has "Data-Intensive" in its title. This is quiet easily fixable. The author can describe the typical scenario they are working with on a day-to-day basis. Do they work with 10 users a day? 100? 10,000,000? What is the traffic? How many engineers? What's the situation of the team/company; do FIXMEs turn into fixes or they become it's a feature? And so on. In the end, without setting a baseline, a lot of engineers will start pointing fingers at each other dismissing the opposite ideas because it doesn't fit their situation. The reasoning might be true, but before that, it is "irrelevant", hence any opposition to or defending of it.
- projektfu 1y agoI think it's just anodyne prose. If he just says, make tables human-readable, the comments will call that out for the specific cases where that's not good advice or it didn't work out. For example, if you need a bunch of user-controlled metadata, the database can't necessarily accommodate it as values in a single table. But you must be careful. https://thedailywtf.com/articles/Soft_Coding https://thedailywtf.com/articles/Soft_Coding
- lutzh 1y agoThe only thing I know about “good system design” is that it doesn’t exist in the abstract. Asking whether an architecture is good or bad is the wrong question. The real question is: Is it fit for purpose? Does it help you achieve what you actually need to achieve? I could nitpick individual points in the article, but that misses the bigger issue: the premise is off. Don’t chase generic advice about good or bad design. First understand your requirements, then design a system that meets them.
- msiyer 1y ago... that is how you achieve a good design (for the time being).
- vishnugupta 1y agoI highly recommend Boring Technology[1]. It is an enjoyable read and most of the advices are actionable. [1] https://boringtechnology.club https://boringtechnology.club
- feyman_r 1y agoIf you want to learn more about good system design at an abstract level (not just online), cannot recommend Systemantics[1] by John Gall enough. I wish all engineers get an opportunity to read it. [1] https://en.m.wikipedia.org/wiki/Systemantics https://en.m.wikipedia.org/wiki/Systemantics
- dondraper36 1y agoI enjoyed reading this book (it's a short one), even though the prose is very, well, special :)
- lysecret 1y agoThe designing data intensive Applications we actually needed (nothing against the original but this one is definitely more practical)
- agentultra 1y agoGreat article. A lot of very standard practices. Or at least… should be. One thing that I often add is the people interacting with the system. They’re a part of it too. Most people don’t operate in an atomically consistent world; a lot of business processes are eventually consistent. But you do need to know where you have to have atomic operations! It depends on where the user expects it. Systems thinking is very useful. From how your software is deployed to how the people using it in their work. Always be thinking about these things.
- robsalasco 1y agoanyone can recommend me a good book about systems design?
- hungryhobbit 1y agoThis is nonsense masquerading as advice. "Add indexes ... but don't add too many" is a perfect example. It's 100% correct ... and also 100% something no one can actually change their actions based on ... which means it's also 100% worthless advice.
- sfn42 1y agoAn intelligent reader might read that advice, realize they don't really know what indexes are nor how to use them, then do some research and learn what they need so that they can use indexes to their advantage. It would be a extremely long article it it went into detail on everything.
- QuadrupleA 1y agoAlso be careful not to reach for system design when you only need software design: https://lukerissacher.com/blog/optimizing_your_web_app https://lukerissacher.com/blog/optimizing_your_web_app
- necessary 1y agoExcellent article. In this vein, are there any books, articles, or other media that we can learn more of these sorts of principles from?
- ramon156 1y ago> I’m often alone on this. Engineers look at complex systems with many interesting parts and think “wow, a lot of system design is happening here!” Whenever I read something like this I feel so confused. Who actually calls themselves an engineer when they have no idea what they're talking about. Ignorant confidence is such a useless personality trait.
- dondraper36 1y agoUnless it's encouraged by the modern technical interviewing culture, which it partly is.
- mrkeen 1y agoHackers can add weight to an aircraft as fast - if not faster - than formally-trained engineers. As long as we keep measuring LOCs or features added, there will always be jobs for them.
- ninetyninenine 1y agoA lot of backend engineers are obsessed with infrastructure. I've seen engineers have servers spin up lambdas to do async jobs that are just database calls. So the server essentially waits for lambda which waits for a database. Why? Why can't you just have the server wait for the database? It's like I'm going to pay a person to wait in line for me while I wait for him. Why? You're waiting anyway!? And you just paid to involve an additional person to unnecessarily wait with you for what? When I told the engineer that you can just spin up a coroutine or like maybe you can allocate some cores before you spin up a new server... he looked at me like I was crazy. He said I was doing things so low level it was like assembly language programming. Going to low level and that lambdas were so cheap it was inconsequential. If you're reading this and you're thinking, wow that other engineer is right, well this quote from the article refers to you: "I’m often alone on this. Engineers look at complex systems with many interesting parts and think “wow, a lot of system design is happening here!” In fact, a complex system usually reflects an absence of good design."
- jagged-chisel 1y agoLots of comments here decrying unnecessary complexity and the depressing reality of job interviews around the subject. And I’m wondering why investors tolerate the expense: it’s surprising how much you can get done with simplicity and a small focused team. It must be something about perception when they’re ready to sell the company.
- firesteelrain 1y agoAs a system architect, software engineer and systems engineer, I see these posts and what is called system design seems to intermix systems design with software design (being that software described herein is a lower level component of the overall system)
- StevenWaterman 1y agoI'm glad I'm not the only one thinking that. These are such minutiae. Where's the discussion about humans? They're probably the most important part of your system, and the most chaotic, and the part that needs the most careful design. It's hinted at a little bit in the OP, with: > What does good system design look like? I’ve written before that it looks underwhelming This is because there are humans in your system! Other developers! You in the future! You have to resort to heuristics like "simple == good" because you're only looking at a small part of the whole system. And zoom out even more, you get to the actual users. How do they interact with the system? If you implement a rate limiter, how do the users respond when they hit it? Do they just spam-refresh the page? Open more tabs? Use their phone? Do they develop weird superstitions about it? Do they spam-call your phone support lines? Does your response to a thundering herd anticipate the second-order impact of your phone support lines being DDOSed?
- firesteelrain 1y agoExactly. Systems design at the SE level is socio-technical: the humans and their feedback loops matter as much as the tech. Many “good” designs fall apart because they forgot that the system boundary doesn’t stop at the codebase.
- gmm1990 1y agoI seem to gravitate towards nosql type databases, defining tables in a ddl and then again in the code seems repetitive, and slows down changes. But the idea would be that the code is what defines the table. It'd be nice though to hear some of the drawbacks of this. Maybe for very relational things it makes sense to be able to write join queries so data is completely repeated, but my understanding would be that most data base engines would already compress that repeated info pretty well.
- fogzen 1y agoI think there's a tension between databases as programs to store and retrieve data quickly, efficiently, and reliably vs. programs to enforce business rules and domain modeling. I am firmly in the former camp. In my opinion databases should be for storing and retrieving data as quickly and efficiently as possible. But the consensus in the database world seems to be that databases are primarily for enforcing business rules and domain models with foreign key constraints, triggers, views, transactions, type safety, domain modeling of relations, and on and on – some of which is at odds with storing and retrieving data efficiently.
- gmm1990 1y agoThat makes sense. Maybe it’s easier in an organization/ some people’s mental model to put guards around changing database because it’s separate from the code, standard across many organizations and in my opinion just harder to change.
- 0wis 1y agoIt is exactly what makes the difference between good and bad experience, both for users and engineers. A well designed system is both easy to use and to maintain or improve. It looks simple, but it is not. It’s both leadership and craftsmanship at its peak.
- jpitz 1y ago>You have two options: fail open and let the request through, or fail closed and block the request with a 429. If the metaphor of a software circuit breaker is meant to emulate an electrical circuit breaker, then it seems to me that these two are inverted. Whenever a physical circuit breaker is open, it is not dangerous and not passing current.
- sgarland 1y agoAgreed, and I don’t know why you’re being downvoted. If someone told me their virtual circuit breaker “fails open,” I would assume that it stops processing data upon failure.
- rekabis 1y ago> you’re a terrible engineer if you ever store booleans in a database Anything like this is trivially dismissible as absolute hogwash. It’s a shame that titles like this actually get the clicks needed to encourage more bullshite in the same vein.
- hn8726 1y agoWhat's the best resource to learn those things in practice? Other than _just trying_ yourself? I'm a senior dev in another area who wants to get into the backend development, but unsurprisingly has limited time to spend on learning a completely new thing,
- __turbobrew__ 1y agoThere are some books out there like “Distributed Systems” by Tanenbaum. Then there are the various company publications like “The SRE Book” by google and eng blogs. You could also contribute to an open source project like kubernetes or postgres to get your feet wet. Like all things, the best way to get better is to do it.
- maest 1y ago> But in most cases replication lag can be worked around with simple tricks: for instance, when you update a record but need to use it right after, you can fill in the updated details in-memory instead of immediately re-reading after a write. I get it, but that sounds very finicky code to get right and a good source of hard-to-debug bugs.
- sgarland 1y agoTwo things: 1. Essentially every RDBMS except MySQL has a RETURNING clause, so you can read your write essentially for free (and even MySQL lets you read the auto-increment value it inserted). 2. Barring that, how is this finicky to get right? If you get an ack back from the DB, assuming you haven’t done silly things to fsync settings, it’s written. It’s durable. So, within the same try/except or similar, take what you just told it to write, and use it. If you don’t get an ack back, it did not write, so don’t use what you had stored.
- 5Qn8mNbc2FNCiVV 1y agoYou'd have to reach the same instance for your next request. Easier to just tell clients that they get returned a header that they should attach to their requests so the backend can route their reads to the primary for some time until replication caught up.
- tra3 1y agoLots of good advice that I could’ve used 20 years ago. And like all good advice I would’ve ignored it back then.
- rekabis 1y ago> But in most cases replication lag can be worked around with simple tricks: for instance, when you update a record but need to use it right after, you can fill in the updated details in-memory instead of immediately re-reading after a write. I found myself truly confused by this one - does this actually need stating? Do people actually re-read immediately after a write? Provided you got confirmation that a write was successful and the data doesn’t have anything that an SQL trigger would change, what would be the point of an immediate read instead of just using the DB “successfully written” response as a go-ahead to just update the in-memory data?
- sethammons 1y agoThey do all the damn time. We have write pressure and teams insist they must read their writes. No, no you don't. There are options and they are not rocket surgery.
- deleted 1y ago[deleted]
- porridgeraisin 1y agoHere's why it happens: A lot of the time the datastructure you pass into writeAPI(obj) is different from the datastructure that is returned from readAPI(obj) -- even if the information contained is the same! No one wants to do that data structure transformation and potentially miss an edge case/break some implicit assumption about the data structure some fuckall downstream consumer has. However it is already done in the readAPI() function. So latency and throughput be damned, let us do: writeAPI(objects) objects = readAPI(objects) To be clear, I'm talking about the typical bloated data structure we all know and love: 20+ fields, different fields redundant for different services, sometimes they are empty in which case we have to fall back to calculate that field differently. And it is this hacky due to a quick bug fix during a sev1 last year that was never revisited to be fixed "correctly". There is a ticket hanging around somewhere to do this, but the assignee has left the company.
- sfn42 1y agoAt the very least you need to read the DB-generated ID, otherwise how do you provide stuff like editing? Simply reading after a write is a perfectly adequate solution for I'd say most systems. If you have performance issues and a change like this may solve them, sure. Switch to app-generated IDs and add workarounds for whatever other issues arise, and skip the read. But if you don't need to I don't see why you'd go through this trouble.
- voidhorse 1y agoGood system design is boring, obvious, and completely uninteresting. This is why a lot of flashy or trendy techniques end up leading to bad systems—they rope people in because of their cleverness or intellectual content, which often is interesting, but stuff that's new/intriguing/intellectually stimulating is often not what you want in a system. A good system needs to be as easy to understand and interpret as possible, A good system design is so mind-numbing my simple that a nincompoop can understand it. The only deviations from this policy should stem from other requirements like storage, performance, etc.
- UncleFullstack 1y agoIt can be kind of horrifying at times. A couple of years ago I was interviewing for work and ended up talking to a big liquor distributor - their challenge was dealing with a bunch of text files over FTP, and it was comically bad, they had a $300k annual spend on AWS, Kubernetes, the works. And they could have done the whole thing on a single EC2 instance with a couple of shell scripts. Needless to say I was laughed out of the room.
- axpy906 1y agoIt’s under appreciated how many of us fall for over engineering. I’ve been there. Just the other day I had coworker suggest that our startup not use cloud blob storage because it’s unreliable and we should build our own. Maybe it’s just harder to design with something simple that is possible to extend and build on. Maybe I am missing something. That said I agree with the author from my decade of experience.
- YZF 1y agoIt seems that there is no appreciation of a good architecture/system design any more. Nobody cares. Leadership can't tell the difference. If anything the worse designs seem more impressive. Engineers often enjoy a bad design because it creates more work and job security. When there's a lot of work it seems like you're getting things done. Managers enjoy a bad design because it helps empire building, now that there is more work we need to hire more people and do more manager-y things. There are also a lot of inexperienced engineers in the work force who have never seen a well designed system. In an organization running these badly designed system it's a political suicide to argue the design is bad. If the business is successful even more so because everyone will assume that a successful business means well designed software. A successful business will directly reward a bad design.
- rawgabbit 1y agoI wonder why the author views CQRS negatively and then later gives this classic CQRS advice: >What this means in practice is having one service that knows about the state - i.e. it talks to a database - and other services that do stateless things. Avoid having five different services all write to the same table. Instead, have four of them send API requests (or emit events) to the first service, and keep the writing logic in that one service.
- alphazard 1y agoAn entire post about "good system design" that completely fixates on the solution domain, and doesn't talk about the problem domain at all. The hardest part of system design is the interface that the system presents to users. That determines how they will use it and what they can use it for. A software system trades problems for different problems. e.g. We will manage your TODO list, provide consistency, durability, security, better than you could do yourself. But in order to get these benefits you have to understand our model, we have TODOs, users, lists, permissions, etc. Decisions about the interface (what problems the system presents to the users) are the most consequential, and the most costly to get wrong. If you aren't spending most of your time arguing about the interface, then you are wasting your time arguing about things that are comparatively easier to change later. Literally everything else about the system can be changed without bothering the users.
- deleted 1y ago[deleted]
- branko_d 1y ago> Indexes work like nested dictionaries If the author meant “dictionary” in a sense of a hash map, that’s not quite correct. In relational databases, indexes are usually B-trees, which are ordered, unlike hash maps. A B-tree can help with range-searches, ORDER BY and even merge joins, not just equality-searches.
- belZaah 1y agoWhen systems people talk about system design, they talk about the whole value-generating system and not just software or hardware. This typically involves people and brings in issues like the Conway’s Law. A tightly coupled team with little independence in terms of workstreams will produce a monolith regardless of what the architect dreams of, for example. If you have two sets of users with diverging regulatory and organizational needs (people who maintain the registry of citizens vs people who issue identity documents), you will have two separate data stores regardless of how much sense it makes to have just one.
- spectraldrift 1y agoIn general, I agree with the author's presumption that simplicity is better in system design. The frequent and tired allergy for managed queues, however, does not follow. > Sometimes you want to roll your own queue system. I have never wanted to do this. > For instance, if you want to enqueue a job to run in a month, you probably shouldn’t put an item on the Redis queue. This specific requirement sounds more like a cron job use case, not a queue case. > In this case, I typically create a database table for the pending operation with columns for each param plus a scheduled_at column. I then use a daily job to check for these items with scheduled_at <= today, and either delete them or mark them as complete once the job has finished. At this point I've decided the author doesn't understand when or why to use a queue. For a strictly scheduled event which must occur on day = (today + N), the proposed approach seems fine. However if you use this for a typical queue use-case you will end up reimplementing a queue in your database (poorly). This is typically far more complex than just using a managed queue service if it's available to you. By more complex, I mean more lines of code and more oncall burden. Furthermore, if you grow, a managed queue is often something that needs very little hand holding. Queues are great because they are simple- both at low and high volume. They are a technology that "just works", and often provide a lot of helpful features out of the box. I don't know the author, but I've encountered this engineering philosophy before. It's a perspective often held by engineers whose ideas haven't been stress-tested by the long-term realities of a successful business. It's the kind of advice you can sell to the 99% of startups that fail, and because they fail for other reasons, you never have to be proven wrong.
- manoji 1y agoExcellent article . Few more that come to mind Think real carefully about breaking transactionality unless absolutely needed . It has been the single source of most problems i have seen over the past few years. Keeping 2 different systems in sync is really hard do not do it if you dont have a real need to do it. Monoliths are really good , there is absolutely no need to run microservices or any services for that matter other than a single monolith. The place where i work at is generating billions of dollars with a single monolith.Having said that there will come a time when some logic has to go to a different services , If you get there your company is really really successful :) . Relational databases can do a lot more than what you think and they can absolutely scale well.
- elcdodedocle 1y agoGood article describing a boilerplate framework & techniques for (web) backend systems architecture implemented with off-the-shelf components. But like every systems architect I have ever met, it does not even put a thought into security or data governance (a priori).
- avinassh 1y agoOP has this first: > If a system has distributed-consensus mechanisms, many different forms of event-driven communication, CQRS, and other clever tricks, I wonder if there’s some fundamental bad decision that’s being compensated for (or if the system is just straightforwardly over-designed). then later down in the article: > Send as many read queries as you can to database replicas. A typical database setup will have one write node and a bunch of read-replicas. The more you can avoid reading from the write node, the better - that write node is already busy enough doing all the writes. isn't this same as CQRS
- frankc 1y agoSomewhat, though CQRS might advocate for separate read and write models. It might be something like the writer publishes an event that is consumed by something that updates the read model and queries go directly to the read model. Whether or not that is overengineering depends on the problem and load.
- quails8mydog 1y agoI thought cqrs was a code pattern where you segregate query models from command models, rather than something that specifies where the data is read from or is written to.