19 ms·
Learn how to design large-scale systems
- nwsm 8y agoThis is a nice followup to the web architecture post yesterday
- tryonqc 8y agoHe/she means this one: https://news.ycombinator.com/item?id=17517155 https://news.ycombinator.com/item?id=17517155 Today's post is way more in-depth. Good follow-up indeed.
- peterwwillis 8y agoObligatory pedantic HN grammar comment: on the outside chance that the gp's gender is not binary, the word 'they' is a good stand-in gender neutral pronoun to 'he/she'. You also have at least 14 alternatives to choose from (https://en.wikipedia.org/wiki/Third-person_pronoun#Summary https://en.wikipedia.org/wiki/Third-person_pronoun#Summary) and two more if you're at a Renaissance faire (https://en.wikipedia.org/wiki/Third-person_pronoun#Historical_and_dialectal_gender-neutral_pronouns https://en.wikipedia.org/wiki/Third-person_pronoun#Historica...). For the grammar snobs, this convention has existed since the 16th century.
- spraak 8y agoI wouldn't say it's obligatory, especially since the poster already was aware of not assuming gender by using "he/she" (though I know some people identify as neither of those). I do prefer singular they; it's very natural and yes, it's been around in English for a long time.
- tryonqc 8y agoI thought I did a good thing :( The use of their / they refering a single person doesn't come naturally to me as english is my 2nd language and we're taught its plural. (it can indeed be used as "third person plural singular" according to oxford dict.) Since its the "least bad" (to my ears) of the gender-neutral pronouns on the wiki page I'll try to use the "they/their" instead.
- geggam 8y agoNo database access layer ?
- cirgue 8y agoWhat's the distinction between a database access layer and read/write apis? Is that a semantic distinction or do they accomplish different things?
- geggam 8y agofrom my understanding you get the ability to put the DAL into a "pause" mode where it queues all the api requests allowing you do to updates / upgrades to the database with no downtime. It also gives you a way of controlling what queries are used by the API servers preventing a developer from doing silly things and creating a production outage
- gnahckire 8y agoIt also makes changing your DB a lot easier since APIs using the DAL don't need to be updated since they're DB agnostic -- you "only" need to update the DAL API.
- pc86 8y agoHow often does one change the DB backing a live production application?
- Lunatic666 8y agoI also don't think it's a good idea. If you don't use the database specific functions out of fear you aren't able to switch anymore, you are probably wasting a lot of potential performance.
- blattimwind 8y agoI think "changing the DB" likely referred to schema changes, not swapping out the DBMS.
- agentultra 8y agoI'd add a section on using TLA+ as a design tool. Diagrams and rules of thumb are useful but they don't catch errors or help you discover the correct architecture. See the Amazon paper [0] on their use of TLA+ in designing (and trouble-shooting) services. [0] https://lamport.azurewebsites.net/tla/formal-methods-amazon.pdf https://lamport.azurewebsites.net/tla/formal-methods-amazon....
- sytse 8y agoWhy is an AWS paper on azure websites? :)
- corobo 8y agoRedundancy. Real answer though it's a Microsoftian's (that's not a word) website https://lamport.azurewebsites.net/ https://lamport.azurewebsites.net/
- mxschumacher 8y agoCould you please talk about your experiences with TLA+? The "tools of thinking" for designing and verifying systems really interest me.
- agentultra 8y agoI could fill a blog post about it but in my current project we're using TLA+ for two things: 1. Helping us design features whose requirements are vague. The more hand-waving required to explain a particular feature the more likely we are to use TLA+ to model our assumptions and verify our understanding. This has led us to ask some interesting questions of our design team to help us build a better feature. 2. Requirements that are really hard that we need to ensure are implemented correctly. We use TLA+ to ensure the properties and invariants are correct with respect to the requirements and validate our model. This is really helpful in the case of concurrency and consistency. For our application we're using event-sourced data and it's imperative that our event store is consistent in the face of concurrent writers, can be replayed in a deterministic and consistent order, and that our assumptions will hold between versions of the events.
- sillysaurus3 8y agoNote that HN, a top-1000 site in the US, runs on a single box via a single racket process. "The key to performance is elegance, not battalions of special cases."
- nlawalker 8y agoHeh, elegance like "There is a story on the front page getting lots of attention, please log out so we can serve you from cache."
- tambourine_man 8y agoAdmittedly, that’s very rare.
- nlawalker 8y agoI know, just a good natured poke :) Plus you could probably take that comment at face value - making use of web caching is definitely an important tool when building a large scale system.
- copperx 8y agoWhy is caching out of the window when logged in?
- dmoy 8y agovotes, things to hide, etc all change the page and prevent serving everything wholesale out of cache I'm not actually sure, just guessing...
- paulddraper 8y agoIt's a page-level cache, and your view on a page depends on username, hidden submissions, point counts, etc.
- toomuchtodo 8y ago
- madethemcry 8y agoOh interesting, I have never seen Anki (https://apps.ankiweb.net/ https://apps.ankiweb.net/) being used for large blocks of source code. Anki is an open source application (desktop + mobile) for spaced repetition learning (aka flashcards). It's a very popular tool among people who want to learn languages (and basically anything else you want to remember). There are many shared decks (https://ankiweb.net/shared/decks/ https://ankiweb.net/shared/decks/). Creating and formatting cards is also possible and pretty easy. If you are planning to learn a language or anything else give Anki a try. I used it for all of my language learning efforts. With this least my vocabulary is rocking solid.
- byebyetech 8y agoI don't think Anki supposed to be used that way. Each card should be recallable under 10 seconds. So it should be only few lines of content. More content you put in one Anki card, it will take you more time and eventually you will stop looking at the card. A failure scenario.
- dan7678 8y agoYeah, I used Anki a ton in college, and doing things like this was always futile and frustrating. Flashcards are fantastic for learning short bouts of things, but not large structures like many lines of code. Additionally I'd say even if you succeeded in memorizing it this way, it's not making you a better problem solver, which is what actually matters for that particular subject; you're just (temporarily) better at regurgitating some lines of code.
- Twisol 8y agoI agree in general, but it seems like there might be a particularly constrained situation where it makes sense. I can usually look at a medium-sized block of code and suss out its intent in a short amount of time, even if I don't know all of the details about its behavior or how it works. A deck of flashcards curated from examples like that might be useful for recognizing the higher-level patterns that drive that intuition. I wouldn't really know for sure; I took the long way 'round (time and experience) for gaining that skill. But it at least seems plausible. Of course, the linked article isn't this.
- d-- 8y agoI'm teaching an intro distributed systems class and would like to share this with my students. I was wondering about how general the linked interview prepwork is. Are the Anki cards and sample interview questions mostly from large companies (FB, Google, MS) or also applicable to interviewing at smaller places? At first look, seems like these are fairly general questions, which is great.
- bandwitch 8y agoIf you liked this page, you might also like the excellent book "Designing Data-Intensive Applications" that among others surveys many characteristics of large-scale systems and presents some. Note that it's not a book for preparing you on system design questions, but it can definitely help.
- zjaffee 8y agoJust wanted to give a +1 to Designing data intensive applications, it's really one of the best resources out there in terms of touching really most of the areas necessary for building out big data applications, where you can then know which areas you'd prefer to dive further into.
- crystalPalace 8y agoI've been reading this and it's great so far. Are there other similar books that describe modern enterprise architecture at scale?
- jordanab 8y agoI'm currently reading "Building Evolutionary Architectures", and I'm liking it so far.
- hrshtr 8y ago+1 to the book. It has been a great resource for me to understand lot of concepts on Distributed Systems
- mozumder 8y agoWhat kind of numbers are they talking about for it to be "large-scale"? One well designed fast app server can serve 1000 requests per second per processor core, and you might have 50 processor cores in a 2U rack, for 50,000 requests per second. For database access, you now have fast NVMe disks that can push 2 million IOPS to serve those 50,000 accesses. 50,000 requests per second is good enough for a million concurrent users, maybe 10-50 million users per day. If you have 50 million users per day, then you're already among the largest websites in the world. Do you really need this sort of architecture for your startup system? If anything, you'd probably need a more distributed system that reduces network latencies around the world, instead of a single scale-out system.
- matachuan 8y agoWhy not have a scale-up system?
- marcosdumay 8y agoBecause it costs money and slows development and ops down. Is there a good reason for getting it when you are not one of the ~200 companies in the world with enough scale to use it?
- NightlyDev 8y agoAnd 1K rps/core isn't something that's necessarily hard to achieve, if someone thinks otherwise. I'm seeing about twice that on higly dynamic PHP pages with ~10 read/writes from/to MariaDB(running on the same machine).
- mabbo 8y agoThis design, roughly, is being used very widely and is well-documented everywhere. But does anyone know of any lesser-known yet equally functional designs that work at the same scale? Are there cases this design does not work for?
- deleted 8y ago[deleted]
- bsenftner 8y agoYes. One can use a C++ library like Restbed and embed the web server directly into a compiled executable that uses SQLite as an embedded database. The "large-scale, multi-system architecture" in such common use today is completely unnecessary when faced with this setup. I have multiple Restbed integrated applications whose entire disk footprint is 7MB; they can run on a $99 Intel Compute Stick, perform industrial grade facial recognition with multiple HD video streams, and still overwhelm traditional web stacks with events and data when pertinent events the software needs to report start emitting over the wire. The "only catches" are the developer(s) need experience working in multi-threaded C++, and they need to understand the traditional web stack they are eliminating.
- s_ngularity 8y agoWhat about fault-tolerance though? That's definitely a single point of failure scenario.
- bsenftner 8y agoRun as many instances as your fault tollerance requirements needs. The expense of adding another physical box is trivial when that physically box is $99 to $250 total to own. They "pay for themselves" in their first month of use, versus any cloud configuration running any 'amp or node or Mean or simply "traditional" web stack.
- deleted 8y ago[deleted]
- squegles 8y agoThis is a great outline for studying before interviews. I recently studied off of this and can say it contributed to my success in SRE/Infra interviews. Highly recommended!
- yread 8y agoI hoped this would help me with this problem I have - I'm coding a web app with a smallish database (<1GB for the next few years, <1% writes). I need low latencies for accessing it. And I would like to have multiple servers over the world sharing the database.
- lalwanivikas 8y agoyou need to provide more details to get any useful advice. but just based on what you have described, any db would do the job. add a caching layer and you have your low latencies. again, what is the traffic and bandwidth load like? peak and average values? what kind of data are you planning to store? small values but huge volumes or the opposite? a lot will change based on your system requirements.
- yread 8y agoTo clarify: let's say I have servers in two locations A and B that are 200ms from each other. When I issue a write to the db in A I don't want to wait (multiples) of the 200ms before it returns. I don't really care whether the write appears to a reader at B in 5s or 50 minutes but of course the writes have to be at least causally consistent. I won't have millions (realistically not even thousands) of users and the database will be comparatively small. I've looked at NDB cluster but it feels quite complicated to setup and maintain
- e12e 8y agoCan you write to a single rdms C, and maybe cache reads at A and B?
- Nican 8y agoLook at MySQL asynchronous replication.
- mindcrime 8y agoConsider Couchbase. It uses a combination of asynchronous writes and automatic replication to do a pretty good job of giving low latency writes even at high volume, while also ensuring data integrity. And since reads are served from the cache if possible, you usually get really good read performance as well.
- samirm 8y agoDid you _have_ to use comic sans for the diagrams? -_-
- cjhanks 8y agoI see something comparable to these diagrams (it feels like) a half-dozen times a year. The architecture is in general 'fine'. But communication paths of subsystems is probably the easiest part of the problem. And in general, re-organizing the architecture of a system is usually possible - if and only if - the underlying data model is sane. The more important questions are; - What is the convention for addressing assets and entities? Is it consistent and useful for informing both security or data routing? - What is the security policy for any specific entity in your system? How can it be modified? How long does it take to propagate that change? How centralized is the authentication? - How can information created from failed events be properly garbage collected? - How can you independently audit consistency between all independent subsystems? - If a piece of "data" is found, how complex is it to find the origin of this data? - What is the policy/system for enforcing subsystems have a very narrow capability to mutate information? If you get these questions answered correctly (amongst others not on the tip of my tongue), you can grow your architecture from a monolith to anything you want.
- pvarangot 8y agoAlso this architectures assume there's no need to do the dreaded "network locking", which for some problems regarding dispatch and avoiding triggering expensive/non idempotent batch jobs on background needs to be done. If you want to rely on SQL to do all the locking for you this usually doesn't scale.
- ris 8y ago> If you want to rely on SQL to do all the locking for you this usually doesn't scale. I suspect you have a mis-adjusted notion of "usually". "Usually", as in, for the majority of systems designed and in-use in the world, a well tuned, reliable RDBMS will be able to do this absolutely fine. The scale of systems that the world needs vs the quantity of them is an extremely long tailed curve.
- deleted 8y ago[deleted]
- 8y ago
- Zeebrommer 8y agoCan we please come up with a more specific name for this type of expertise? A large-scale system can mean anything from a social security system to a rocket. I was a bit disappointed that it only concerns websites here (though I'm aware that I'm browsing HN).
- cc-d 8y agoThe label is fine. Nobody is confused as to what a "system administrator" is, even though technically the word "system" itself can have a much broader range of meaning.
- mcqueenjordan 8y agoI'm not saying the label is wrong, but I agree with the parent's sentiment for a more specific label. "How to design a large-scale CRUD system" seems more precise. Large scale systems come in many different shapes and forms; this is an instance of one of them. Its learnings are interdisciplinary and cross-functional, but this isn't the roadmap for other types of systems, especially asynchronous reactive systems.
- s-shellfish 8y agoI agree. From my inferences in reading the usage of the label, large scale means not only users interacting with defined components that operate in predefined, predictable, static ways, but also components that involve the automation of development. This can be anything from the development of APIs, testing frameworks, parsers, code generation - all the computer science stuff basically. Large scale usually means some aspect of the business is focused on catering to developers, because the systems have become that complex that they require some form of automating existing automation.
- pier25 8y agoAnyone knows what software is being used to draw the diagrams?
- poxrud 8y agoOmniGraffle
- stvnw 8y agoIs there something similar to designing scalable front-end systems and going into deep discussions about how certain companies resolve similar issues at scale? I'd be interested if there is a resource like that out there. Everything out there tailored to systems design and architecture are entrenched in backend components.
- robax 8y agoAs a junior dev who one day wants to be in a senior position, this is super helpful. I failed the system design portion of the triplebyte interview and this would have been invaluable to me. Thank you!
- e12e 8y agoInteresting how the write api doesn't appear to invalidate/update the memory cache in the first diagram. Still recommend people read Fielding's REST thesis - as it demonstrates a lot of possible architectures (eg fat client or what we today call SPAs) - not simply REST. Along with some trade-offs. (REST is mainly motivated by simplicity of a simple hypertext application coupled with easy multi-level caching). https://www.ics.uci.edu/~fielding/pubs/dissertation/top.htm https://www.ics.uci.edu/~fielding/pubs/dissertation/top.htm For a preview of SPAs before the prevalence of Javascript, see 3.5, in particular 3.5.3 "code on demand": https://www.ics.uci.edu/~fielding/pubs/dissertation/net_arch_styles.htm#sec_3_5 https://www.ics.uci.edu/~fielding/pubs/dissertation/net_arch... And keep in mind the text is from 2000. Early Ajax was introduced in IE in 1999, and late 2000 in Mozilla - but it took a while for Ajax to become standardized...
- ris 8y agoI'm quite tired of everyone wanting to build "large scale systems" and play at being Netflix. The truth of the matter is the vast vast majority of people will never need to do this with their project and instead will just end up making an expensive to maintain mess with way too many moving parts. At least as important as designing something that can scale up is designing something that can scale down. You don't know when the organization's going to need to deprioritize this project and be able to keep it running without burning a couple of million in resources every year. See: microservices. (as in, for the problem, not the solution)
- deleted 8y ago[deleted]
- amorphid 8y agoI never thought about scaling down as a skill until just now. I kind of assumed "scaling up" implied up && down, our maybe "scaling out" implied out && in. Interesting thought.
- dmarlow 8y agoIt is indeed interesting to consider things like connection draining and playing nicely with the LB. Even in scenarios where machines are just removed for non-scheduled reasons.
- s-shellfish 8y agoBeing able to whittle down and simplify is an excellent skill to learn as a developer. It's my favorite and the one I find most fun. It allows everyone to focus on their specific components without leaping ahead in assumptions about how each developer will use each piece in the future. Lots of those kinds of problems are more easily solved in a room together, planned out, and done together. At least, that's what I've learned from how NASA developed their most important, complex parts. It's very easy to get ahead of oneself. Complexity grows by factors that are incredibly difficult to manage. Being able to simplify down to a context of parts that are moving and parts that are stable is a serene state of coding. Everything flows much easier that way. There will likely always be bugs and issues, but minimizing them to the smallest number there can be is an ideal value to maintain in software development.
- tanilama 8y agoLarge-Scale in what sense? A web service runs many instances isn't really instantly indicating its complexity.
- bovermyer 8y agoYou know what hasn't been done? A blog post about how to make a service that fulfills the needs of most people most of the time. All of the online and print material about such things focus on how to achieve massive scale correctly. Don't get me wrong; this is valuable and, generally, sound advice. However, it also ignores the majority of use cases for software. I would love to see a blog post here from someone who has solved a very specific problem for a very small audience, and gotten a very enthusiastic response. That would be meaningful on a larger scale for me.
- 0xFACEFEED 8y agoThis is because the people solving real world problems aren't writing books/tutorials/guides. Real world system design is dirty. Mostly this is due to constraints (time, cost, etc). And no one starts with zero architecture and 10 million users. Guides like this serve no purpose other than to fatten vocabularies and promote the "brand" of people who aren't actually doing the work (speakers, educators, etc).
- econochoice 8y ago> Guides like this serve no purpose other than to fatten vocabularies and promote the "brand" of people who aren't actually doing the work (speakers, educators, etc). Yep. They're often ghostwritten, too.
- bovermyer 8y agoI didn't say I was looking for a guide. I'm looking for a story. Surfacing things like this elevates the entire practice, since it illuminates what that "dirty" work looks like.
- NightlyDev 8y agoI find it fun to thinker with high performance and high scalability designs, but I, as most others, have no need for it. Start out small, make efficient systems and have scalability in the back of your head when doing so. Don't do as so many others: "Oh, this lib seems popular, let's just use that! Heck, the cart sometimes takes 8 minutes to load, we need to add more nodes on AWS!" Yeah, stuff like that happens. At least in my book optimization usually beats scalability as the place to start for more performance.
- visviva 8y ago*Software systems
- ex_amazon_sde 8y agoMost of this stuff would not pass a design review at Amazon. Anything that requires a fleet of (relational) databases to ensure consistency will not work on a global scale.