11 ms·
Distributed Systems Reading List
- kohlerm 6y agogood list but for me the reference is now Martin Kleppmanns new lecture: https://www.youtube.com/playlist?list=PLeKd45zvjcDFUEv_ohr_HdUFe97RItdiB https://www.youtube.com/playlist?list=PLeKd45zvjcDFUEv_ohr_H...
- rockanj 6y agoHis book “designing data intensive applications” is one of my favorites. I didn’t know he also had video lectures. Thanks for the link.
- Fiahil 6y agoThe book is absolutely amazing, it should definitely be up there in the list!
- MehdiHK 6y agoThank you for the link!
- giu 6y agoThank you very much for the link! Dr. Kleppman also provides notes for that specific lecture [0]; they contain all the slides with the corresponding detailed text, which is really awesome! [0] https://www.cl.cam.ac.uk/teaching/2021/ConcDisSys/dist-sys-notes.pdf https://www.cl.cam.ac.uk/teaching/2021/ConcDisSys/dist-sys-n...
- ilolu 6y agoIs the first part of the course (Concurrency part) available online too ?
- nwsm 6y agoThanks for sharing! I love having my copy of DDIA and am glad to have this resource from Martin as well.
- maurys 6y agoI've enjoyed listening to the MIT Distributed Systems lectures too. They do paper reviews and I pick the papers/topics which seem the most interesting (rather than doing them in order). They are great if you've already gone over the basics using "Designing Data Intensive Applications" or "Distributed Systems for Fun and Profit". https://www.youtube.com/playlist?list=PLrw6a1wE39_tb2fErI4-WkMbsvGQk9_UB https://www.youtube.com/playlist?list=PLrw6a1wE39_tb2fErI4-W...
- Zaheer 6y agoHighScalability is also a great practical blog for this. I've learned a ton and implemented learnings based on some of their write-ups: http://highscalability.com/ http://highscalability.com/
- DJBunnies 6y agoFollowing.
- dumb1224 6y agoI don't know if it belongs to the same category; but I find distributed learning very interesting and related. It has been a trending topic in biomedical field. In an area you can't gather enough data ethically, distributed systems are key. Example: https://ascopubs.org/doi/10.1200/CCI.19.00047 https://ascopubs.org/doi/10.1200/CCI.19.00047
- deepGem 6y agoI'm surprised there is no mention of the Bitcoin white paper. One of the most practical consensus protocols out there. Sure there is a lot of emphasis on the currency itself but I found the consensus protocol explanation so simple and easy to understand, as opposed to the Paxos algorithm.
- ahelwer 6y agoMurat Demirbas has an interesting analysis that Bitcoin can be viewed as a Paxos variant with very, very expensive leader election: https://muratbuffalo.blogspot.com/2018/03/anatomical-similarities-and-differences.html https://muratbuffalo.blogspot.com/2018/03/anatomical-similar...
- denkmoon 6y agoIt's missing the bible. Andrew Tanenbaum's Distributed Systems: Principles and Paradigms.
- projectileboy 6y agoGreat list. Only thing I’d add for the other enterprise developers out there is to first default to not building a distributed system at all, but rather build a much smaller monolith. In 25 years I’ve worked for so many orgs that wanted to build The World’s Most Scalable System for what would maybe be a few hundred concurrent users. Not surprisingly, those projects tend to tank.
- golergka 6y agoThat's great advice, as long as different systems in this monolith are loosely coupled both in terms of code and data, and you keep in mind that you will need to move them to separate services later on.
- lightlazer 6y agoThat is a great advice. However if scalability must be introduced later on, it can be really hard as there are many features that have been added to the monolith, and refactoring it to become scalable can be a huge task. The conditions where the monolith must be converted to a scalable system should be defined as early as possible.
- herodoturtle 6y agoThis too is excellent advice. Start off with a monolith if it makes most sense (in terms of simplicity / MVP) but from the start keep future scalability in mind and plan for it. I like this.
- bovermyer 6y agoUnless, of course, you intend to cap users at a specific amount for another reason.
- dasil003 6y agoDesigning something to be “scalable” before it’s actually established is a recipe for premature optimization at best and a poorly baked SOA with service boundaries that make change incredibly difficult at worst. It’s important to keep mind that SOA is about scaling teams first, code second and not really about throughout per se. A share-nothing web tier plus a couple judiciously applied databases and background job queues can effectively scale a huge proportion of applications without the overhead of a full SOA.
- mav3rick 6y agoMost of the links in comments are useful. But after a point it's hard to find small projects to implement to practice these. I had to do 2PC for a class.
- hb4ch 6y agoGreat list, but half of the entry link is dead.
- shred45 6y agoI find "distributed systems" to be a huge source of imposter syndrome. Despite having worked almost exclusively with distributed applications for several years now, it is difficult to consider myself experienced. When I'm asked if I've worked with distributed systems, I don't think they are asking me if I've managed a Hadoop cluster. They are interested in building new applications using some of the primitives discussed in this post. All of these links are great, but the fact is that building and operating tools like this is hard. In addition to consensus primitives, your system may need very precise error handling, structured logging, distributed tracing, resource monitoring, schema evolution, etc. In the end, I probably pause for a second too long when answering that question, but I don't think its because of a lack of experience, quite the opposite!
- calcorbin 6y agoI'm realizing I've only worked in distributed systems as well, but I'd never feel comfortable telling potential employers I'm an expert. Being an expert in distributed systems seems almost too broad. At a high level couldn't it be expertise at integration, accessible logging, and configuration?
- shred45 6y agoIt is definitely a broad term, and I think that disciplined implementation of the things that you mentioned is the real key. It just isn't as exciting to talk about.
- ystad 6y agoYes. The area is fairly broad. In my opinion - I think the question to ask for is: Do you have the distributed systems mindset? Not Are you an expert on Distributed Systems
- 0xbadcafebee 6y agoDistributed systems research has been going on since the 70's and Unix Neckbeards have probably forgotten more about them than we have learned, so actually I think impostor syndrome is a bit warranted with them. The actual hard stuff is not even these papers, it's the implementations that are way more complex than some algorithm or architectural pattern. Anyone who says "X is better than Y" is fooling themselves because it's only the implementation context that matters. The only thing you can say for certain is that reducing the amount of components and complexity in the system often results in better outcomes.
- ahelwer 6y agoI guess I'm pretty opinionated about this, but it was odd the author talked about the necessity of changing the way you think without also including anything about TLA+. IMO the "way you think" about distributed systems - if you want to be effective - will basically end up looking exactly like you think when writing a TLA+ spec, and learning TLA+ is a fast-track method of thinking like a distributed systems engineer. This is much, much more important in day-to-day work on distributed systems than knowing how a bunch of distributed systems algorithms work.
- tyu2 6y agoIs it though? The hard part about distributed systems is performance in our crappy real world environment with unreliable poorly performing and faulty public internet, unreliable hardware, OSes, etc. Which is directly at odds with needing TLA+, because if you do need it, it means the complexity of the algorithms is so great, that you won't be able to keep them in your head and understand every aspect of their performance to make something work well. It's similar how people think they can just peek a random consensus algorithm they've heard is correct and easy and make a decently working distributed storage, which is silly of course, it's only good for educational purposes. EDIT: (Ah, I see you are a TLA+ promoter, that's why you made a comment like that)
- ahelwer 6y agoI think the hard part about distributed systems is the combinatorial explosion of possible system states, which is also common to any concurrent program. Really distributed systems is just concurrency on hard mode, where failures are basically guaranteed instead of being very rare. I wouldn't particularly say I'm a TLA+ promoter (it's a FOSS project), any more than anyone who has a great fascination with a language/framework/algorithm/viewpoint is a promoter. We're all promoters of the memes that live inside our heads!
- smiths1999 6y agoI am not a "TLA+ promoter" but think it is a very valuable tool for anyone building distributed systems. The value of TLA+ is that it forces you to carefully consider your algorithm, which is certainly important if the algorithm is complex but equally important if the algorithm is simple. Most people will struggle to correctly specify even a simple algorithm in TLA+ because they will miss a lot of things they had assumed without ever thinking about. Real world systems need to handle all the things you mention. TLA+ helps you consider all these issues with spelling them out individually. There is no point in building a complex system if you haven't taken the time to validate the correctness of the target system in the first place.
- rlewkov 6y agoSeveral links no found
- danesparza 6y agoI would add the 'Secret Lives of Data' presentation about 'Raft" to this list: http://thesecretlivesofdata.com/raft/ http://thesecretlivesofdata.com/raft/ It's a fantastic visual walk through of the Raft consensus protocol (used in many modern distributed systems like Consul).
- libraryofbabel 6y agoGood list, but is it still being actively updated? Not having Kleppmann’s seminal Designing Data-Intensive Applications (2017) on it would indicate no. Alex Petrov’s Database Internals: A Deep Dive Into How Distributed Data Systems Work (2019) is another essential recent reference that should be here. Not as broad as Kleppmann but dives a lot deeper into certain topics.
- markc 6y agoAgreed, I looked for Kleppmann and didn't find it. Red flag.
- linux2647 6y agoEDIT: I was looking at the wrong branch, but still it hasn't been updated since 2018: https://github.com/dancres/Pages/commits/gh-pages https://github.com/dancres/Pages/commits/gh-pages Looks like it was last updated in 2013: https://github.com/dancres/Pages/commits/master https://github.com/dancres/Pages/commits/master (there's a commit from 2018, but it doesn't touch the actual page)
- avinassh 6y agoWhat are some more books / resources recommendations on the same?
- abledon 6y agoHow long would it take to realistically get through all these books? for the average developer... i'm thinking 3-4 years? and this includes taking weekends outside of work to painfully work through every chapter (+ break weekends since working+studying this in evening would probably lead to burnout in 1 year)
- johnnujler 6y agoI would wager around a year at max. Most materials related to Distributed Systems are information heavy(requiring you to read and remember) as opposed to being math heavy(with the possible exception of graph theory, which usually none of the fundamental texts expect the readers to know). Even many popular papers like raft, paxos, chubby, BitTorrent, gfs etc do not require much mathematical knowledge beyond basic arithmetic. You can read them like a non-fiction and understand most of the things. Not saying that Distributed Systems is easy, doing research in Distributed Systems is very challenging, but I don’t think reading and understanding these materials should be that difficult for an average software engineer.
- ruang 6y agoIs linear algebra useful to learn for distributed systems? I'm taking graph theory next semester and wondering what else might be useful. I took a distributed systems course already and we used no math at all.
- johnnujler 6y agoDefinitely, may be not directly, but it lies at the heart of software engineering(or more appropriately computer science). It helps with thinking differently, for example I don’t know if ndergrad courses teach distributed search algorithms using eigen vectors and values, but it helps to understand invariants and transformations better.
- stuxnet79 6y agoIMO a better approach would be to read Designing Data Intensive Applications (mentioned a million times on HN) which is more like a high-level map of the field. The references in DDIA are also a goldmine of information. You don't have to read DDIA front to back. Just picking a topic (for instance "Distributed Transactions") is enough to get you started building an intuition about these issues.
- cyberfart 6y agoAnother really good reading about distributed systems is Kyle Kingsbury's (aphyr) distsys-class notes: https://github.com/aphyr/distsys-class https://github.com/aphyr/distsys-class