11 ms·
Systems Design for Advanced Beginners
- dreamcompiler 6y agoNice tutorial! One small nitpick: > Calculating a hash value from an input is computationally very easy, but reversing the transformation and recovering the original input from its hash value takes so much time and computing power that it is, practically-speaking, impossible. The above is true for encryption but not for hash codes. Recovering the original input from a hash code is not just practically impossible; it's provably impossible -- even with an infinite amount of computing power -- because in general a hash code contains less information than the original text.
- lamida 6y agoI think the author means recovering the hash by using dictionary attacks
- exdsq 6y agoYou’d still have hash collisions
- deleted 6y ago[deleted]
- hansvm 6y agoMaybe not enough to matter though. If the input space is sufficiently smaller than the hash space (e.g., you aren't hashing arbitrary strings, but instead just arbitrary English words) then the probability of any hash collision that also lies in the realm of interest is vanishingly small.
- exdsq 6y agoTrue, but that'd be very esoteric - a string input field that only accepts values from a dictionary. Although also if the input space was only these values but the input size was infinite you'd have a similar issue :P
- hansvm 6y agoIt'd be esoteric for a system to enforce only containing values from a dictionary, but every sentence ever spoken by anyone alive fits in a 64-bit space by a wide margin. For real-world input you could still reverse most passwords uniquely given enough time. It's actually kind of a fun problem -- the fewer bits you have in your hash the easier it is to find _any_ collision that gives access to the current system, but the harder it is to uniquely reverse the hash into a plausible password for stuffing into other systems.
- sz4kerto 6y agoActually, it might be occasionally possible. Not perfectly, of course. If the input space is limited and you have other constraints on it, then you might be able to find the original text. For example, let's say your input is a 280 character long tweet, the hash is only 16 bytes long. If you find that "I AM A STABLE GENIUS" is one of the inputs that hashes to the given hash value, then you can be reasonably sure that this was the original text, given that most of the other inputs with the same hash won't be meaningful English sentences.
- talaketu 6y agoI'd say the count of believable English tweets would be many orders of magnitude greater the count of 16 byte hashes, so in general we can't just find a believable tweet and declare that we found the plaintext. There would be many possible collision. The phrase you mentioned is notable (with a slight edit) - so I suppose you mean you could index phrases by notability, then it may be tractable find tweets of notable phrases, since the number of notable phrases is relatively low.
- KMag 6y agoI think the author is just using imprecise language for a preimage attack.[0] [0] https://en.wikipedia.org/wiki/Preimage_attack https://en.wikipedia.org/wiki/Preimage_attack
- oconnor663 6y agoInformation is sometimes destroyed, but we can't count on it. An algorithm like HMAC might be used to authenticate a short message. Or even if the message is long, it might also be known, so perhaps the only unknown input (the key) is short. HMAC still provides security in this case, but it's not provable-because-information-is-destroyed in the sense you're describing.
- netsectoday 6y ago215D70DA68DBDC2217D13916E40F89B87DDC3DAACA99892C943E0DCC76C8C1AF68097C86CC4CB53F15BD6C3870418A1E3577A076611A5104C6D4FC33939A21FB
- dreamcompiler 6y agoI said "in general".
- netsectoday 6y ago> in general a hash code contains less information That is true, but has nothing to do with your ability to use a rainbow table or dictionary attack on my hashed statement. The specific reason why you can reverse my hash is because it's unsalted. The time-complexity of my example hash is somewhat low. 'you are wrong' is 13 characters that spans the lowercase alphabet with spaces, in all english words that can be found in a dictionary. The actual key-space is 13^10 because of the limited characters I used, but when creating a rainbow-table it would probably take a key-space of 13^27 because you don't know the specific characters I'm using. (There are 13 characters in my statement, then there are 26 lower-case letters in the alphabet - plus space - which makes 27... 13^27) Now, if I followed general op-sec suggestions I would have added a salt to my statement, so the process looks like SHA512('you are wrong' + '7aomAxgeVjAvyDXGrdmNJNKuuiumYbkG') which gives you the possible key-space of 45^63. Also, the salt should be something not found in a dictionary. 2812CDA67BF0EA2D5FE125C2637466FA9A81C12AE9D101771C581DB87EBEB05D7724C9DCFBC6F905B99F9737948543EC64CAC5D89C785125DDA2E3297214CC58 On the upper-side of the estimate there are 10^82 atoms in the universe. There are about 14^103 possible solutions to that salted hash above, with many possible collisions. Good luck with creating a database that needs more rows than available atoms, or waiting for a brute-force to complete on that hash, then sorting through the collisions. This is the "impossible" reversible hash you are talking about. If you are interested in this topic and would like to do this in your code: use the HMAC function which is built for this, or specifically for passwords use bcrypt/blowfish because that algorithm is designed to run slowly (brute force protection).
- imvetri 6y agoTitle corrected - Web application system design for advanced beginners
- fouc 6y agoYeah, it's not true "Systems design" since it's specific to programming
- rustamm 6y agoRegarding DB backups, it would be useful to add that it is not enough to simply make backups, one should also test that the backups are actually restorable.
- intpx 6y agoreal developers never concern themselves with the realities of operations /s
- chasd00 6y agoalso the restore/recover process has to be laid out step by step and be as simple as possible. No one is doing a restore in a relaxed, easy going state of mind. Restores are done in full panic mode with the team thinking about how they're going to break the news to their families that they've lost their job.
- pc9 6y agoHow can I gain practical experience in these things? The jobs I've had mainly revolve around adding new features, not doing any of the things described in the article. Does working at bigger companies actually give you experience with this?
- qppo 6y agoWork at a startup as the systems person, because they don't already have one
- taigi100 6y agoYes and no, depends. It gives such experience to a software architect, otherwise not really. Tho, software architects do a lot more than just systems design (tho, that is pretty much their main task). But yes, you get a job as a software architect or something like an "in-training" one, helping one, etc. at a big corporation to get such experience (I'd argue that small companies only need some code design, you want a lot of systems/software design you should go for a big corporation)
- mav3rick 6y agoYou evolve into the architect role as an IC. No new grad is hired like this. You design things and keep going up in scale.
- taigi100 6y agoIt happens from time to time, but yes - there is no specific job / hiring for this. The default/common path is evolving into the architect role which seems like a "bad" process of developing an architect. The best way to develop a new architect, is to have him learn alongside a mentor/teacher who is an architect himself. Developing and architecture are very different jobs which require different mindsets and skills. Also, I've seen many places, teams, etc where people are mostly made to implement feature after feature with no time in between for learning, self-development, courses, etc. Otherwise, after you've just "arrived" to that architect role you start learning on your own what architecture really is and means.
- Hyperborian 6y ago> How do they store their data? SQL. > How do their different applications talk to each other? Proprietary APIs. > How do they scale their systems to work for millions of users? T H E C L O U D > How do they keep them secure? They just... don't. > How do they make sure nothing goes wrong? They just... don't. > What are APIs, webhooks and client libraries, when you really get down to it? Easily outsourced to India.
- mberning 6y agoThat last point is so painful and true. I work with a product that has dozens if not hundreds of “connectors” to interface with other systems and you can tell that most of them were farmed out to the lowest bidder.
- amdelamar 6y agoThis basically outlines my current job. Huge system with many, constantly moving/upgrading parts and services across teams, all while utilizing dozens of internal services, tools, and navigating corporate policies. I’m thriving in it, but totally recognize it’s a steep learning curve and takes longer to onboard newcomers. Eventually you get to enjoy deprecating old services as much as building new ones, simply because you never have to teach others about them again.
- dzink 6y agoAs a solo tech founder of several sites, I’ve had the pleasure of digging into each of these problems and more over the past few years. Some topics worth adding to the list for consumer sites: 1. Prototype and benchmark each of your stack pieces before you pick a stack. It is far easier to fix architecture mistakes when you don’t have 10000 users expecting overnight customer service. If you are using new technology, find good open source products to see how they’ve structured their projects. Your architecture, designed for speed and experience, will be your key differentiator. Inherent speed at core task, design to user needs, and name choice are the 3 musketeers of a solid growth. 2. Prepare for abusers to attack your system from every direction, especially if you enable users to publish content under your domain. You will see bots looking for Wordpress installations, users trying to fill content with SEO links, users trying every hacking vector known to market. Collect known vectors and test for them and never refuse a legitimate bug bounty request. 3. There is an eternal debate about the trade off between building a quality product at start or opening early to get user feedback before you get too deep into features. There is merit to both choices as a solo founder. The moment you open the gates to users, your ability to make changes comes with very high friction. With time, trying new features becomes a tremendous luxury hidden under bug requests, roadmap, customer service replies, etc. Build your biggest riskiest assumptions first. 4. Testing is hard as a solo founder. Selenium is your friend. If you don’t spend time with them, your users will take that time in multiples after a mistake. The best way to learn is to come up with a product you really want and build it in your own time. You can test launch in a weekend. When I started launching consumer sites solo as an engineer, I went from a tight specialization to being unafraid to try any tech if it gives me an advantage in solving a problem. Once you’ve simulated enough problems and dealt with the consequences of your choices personally, you can play that 10 level chess game with the architecture of each new feature much faster.
- cactus2093 6y agoThis is interesting advice because every company I've come across that has been successful enough to get past the early startup phase has by and large ignored all of this. The companies I've seen succeed were 100% focused on shipping their product to customers. Not 90% focused on customers and 10% focused on code quality, but 100% focused. They'd rather have to spend 30 engineer-days a few years from now fixing an issue if they get to that point than spend 3 hours getting it right upfront. As an engineer that goes against every instinct I have, it really seems like spending a couple hours upfront must be a better use of time. It seems like it should be possible to spend 10% of your time setting yourself up well for the future, that's still just a rounding error of your time. And then if you do survive another few years, you'll have a huge leg up on other series B or C stage competitors if you're not hindered by a lot of tech debt at that point. But from a capitalist perspective, it's probably not so crazy. If you are working with a $200,000 seed round in the beginning, and an engineer costs $80,000 a year, 3 hours of their time costs $115 which is 0.06% of your funds. And more importantly, that $200k is maybe enough for a year of runway, so 3 hours is 0.14% of the time you have to live given a 40-hour a week year (or 0.07% of an 80 hour a week year). Every bit of that starts to add up. Whereas by the time you're a later-stage company and you've raised, say, $40 million dollars and are paying engineers $150k, 30 engineer-days of work is $17,300 but that's only 0.04% of the money you've raised plus your runway is now approaching infinity if you're close to profitable. I'm still kind of playing devil's advocate here, my instinct really wants to believe that a better balance than what I've seen is possible. One huge missing factor is that people have a strong tendency to ignore spread out costs like the time wasted fixing bugs that pop up later whereas the upfront cost of writing a bunch of tests is more visible. But it has been interesting for me to consider that maybe most founders really are acting pretty reasonably, even though it originally seems pretty careless and lazy to let your startup build up a ton of tech debt early on.
- semicolonandson 6y agoThought I'd post this here since I imagine there's a strong cross-over of interest: If you'd like a guided video code-tour of how all the pieces of a production web-app fit together, I publish detailed weekly screencasts showcasing the code, systems, and architecture behind Oxbridge Notes, the business that's supported me for the past decade. So far I've covered: — software dependency vetting — data integrity systems (constraints, foreign keys, transactions etc.) — integration testing systems — trade-offs in software quality between customer-facing and admin areas — softer stuff, like designing for SEO (marketing ease is a critical part of any system I design, as an indie-hacker) https://www.semicolonandsons.com/series/Inside-The-Muse https://www.semicolonandsons.com/series/Inside-The-Muse
- bradleykingz 6y agoInteresting. Thanks.
- sciencewolf 6y agoHey Jack, Just a heads up - I tried to sign up for your mailing list but kept seeing "ERROR: Did you forget to type your email?" even after typing the email. I've tried a few different emails and also tried viewing the page in incognito mode. Hope you get this fixed soon, as I'd really love to get more of your content!
- igneo676 6y agoI get the same error :/ Also, it'd be great to be able to turn off the mailing list sign up on the videos. I'm already technically subscribed via RSS and it'd be great if the videos play through without intervention
- semicolonandson 6y agoThanks for the head's up! I checked the logs and it was just a case of an incorrect error message. Fixed now.
- iudqnolq 6y agoYour site appears blocked on Vodafone UK's mobile network. I get a connection interrupted error, the same error I get if I try to access a site I know is blocked (such as archive.org). If I use a VPN I can load the page, even if the VPN connection terminates in the UK.
- bibabaloo 6y ago> Steveslist currently has a very simple and slightly fragile cron setup. We have a single “scheduled jobs server”. We use crontab on this server ... This setup isn’t scaling very well ... We’re considering setting up a new system using a modern tool like Kubernetes. Do people use Kubernetes for running scheduled jobs like this? It seems like it'd be overkill but in saying that I'm not sure if I know of anything that can be used for running scheduled tasks in a reliable and observable way that's scalable. Maybe Jenkins?
- chucky_z 6y agothis is a pretty common use-case for ci/cd tools. azure devops, jenkins, concourse, drone, gitlab, ... also a lot of etl tooling has this stuff (airflow comes to mind immediately) there's also nomad which is easy to run and can schedule a lot of jobs pretty quickly, with decent observability. there's... a surprising amount of tooling in this exact space!
- harpratap 6y agoThe biggest Kubernetes cluster (arguably) is basically one huge batch job - https://cloud.google.com/blog/products/containers-kubernetes/google-kubernetes-engine-clusters-can-have-up-to-15000-nodes https://cloud.google.com/blog/products/containers-kubernetes...
- cyberdrunk 6y ago> Do people use Kubernetes for running scheduled jobs like this? We do, mostly because we already have Kubernetes and automation to deploy to it, so why not? The reliablity and monitoring are better than something we'd cook up on our own.
- sciurus 6y agoFor better or worse, sure. Here's one case study- https://stripe.com/blog/operating-kubernetes https://stripe.com/blog/operating-kubernetes
- seneca 6y agoKubernetes include cronjobs as an object type. Meaning, this functionality is built in. If you're already running Kubernetes, it's a very simple way to do schedule jobs without needing new management tech.
- boredatworkme 6y agooff topic: I like the simplistic design of this website. Can someone please help me understand if this is hosted on WordPress or something like that?
- cricalix 6y agoLooking at the page source, it doesn't appear to be Wordpress; none of the usual cruft that WP scatters through the source (though it could have been processed via a filter). Looking at the request/response headers in Firefox's web developer tools, it's gone through Cloudflare, but there's also a x-github header, so maybe some custom tooling that's pushed content to github?
- tonyedgecombe 6y agoIt might be Jekyll according to https://robertheaton.com/2014/07/26/lessons-from-a-surprisingly-successful-blog/ https://robertheaton.com/2014/07/26/lessons-from-a-surprisin...
- iamAy0 6y agoNowadays, many sites are built using a static site generator. Have a look at Hugo, Jekyll, Gatsby, etc. Usually you have plenty of themes to choose from. I use Hugo and can tell you there are many themes similar to the one in this website.
- swordbeta 6y agoIt's jekyll: https://github.com/robert/robert.github.com https://github.com/robert/robert.github.com
- moonchild 6y agoThey have this diagram: +-----------+ +--------------+ +-----------------+ |Web Browser| |Smartphone App| |Client Libraries/| +-----+-----+ +------+-------+ |Other API code | | | +-------+---------+ | v | | +-----+------+ | +--------->+ Steveslist +<-----------+ | Servers | +------------+ Why not this? +-----------+ +-----------------+ +--------------+ |Web Browser|-->|Client Libraries/|<--|Smartphone App| +-----------+ |Other API code | +--------------+ +-------+---------+ | v +-----+------+ | Steveslist | | Servers | +------------+
- monocasa 6y agoDifferent ergonomics of each of those environments, at least in the ideal case for each combined with the server API being a clean enough API boundary if I had to guess.
- filipn 6y agoYes, I think the second diagram that you showed is the preferred way of doing things, since using or "dogfooding" the client libraries will ensure they are always tested and correct, plus you'll get immediate feedback from the other developers who work on the apps as well.
- amarant 6y agoIn this day and age, your client libraries should be generated automatically. And of course used in your apps/webpage to save time. See openapi/swagger for good examples how client generation works
- awofford 6y agoI've been using openapi/swagger for generating my clients and honestly think it creates a mess of files and manually keeping clients up to date is not a huge task. I'm not sure that I buy into "your client libraries should be generated automatically" as a blanket statement.
- Roybot 6y ago> To do this, we need to write database queries that aggregate over the entirety of our data. We don’t want to run these queries against our production SQL database, because they could put an enormous amount of load on it. We don’t want a huge query issued by an internal analyst to be able to bring our production database to a grinding halt What kind of query would you have to write to bring down a production db? What makes a solution like hive much better - I guess its optimized for this?
- sradman 6y ago> What kind of query would you have to write to bring down a production db? Scans and Sorts, seen in a query plan, are relatively expensive to run in a production row store. OLAP queries (GROUP BY with aggregate functions like COUNT, SUM, and AVG) do large^/full table scans by definition. They take seconds to run while your goal in a Cloud OLTP system is thousands of requests per second. An automatic sort issued per query in an OLTP system is pathological and represents a vector for a DoS attack. > What makes a solution like hive much better - I guess its optimized for this? Column stores use compressed bitmap indexes that are optimized for scans over a small number of columns. Hive is SQL over Hadoop, and is inherently slow but it does offload the processing from your Production OLTP server. Hive supports the RCFile format which is partially column oriented. The ORC file format is fully column oriented, replaces RCFile format, but requires Presto (or equivalent). Hive is brownfield for existing Hadoop clusters but it has no place in a discussion about greenfield architecture other than discussing historical systems. If you have a need for GROUP BY style analytics, a true column store like Presto, Impala, or RedShift is a necessity. ^EDIT: based on zbentley's comment
- zbentley 6y ago> OLAP queries (GROUP BY with aggregate functions like COUNT, SUM, and AVG) do full table scans by definition Isn't it only a full table scan if your query isn't otherwise filtered? Those functions have to read every row of "something", but that something might not always be a whole table.
- 6y ago
- rhlsthrm 6y agoJust recently started messing with AWS Amplify. It's crazy how easy it makes basically all this stuff. It really makes it simple to be a solo founder and manage every piece of the stack.
- pantulis 6y agoGreat post!
- gmanis 6y agoI have been doing a greater part of the things elaborated in the post as a solo founder/freelance software person. Albeit the scale of things I need to handle are relatively small. What kind of a job profile should I be looking at if I am in a position where I absolutely need to have one? I don’t slot particularly well in any one thing I feel, neither a great developer nor a great systems person. And I did do a bunch of recruiting and tech consulting too:
- tda 6y agoPerhaps a job at a medium size non-tech company. There are plenty of problems that can be solved with small scale (web) apps everywhere, that don't need webscale levels of scalability, but are well suited for a single full stack developer
- intpx 6y agothe mythical 'full stack' engineer? or just sell the AL/ML snake oil and get hired to brute force tech debt.
- gmanis 6y agoThere’s nothing mythical about it. Just have been doing this for a long time across different parts of the stack and few indie products under the belt. The newer javascript frameworks are a bit tricky to wrap your head around though.
- deleted 6y ago[deleted]
- banq 6y agono DDD?
- avipars 6y agoI love the phrase "Advanced Beginner"
- forgotmypwbctbi 6y agoanyone else just getting a blank page here?
- withinboredom 6y ago> We’re considering either splitting up our cron jobs into multiple servers, or setting up a new system using a modern tool like Kubernetes. :facepalm: beanstalkd can handle a pretty massive amount of jobs before you need to start worrying about scale. I love the huge jump from simple to complex.
- zbentley 6y agoTo be fair they offered a simple option (splitting crontabs) as well.
- rataata_jr 6y agoThis is amazing. Thank you Robert.
- zeckalpha 6y agoDon’t forget a system for organizing contributions to all this code, a system to ensure compliance with the many legal regimes “Steveslist” is offered in, a system for onboarding new hires, a system for expanding market share, a system for keeping money and reporting on it to wouldbe investors, and a system of systems designers that know how to respond to changing and ever expanding requirements. Don’t sell yourself short: there are systems everywhere you look, not just where there’s convenient precedent to make a web service.
- brianzelip 6y agoReally like the prose/Adventure-like writing style of this "tutorial"!
- giggl 6y agoGreat post! Thanks for sharing.
- iblaine 6y agoSystems Design is a formidable topic because it can go in so many directions. This is a good guide. Might I also recommend something that includes caching, partitioning, indexing, and NoSQL vs SQL. More detail can be found here https://www.educative.io/courses/grokking-the-system-design-interview https://www.educative.io/courses/grokking-the-system-design-...
- humanlion87 6y agoI remember using the course you have linked to some time ago - it was very useful. I noticed now that the course is sold on a subscription model. I don't remember being that the case. Would you happen to know when that changed?
- iblaine 6y agoI don't. I came across this in 2018 and recall the course being for sale for $80.
- tinalumfoil 6y ago> database engines that are quick at small queries are typically unacceptably slow of answering giant queries Even with proper indexing? I haven't seen this issue with Postgres but maybe I wasn't working on large enough data sets.
- zealsham 6y agoThis is the best thing I have read in 2020. As a bug bounty Hunter this is insanely helpful to me .