14 ms·
Genuinely curious: how does a project like Dropbox grow to four million lines? In my (albeit naïve/inexperienced) mind, I could see how a file syncing app coul
by linux2647 7y ago
Genuinely curious: how does a project like Dropbox grow to four million lines? In my (albeit naïve/inexperienced) mind, I could see how a file syncing app could grow to several thousand lines (even tens of thousands), but what all is in those four million lines? Is it a lot of UI code? Backwards compatibility for older data models? Building components themselves instead of just `pip install`ing?
- jdm2212 7y agoPer the blog post, it's both the client code and server code. The server code is >75% of the annotated Python. Which makes sense, given the amount of stuff you have to do to run one of the world's larger distributed file system with tight SLAs (and auth, and your own data centers, etc).
- paggle 7y agoClient support for every random configuration you can imagine (Windows XP in Arabic for example) and a ton of server side performance optimization.
- ampersandy 7y agoLet's just dig into one aspect of a file syncing app: storing the files at Dropbox scale. Without spending too long thinking about it, I'd say we need: * An API to request files/upload files * Authentication/encryption (how do we ensure DB employees aren't arbitrarily reading data on servers?) * Service that shards data and handles multi-region replication * That entails multiple datacenters (how much code does it take to keep a DC running?) * Automated backups to cold storage * Automated restoration + testing of cold storage volumes * Soft + hard deletion (deleting hot/warm/cold storage volumes reliably) * Error handling (retries, host errors, network failures, filesystem corruption, finding data when a host dies) * Fail-safes like blocking requests when hosts can't handle them, shedding load, etc. I'd also add that at this scale, very few libraries start handling the types of problems you have. For example, there might be a memcache library that makes writing queries easy. However, when you have thousands of memcache servers globally you can't just hardcode IP addresses anymore. You know have to write your own custom 'service-service' that lets you look up what host you should talk to for something.
- abernard1 7y agoI could imagine a list 10X that size that still does not need even 400K lines of Python code, much less 4M. I agree that people generally underestimate the lines of code needed to operate systems at scale, but having used dropbox, I'd still say that's excessive. Note, this is just the Python code, not the UI stuff or the golang or Rust. I've been in multiple 1M-10M line codebases at scale and I just cannot fathom with how with their product simplicity (not engineering simplicity) they could be at that size with a language as expressive as Python. My guess is this is generously counting a lot of forked libraries. Ungenerously, it makes me think there's a lot of NIH syndrome.