9 ms·
> Designing scalable systems when you don't need to makes you a bad engineer. > In general, RDBMS > NoSql These two bullet points resonate with me so much rig
by cbdumas 6y ago
> Designing scalable systems when you don't need to makes you a bad engineer.
> In general, RDBMS > NoSql
These two bullet points resonate with me so much right now. I'm a consultant and a lot of my client absolutely insist on using DynamoDB for everything. I'm building an internal facing app that will have users numbering in the hundreds, maybe. The hoops we are jumping through to break this app up into "microservices" are absolutely astounding. Who needs joins? Who needs relational integrity? Who needs flexible query patterns? "It just has to scale"!
- virmundi 6y agoAre there other constraints that might make DynamoDB a good fit? For example I made an app at a client. We could use RDS or we could use Dynamo. I went with Dynamo because it could fit our simple model. What’s more, it doesn’t get shut off nightly when the RDS systems do to save money. This means we can work work on it when people have to time shift due to events in the life like having to pick up the kids.
- tomnipotent 6y ago> it doesn’t get shut off nightly when the RDS systems do to save money If your company needs to shutdown RDS to save a couple of bucks a month, there's a much larger problem at hand than RDS vs Dynamo.
- Aeolun 6y agoAt scale it’s a little bit more than a few bucks. Across the board, we spend hundreds of thousands on ec2 instances for dev/test, so turning them off at night when nobody uses them saves you quite a lot of money.
- martinald 6y agoThe problem with NoSQL is that your simple model inevitably becomes more complex over time and then it doesn't work anymore. Over the past decade I've realised using a RDBMS is the right call basically 100% of the time. Now pgsql has jsonb column types that work great, I cannot see why you would ever use a NoSQL DB, unless you are working at such crazy scale postgres wouldn't work. In 99.999% of cases people are not.
- xupybd 6y agoThere are specific cases where a non SQL database is better. Chances are if you haven't hit problems you can't solve with an SQL database you should be using an SQL database. Postgres is amazing and free why would you use anything else.
- maccard 6y agoPeople keep saying there are specific cases where NoSQL is better, but never what any of those cases are.
- throwaway870923 6y agoTime series is one. Consider an application with 1000 time series, 1 hosts, and 1000 RPS. You are trivially looking at 1M writes per second per host. This usually requires something more than "[just] using a RDBMS".
- kfir 6y agoHere you go, this is from a system I helped building 10 years ago that is an eternity in tech - https://qconlondon.com/london-2010/qconlondon.com/dl/qcon-london-2010/slides/GeirMagnusson_ProjectVoldemortAtGiltGroupeWhenFailureIsntAnOption.pdf https://qconlondon.com/london-2010/qconlondon.com/dl/qcon-lo...
- kfir 6y agoa bit more context: High-velocity transactional systems (e.g any e-commerce store with millions of users all trying to shop at the same time), I helped to build such a 10 years ago here is the presentation - https://qconlondon.com/london-2010/qconlondon.com/dl/qcon-london-2010/slides/GeirMagnusson_ProjectVoldemortAtGiltGroupeWhenFailureIsntAnOption.pdf https://qconlondon.com/london-2010/qconlondon.com/dl/qcon-lo...
- Joeri 6y agoWe just ported a system that kept large amounts of data in postgres jsonb columns over to mongodb. The jsonb column approach worked fine until we scaled it beyond a certain point, and then it was an unending source of performance bottlenecks. The mongodb version is much faster. In retrospect we should have gone with mongo from the start, but postgres was chosen because in 99% of circumstances it is good enough. It was the wrong decision for the right reasons.
- cbdumas 6y agoI can't speak to your specific use case, but I can tell you that a relatively small RDS instance is probably a lot more performant than you think. There is also "Aurora Serverless" now which I've just started to play with but might suit your needs. As far as what makes Dynamo a good fit, I almost take the other approach and try to ask myself, what makes Postgres a bad fit? Postgres is so flexible and powerful that IMO you need a really good reason to walk away from that as the default.
- virmundi 6y agoAurora wasn’t allowed at the time. The system is a simple stream logging app. Wonderful for our use case. Dynamo for so far. Corp politics made the RDS instance annoying to pursue.
- bawolff 6y agoPeople always talk about nosql scaling better, but some of the largest websites on the internet are mysql based. I'm sure some people have problems where nosql is genuinely an appropriate solution, but i find it hard to believe that most people get anywhere near that level of scalability.
- cbdumas 6y agoExactly, and from a features standpoint Postgres can do everything Dynamo can do and so much more. I think a lot of software devs don't really know SQL or how RDBMS work so they don't know what they are giving up.
- whymauri 6y agoThis is similar to how I feel about graph databases. Twitter (FlockDB) and Facebook (TAO) built scalable graph abstractions over SQL without a hitch. Why would I want to use a graph DB directly then?
- GordonS 6y agoPostgres even has JSONB support, so if you really want to store whole documents NOSQL-style, you can - and you can still use all the usual RDBMS goodness alongside it! Postgres really is a wonderful database.
- peteradio 6y agoPOC scales easier, thats all that matters to win the idiot match.
- Joeri 6y agoThose very large mysql deployments typically use it as a nosql system, with a sharded database spread over dozens or hundreds of instances, and referential integrity maintained by the business layer, not by the database. For a good example of a high volume site using a proper rdbms approach I would look at stackoverflow. It can (and has) run on a single ms sql server instance.
- 6y ago
- PragmaticPulp 6y agoAs an engineer-turned-manager, I spend a lot of time asking engineers how we can simplify their ambitious plans. Often it’s as simple as asking “What would we give up by using a monolith here instead of microservices?” Forcing people to justify, out loud, why they want to use a specific technology or trendy design pattern is usually sufficient to scuttle complex plans. Frankly, many engineers want to use the latest trends like microservices or NoSQL because they believe that’s what’s best for their resume, even if it’s not necessarily best for the company. It doesn’t help that some companies screen out resumes that don’t have the right signals (Microservices, ReactJS, NoSQL, ...). There’s a certain amount of FOMO that makes early-career engineers feel like they won’t be able to move up unless they can find a way to use the most advanced and complex architectures, even if their problems don’t warrant those solutions.
- blabitty 6y agoI like to half jokingly assert that microservices are a pysop to sell cloud hosting
- naikrovek 6y agoI bet you're half-right as well.
- nitrogen 6y agoIf I were an evil tech giant, I would open source a bunch of libraries that require significantly more effort to use than necessary, and pitch them as the One True Solution. Just to slow my competitors down.
- specialist 6y agoI had a nemesis who would steal all my ideas. So I bought all the XP books, dog eared them, left them on my desk. My team nearly mutinied. I asked them to wait and see. Two weeks later, nemesis announced his team was all in for XP, Agile, pair programming, etc. They never recovered, didn't make another release. I tossed my copies, unread.
- d23 6y agoEh, the second one is probably the only point I was kind of meh on. You should almost always start with an RDBMS, and it will scale for most companies for a long long time, but for some workloads or levels of scale you're probably going to need to at least augment it with another storage system.
- deleted 6y ago[deleted]
- Aeolun 6y agoDepends? If you know you are going to need that scale you can take it into account when selecting your RDBMS technology/setup.
- k__ 6y agoI think, it's mostly a question of education. Universities taught SQL for years, so everyone knows it and its edge cases. NoSQL databases are all different AND they weren't all taught for decades. If you put real effort into learning a specific NoSQL database and it is suited for your problem things work out pretty well.
- specialist 6y agoI was on a team which used DynamoDB for their hottest data set. Which would trivially fit in RAM.
- busterarm 6y agoIf I had a dollar for every senior engineer I've worked with who has never heard of SQLite...
- xxs 6y agoTruth be told I am yet to see a reason to use in-memory database. Datstructures, maps/trees/set - yes. Concurrent/lock free/skip lists/whatever - all great. I don't need a relational database when I can use objects/structs/etc.
- busterarm 6y agoACID transactions, validations & constraints, and the ability to debug/log by dumping your data to disk which can then easily be queried with SQL. All of the same reasons you would store relational data in a dbms...
- xxs 6y ago>ACID transactions, validations & constraints There is no D from the ACID. For the D to happen, it takes transaction logs + write barrier (on the non-volatile memory). Doing Atomic, consistent and isolated is trivial in memory (esp. in GC setup), and a lot faster: no locks needed. Validations and constraints are simple if-statements, I'd never think of them as sql.
- piggubiggu 6y agoIt sounds like you're talking about toy databases which don't run at a lot of TPS. Let me point out some features missing from your simple load a map in memory architecture. You also have to do backup and recovery. And for that, you need to write to disk, which becomes a big bottleneck since besides backup and checkpointing there is no other reason to ever write to disk. Then, you have to know that even in mem database, data needs to be queried, and for that you need special data structures like a cache aware B+tree. Implementing one is non trivial. Thirdly, doing atomic, consistent and isolated transaction is certainly trivial in a toy example but in an actual database where you have a high number of transactions, it's a lot harder. For example, when you have multiple cores, you certainly will have resource contention, and then you do need locks. And last thing about gc, again, gc is great, but there has to be a custom gc for a database. You need to make sure the transaction log in memory is flushed before committing. And malloc is also very slow. I'd suggest reading more into in mem research to understand this better. But in mem db is certainly not the same as a disk db with cache or a simple Hashmat/B+tree structure.
- potta_coffee 6y agoI don't understand splitting an API into a bunch of "microservices" for scaling purposes. If all of the services are engaged for every request, they're not really scaled independently. You're just geographically isolating your code. It's still tightly coupled but now it has to communicate over http. Applications designed this way are flaming piles of garbage.
- Nycto 6y agoThe first thing that comes to my mind is that there are different axes that you may need to scale against. Microservices are a common way to scale when you’re trying to increase the number of teams working on a project. Dividing across a service api allows different teams to use different technology and with different release schedules.
- potta_coffee 6y agoI don't necessarily disagree, but I believe that you have to be very careful about the boundaries between your services. In my experience, it's pretty difficult to separate an API into services arbitrarily before you've built a working system - at least for anything that has more than a trivial amount of complexity. If there's a good formula or rule of thumb for this problem, I'd like to know what it is.
- Nycto 6y agoI agree. From my perspective, microservices shouldn’t be a starting point. They should be something you carve out of a larger application as the need arises.
- captrb 6y agoEspecially true when the services are all stateless. If there isn’t a conway-esque or scaling advantage to decoupling the deployment... don’t. I had a fevered dream the other night where it turned out that the bulk of AWS’s electricity consumption was just marshaling and unmarshalling JSON, for no benefit.
- vmception 6y agoIn my own different comment I highlighted the same two points with the opposite conclusion haha! I find dynamodb to be unnecessary but I prefer nosql systems
- mathattack 6y agoI’ve seen similar issues where people got stuck on Mongo because it’s easy to install.
- jugg1es 6y agoI can't think of a worse decision than trying to use DynamoDB just for the sake of using it.