4 ms·
I can't even count how many mistakes this article makes. A random sampling: - Relational databases were designed for a world where availability is unimportant
by jaylevitt 15y ago
I can't even count how many mistakes this article makes. A random sampling:
- Relational databases were designed for a world where availability is unimportant, like transaction processing. Um, no. OLTP has four letters, not two, and the first are as important as the last. Tandem was providing five-9's systems in the '80s and '90s for ATM networks, lottery systems, airline reservations, credit cards, etc. Tandem is a fault-tolerant relational database in hardware.
- Up-front schema design is a poor fit in a world where data requirements are fluid: I don't think "up front" means what you think it means. You can change schemas on the fly nowadays; your schema design is no more up-front than your coding is.
- You can't have millions of columns in a relational database: True, and you wouldn't; you'd normalize that. This is an important difference, but not a disadvantage of relational databases, any more than saying "In a relational database, you'd join URLs with IP addresses, and maybe five other tables; this design isn't even conceivable in a NoSQL database."
- To optimize relational performance, you "do away with joins wherever possible": 1995 called, and it wants its MyISAM back.
- Two-phase commit is so obsolete, even banks don't use it: Of course they do. You still need a two-phase commit to make sure the other end got your data; whether "got your data" happens in the customer path or during reconciliation is a design decision. How, exactly, do you think they discover that you and your spouse both got the money?
- "relational databases were developed when distributed systems were rare and exotic at best." That's nothing; when Von Neumann machines were developed, we didn't even have transistors. Some legacy architectures keep on working.
- "absolute consistency isn't a hard requirement for banks": see above. Yes it is.
- "So the CAP theorem is historically irrelevant to relational databases: they're good at providing consistency, and they have been adapted to provide high availability with some success, but they are hard to partition without extreme effort or extreme cost." ... Wh... Bu... That's not even wrong.
- "consistency requirements of many social applications are very soft." I like the Facebook example from a recent article on causal consistency: I de-friend my boss and then post that I'm quitting. Certain kinds of consistency are in fact critical to social applications.
There are many good reasons to design around a NoSQL database instead of a relational one. This article provides fewer than zero of them.
- mjb 15y agoI agree with you - this article does a very poor job of communicating the advantages of 'NoSQL' databases over traditional solutions. It spends too many words setting up a straw man (SQL == not partitioned, not highly available, not redundant, etc) and too few actually making it's case. Lines like this are, at best, a gross oversimplification of relational database performance tuning: "But when you need to optimize performance, you look at the queries you actually perform, then merge tables to create longer rows, and do away with joins wherever possible." And this: "We require sub-second responses to queries." Really? Single big-box OLTP systems have many issues, but if you are getting query times greater than a second on a typical online database then you are either doing something wrong or have specific requirements. Neither of those things will be fixed by blindly picking 'NoSQL' over 'SQL'. "any significant database needs to be distributed." I could easily list, off the top of my head, 100 significant databases that aren't distributed. Maybe some of them should be, but saying that "any significant database needs to be distributed" is a real stretch. > There are many good reasons to design around a NoSQL database instead of a relational one. This article provides fewer than zero of them. Yeah, that's the thing. An article of a quarter of the length with more research, less strawman bashing and less breathlessness could have made the case for NoSQL much more effectively.
- opendomain 15y ago- availability means access to the data, not just uptime - I have never seen changing schemas on the fly in a RDBMS programmatically (except admin functions). I have seen overloading of column types however. - "super column" data stores have a deistinct design advantage that you can not get from relational DBS. - everyone that is using relational does the same thing when they get big data: denormalize. From materilzed views al the way to sharding and replication. - banks use transactions, but not as a two phase commit. it is a single atomic record - not partial updates. - we still have some species from the age of the dinosaurs, but the land is rules by those that have evolved. - I agree that banks should be consistent. Not sure what he author was saying here - CAP is relevant to ALL datastores. Consistancy, Availability, Partition tolerance : pick any two. - if your post does not get committed to Facebook, it may be important to YOU, but the application is designed more for availability