16 ms·
What's left of NoSQL?
- h1karu 12y agoRelational databases don't scale horizontally. That's still as true today as it was a decade ago. Therefor engineers who need to cope with Big Data and Web Scale will continue to migrate away from relational database solutions towards persistence technology centered around distributed systems. It's as simple as that. If you build your app around a relational database and you need to scale up big then at some point you're going to hit a brick wall in terms of scaling out storage and/or writes. You either have to build sharding logic into your relational db app from the beginning(which is a pain that NoSQL saves you from), or else you have to re-architect your entire app when the time comes that you need to deal with scale. Many shops end up borrowing VC money to build out a team to re-architect their systems to handle web scale, but this can be avoided by thinking about data access patterns from the beginning and choosing a technology that can handle your future needs. https://en.wikipedia.org/wiki/Scalability#Horizontal_and_vertical_scaling https://en.wikipedia.org/wiki/Scalability#Horizontal_and_ver...
- workhere-io 12y agoAdWords was run on MySQL up until two years ago. The vast majority of developers working on "web-scale" projects won't get to handle projects with larger requirements than AdWords. In that perspective the whole NoSQL trend makes very little sense, especially given the fact that e.g. PostgreSQL and MySQL have a lot of needed features that many NoSQL databases don't, and that the most "hip" NoSQL database a couple of years ago was a database that doesn't even scale that well (MongoDB).
- threeseed 12y agoYou do understand how MySQL et al are used in those cases right ? They are treated as dumb key value stores and sharded horizontally with joins done in the application layer. They are NOT your typical SQL deployment and the features you talk about are often meaningless. And you are 100% wrong about MongoDB not scaling well. The stories you hear of people switching are never going back to PostgreSQL or MySQL they are going to the next level in scalability e.g. HBase or Cassandra.
- h1karu 12y agoLike I said because the relational database doesn't scale horizontally you are forced to build a sharding layer into your application which introduces complexity into the application layer and limits your ability to use the relational features of the relational database. At that point you might as well just be using a key-value store that does the sharding for you and offers greater flexibility.
- baibaichen 12y agothere is still transaction in one shard beyond SQL. I believe that ACID is the main benefit comparing with NoSQL store
- h1karu 12y agoMy NoSQL store of choices gives me document-level ACID semantics along with eventual consistency, MapReduce, and replication. Here's an example that explains how to get transaction-like guarantees from this kind of NoSQL data-store: http://guide.couchdb.org/editions/1/en/recipes.html http://guide.couchdb.org/editions/1/en/recipes.html https://en.wikipedia.org/wiki/BigCouch https://en.wikipedia.org/wiki/BigCouch
- baibaichen 12y agoInteresting, that isn't transaction, it is just a workaround, and I don't think you can design your app like that which treat document as a transaction log. And then using view to generate the real information.
- h1karu 12y agothat's exactly how I design the parts of my app that need a transactional nature. Map/reduce views make it a breeze to query for a consistent aggregate view of the world. This is not something new it's a well established technique from the relational world called "event sourcing" http://www.martinfowler.com/eaaDev/EventSourcing.html http://www.martinfowler.com/eaaDev/EventSourcing.html when everything is a write you don't have to worry about conflicts or locks, so that's nice, plus couch is really good at scaling write thoughput with small documents.
- h1karu 12y agoAdWords was designed with sharding built into it's application layer. This essentially means that the developers were forced to implement a custom app-specific persistence layer that rides on top of MySQL. This limits your ability to write SQL queries because you can only query within a particular shard and that decreases the usefulness of mysql and makes it feel more like a nosql data-store. At that point you find yourself asking why didn't we just use a NoSQL store ? And often times the answer is "because we wanted to use something we were already familiar with". Sometimes people are willing to add a lot more complexity to their application just to allow themselves to avoid having to learn something new.
- balfirevic 12y ago"At that point you find yourself asking why didn't we just use a NoSQL store" Because you still have the full ability to write general SQL within the shard. For many types of application this is useful.
- threeseed 12y agoSure. But only if ALL the data within that query exists on the shard. It is rare that this would happen if you have anything resembling a normalised schema. In which case you would still be doing a lot of joins in your application layer. IMHO Sharded MySQL very much belongs in the NoSQL camp.
- twic 12y agoI don't know that this is true. Imagine you're LeanKit, or Fog Creek, and you run a kanban board as a service. Or a bug tracker, CMS, whatever. You have many customers, each of whom has no more than thousands of users and millions of items. There are many relationships between objects belonging to a given customer, but precisely zero relationships between objects belonging to different customers. Shard using the customer identity as a key, and you have nicely spread-out data and the ability to do any query the application might need to, while still having a normalised schema. There are plenty of other application whose schemas have this property, or almost have it. In my company, we make financial applications, and a lot of the data has very similar siloed ownership structure. The one thing you can't do is reporting queries across your customers. That doesn't seem like a killer, though - it's normal to farm that stuff out to an offline reporting database even in single-server environments.
- rsynnott 12y agoThe usual thing, and I believe what was done with AdWords, was application-level sharding. This is in effect implementing your own database using MySQL as a glorified key-value store (you will NOT be doing any interesting joins, and you will not have cross-shard transactions unless you do those yourself too), and is thus not really a great argument for relational databases.
- ithkuil 12y agoSharding and SQL are orthogonal. However SQL doesn't make the problem of handling distributed databases magically disappear. An example of distributed, horizontally scalable database supporting strong consistency and offering an SQL interface: http://static.googleusercontent.com/media/research.google.com/en/us/archive/spanner-osdi2012.pdf http://static.googleusercontent.com/media/research.google.co... and a layer above: http://static.googleusercontent.com/media/research.google.com/en/us/pubs/archive/41344.pdf http://static.googleusercontent.com/media/research.google.co... It still requires the users to carefully organize their data, according to a hierarchical data model provided by the database. There is a mapping between a hierarchical relational model and a column store model. (EDIT: see one possible mapping in http://www.cidrdb.org/cidr2011/Papers/CIDR11_Paper32.pdf http://www.cidrdb.org/cidr2011/Papers/CIDR11_Paper32.pdf) The article skims over this aspect, as if all that matters is the syntax of the query language.
- gcv 12y agoOracle RAC does scale horizontally. Sort of: it requires putting your data in a SAN, but that is mostly horizontally-scalable as well.
- CraigJPerry 12y ago>> NoSQL is nothing more than a storm in a teacup There's been a difference with this "technology cycle" though, there have been some prominent, well grounded voices from the start of this cycle. That's a good thing. That wasn't the case for other cycles - the thin vs fat client cycle, the DAS vs NAS cycle etc. etc. EDIT: i'm not saying NoSQL has no application. I earn a paycheck working with a huge graph db (and it's nothing to do with "social", yay!), and have previously been a heavy user of Cassandra.
- versusdotcom 12y agoWe switched from SQL to Mongo on http://versus.com http://versus.com one year ago and are quite happy—we couldn't imagine to go back. But I think it heavily depends on the specific use case what DB to use.
- BugBrother 12y agoI agree, Mongo is easy and fast. My problem with NoSQL is that you MUST know there won't be changing requirements: Later you might need to do the equivalent of joining 3-4 tables. Then you have trouble... But I'm a coward that look both left and right before crossing a street.
- threeseed 12y agoMongoDB is actually well suited to changing requirements. Because it is schema less and document based you can trivially add new rich schemas to existing documents. And in the case of doing joins you can always do it in application layer or using DBRef or MapReduce. There are valid options that are still quite performant.
- nebstrebor 12y ago"Because it is schema less ... you can trivially add new rich schemas" made me lol. Maybe this makes sense to Mongo users, but I have no idea what you mean here and the language you use to describe this feature is, well, contradictory to say the least.
- threeseed 12y agoActually it does seem contradictory reading back on it. I meant it doesn't have a fixed, enforced schema like SQL databases and it is easy to add new data structures e.g. lists, maps, sets.
- liveoneggs 12y agoit means a typo in release-15 auto-creates a whole new database without anyone noticing and you sit around and wonder where all of the data went yesterday.
- bowlofpetunias 12y agoAs usual, the next-big-thing-to-replace-the-old-crappy-thing has become nothing more than a useful addition to our toolbox. Which is not a bad thing, but I wish we could do it without the overhyped nonsense. It really is time for this business to grow the fuck up. The "web generation" has brought us lots of great changes, I would have quit IT back in the 90's if the internet hadn't happened. But this immature attitude towards everything from technology stacks to how to run a company is really starting to grate.
- threeseed 12y agoYou are right. They are just tools nothing more, nothing less. For those of us that have been doing software development for a long time we would have heard countless vendors hyping their products. It's what they do and will continue to do. Normal developers have always ignored it and chosen the right tool for the job. I actually find the anti-NoSQL crowd to be the immature ones and are often quite patronising as though I am some idiot for choosing a particular database. It's quite bizarre.
- bunderbunder 12y agoThe "web generation" has brought us lots of great changes, I would have quit IT back in the 90's if the internet hadn't happened. But this immature attitude towards everything from technology stacks to how to run a company is really starting to grate. Reminds me of back in the 90's when object-oriented databases were going to rule the world and software was all going to be write-once-run-everywhere, or the early 2000's when static typing was the downfall of society. Meh. Every generation's got something like that. As skeptical as I am of the latest trends, I sincerely hope that people stay excited about them, and keep using them. That bubbling cauldron of ideas does generate a lot of noise, but it also generates new ideas and progress.
- mathnode 12y ago> NoACID Perfect!
- frik 12y agoWe are stuck with NoSQL in HTML5, because Mozilla and Microsoft refuse to implement WebSQL (http://en.wikipedia.org/wiki/WebSQL http://en.wikipedia.org/wiki/WebSQL ) IndexedDB is fine for storing JSON objects, etc. but a relational database with SQL query syntax, indexes, etc. more powerful and means less code to write. With IndexedDB one has to reinvent the wheel to just get basic query features. WebSQL is not deprecation, the W3C Working Group Note actually says: 'This specification is no longer in active maintenance and the Web Applications Working Group does not intend to maintain it further'. WebSQL is only available in Webkit based Browsers (Safari, Chrome) which means most mobile browsers. As SQLite is in public domain, no company would "loose their face" if they choose to use it. They could fork off SQLite and change the SQL query syntax (parser) to whatever the W3C finds suitable. https://www.sqlite.org https://www.sqlite.org Mozilla Firefox and FirefoxOS both already ship SQLite for years and can be accessed by its internal JavaScript API. And several Microsoft products already use it anyway (e.g. Forza Xbox games). Microsoft has of course also various other SQL database libraries like MS Access JetRed, MS Outlook JetBlue and SQL Express. We had a discussion about it recently: https://news.ycombinator.com/item?id=7574754 https://news.ycombinator.com/item?id=7574754 The new hip things is "NewSQL" (http://en.wikipedia.org/wiki/NewSQL http://en.wikipedia.org/wiki/NewSQL ). For example Facebook, Google Ads, etc. are powered by MySQL's InnoDB database engine. I would go as far as count SQLite to this group. We would need a movement to convince Mozilla to finally add WebSQL to Firefox and FirefoxOS.
- davidjohnstone 12y agoThere are good reasons why work has stopped on the WebSQL specification — everybody was using SQLite, and the specification can't be tied so closely to a single implementation of SQL. The specification even has the line "User agents must implement the SQL dialect supported by Sqlite 3.6.19"[1]. 1. http://www.w3.org/TR/webdatabase/#web-sql http://www.w3.org/TR/webdatabase/#web-sql Edit: here is Mozilla's rationale for not supporting WebSQL: https://hacks.mozilla.org/2010/06/beyond-html5-database-apis-and-the-road-to-indexeddb/ https://hacks.mozilla.org/2010/06/beyond-html5-database-apis...
- mmahemoff 12y ago
- phpnode 12y agoNoSQL has always been a stupid term, especially as there are "NoSQL" datastores that support SQL, e.g. OrientDB.
- crusso 12y agoThe term was perfect. The spectrum of NoSQL implementations out there defies an easy name to cover them all. They aren't all key-value stores, they aren't all json stores, they aren't all big tables, etc. Indicating that they were NOT within the current paradigm of SQL databases was the only way to go.
- danford 12y agoIt's stupid if you don't realize it stands for "Not only SQL".
- zaidf 12y agoHow do people that abstract away their database via an ORM feel comfortable about not dealing directly with the data and accidentally dropping a db column via ORM? I know when I am dealing with the database I am a lot more surgical in my approach than when I am writing code.
- CmonDev 12y ago"dropping a db column via ORM" - prevent this via security configuration, release schema changes to production via plain SQL. ORM is not a complete replacement for dealing with DBMS.
- threeseed 12y agoThe same way I feel comfortable that I won't accidentally write and execute a SQL script that does it. ORM doesn't make it especially trivial to corrupt your database.
- josephlord 12y agoSchema changes are often (normally?) done using a separate mechanism and operations called migrations rather than in the main code. In a production environment they should be well tested before deployed to the live system. That plus backups should help prevent calamities. There are still possibilities of deleting the data whether through the ORM or direct SQL, both could be easily accomplished so test before deployment and keep regular backups in case of disaster.
- StevePerkins 12y agoYou generally only let the ORM create/update your schema during early development. By the time you get to production (or heck, by the time your team is even seriously engaged), you will have disabled that feature. In pretty much every enterprise-grade ORM I've ever seen, schema modification is disabled by default, and must be explicitly turned on. In most Java shops where I've ever worked... unless it's a quick prototype, the database schema will be developed before you start coding anyway. Also, in a typical enterprise scenario, the username with which you configure the ORM lacks admin-level privileges to alter the database. Most of the criticisms about ORM's come from people who have never used them, beyond maybe working through a Rails chat-room tutorial once upon a time. It really has nothing to do with "abstracting away the database". As the name indicates, Object-Relational Mapping is merely about reducing the boilerplate required to map a relational schema to programming language objects. If you do that mapping by hand, then you have to make decisions when a table/object has relationships. Picture a CUSTOMER table, which has a foreign key relationship to an ADDRESS table. When your application loads a "Customer" object: [1] You could "eager fetch", meaning that you go ahead and retrieve all of the ADDRESS rows related to that CUSTOMER, and attach the Address objects to the Customer object. Eager fetching is wasteful and leads to poor performance, because you're hammering your database for values that you often don't ever use. [2] You could "lazy load"... meaning that your Customer object has an "addresses" field, but you wait until some code tries to use that field before you actually query the ADDRESS table to populate it. This is much better design, but complicates things. The lazy load logic has to go somewhere. You either have to ensure that every piece of code using that object is aware of the lazy load pattern, or you have to stuff database logic into the "getter" method for each lazy-loaded field. ORM's give you highly-performant lazy loading, without the buggy boilerplate suggested by #2 above. Moreover, enterprise-class ORM's typically handle caching for you, to avoid hitting the database unnecessarily. Monitoring the state of objects to notice when they've gone stale due to changes on the database side, etc. Lastly, for complex queries, most ORM's have a query language (e.g. JQL, HQL, etc) that is nothing more than a VERY THIN wrapper around SQL. It merely smooths out differences between various vendor dialects. You're not abstracted away from the database, you still very much need to understand SQL and the underlying structure of your data.
- tempodox 12y agoFinally a good writeup that explains the strange phenomenon of NoSQL. Now, all that's left is to introduce a tax for “drinking the Coolade”.
- redwood 12y agoThis is an incredibly short-sighted piece for two reasons 1) the data challenges of yesterday's internet giants are the data problems for just about every 21st century enterprise tomorrow. We're at the start of / in the midst of an irreversible data explosion. 2) Established players not taking new technologies and upstarts seriously is never evidence of the threat not being real. We know that on the contrary our industry is characterized by constant disruption where established players for whatever reason do not consistently stay at the front line of innovation and tend to get leapfrogged. (as an analogy Google+'s lurch into social may be akin to Oracle's lurch into nosql. Does that mean social was not real?) Bottom line: It all depends on what you're doing. But more and more of us will be dealing with more and more complex data in the years to come. Of that, I'm certain.
- jhorey 12y agoI agree that's a bit short sighted, but I think it's one of perspective. The mistake most people make is to believe that Oracle (and other relational DB vendors) are competing directly against the various NoSQL stores (this includes doc stores, scalable key-value stores, etc.). I don't think that's true. I think that relational databases are really competing against industry-specific SaaS. So instead of implementing their own database for inventory & sales, companies may opt to use a service. These SaaS companies, in turn, are more likely to adopt a variety of technologies (including NoSQL and relational). Since the SaaS companies serve more than a single company, they're also more likely to adopt easily scalable technology like Cassandra. So fewer businesses will need to purchase databases (of any sort), but the ones that do will have greater scalability needs.
- redwood 12y agoAgree. Where you see nonrelational popping up analogously is particularly in new SaaS shops, many of which aim to take some Oracle's pie.
- threeseed 12y agoNoSQL is absolutely competing with Oracle and other relational DB vendors. You are seeing it today with the valuations for companies like Mongo or DataStax. SaaS is hardly a big enough market compared to say every enterprise. And surely the companies listed on the client pages agree with me. And people forget something important with SaaS. Data sovereignty. Here in Australia for example there are many enterprise companies who are forbidden from using ANYTHING that is hosted in the US due to the grey legal area e.g. Patriot Act. So in-house databases are absolutely still here to stay.
- Arkadir 12y agoNoSQL discussions always seem to conflate three very different things: storage engines, APIs and architecture. Where do we store the data ? How do we access it ? How do we make sure it scales ? The "traditional" approach is to use Oracle/SQL Server/MySQL for storage, SQL and/or ORM as an API, and single-server tables-with-relationships as an architecture. Back in the early 2000s, everybody did this. Sure, there were a few performance-minded exceptions that went with sharding or master-slave architectures instead, but those were exceptions. And single-server architectures tend to behave badly at medium loads. Spend the market rate for a genius DBA, and they still behave badly at high loads. The next step is a 32-core 128GB RAM monstrosity that costs an order of magnitude more than what eight 4-core 16GB servers would cost. Most NoSQL solutions came with a new architecture. You had the MongoDB flavor of distributed storage, or the BigTable flavor of distributed storage, or the CouchDB flavor of distributed storage, and so on. Properly implemented distributed storage eats high loads for breakfast: just add more servers. This is a good thing. My issue with the NoSQL movement is that they threw away the baby with the bath water. They threw away the single-server relational architecture, which was a nice change, and they also gave up the old battle-hardened storage engines and the highly expressive SQL language and replaced them with only-recently-experimental engines and ad hoc lean APIs. It takes time for a storage engine to mature. To have all its performance kinks ironed out and all its bugs smoked out. I still remember the brouhaha around MongoDB persistence guarantees, or the critical data loss bugs in CouchDB. And the lean APIs just forced back all the querying logic into the application, with all the filtering and the manual indexing and the joins and the approximate but ultimately incorrect implementations of whatever subset of ACID was required at the time. This wasn't an entirely bad thing: it certainly made many developers aware of the performance implications of some joins or transactions. But when you need to write a JOIN or GROUP BY or BEGIN TRANSACTION that you know will scale properly, and there's no API support for it ? Feh. I'm a huge fan of the CouchDB architecture. Distributed change streams, with checkpointed views and cached reductions. But I have been burned by the CouchDB storage engine (can you say "data corÊ–NÑ %ñXtion" ?) and I see no point in bending knee to the laconic CouchDB API. So I took the CouchDB architecture and reimplemented it with a PostgreSQL back-end. It's _faster_ (don't underestimate the cost of those HTTP requests), I have trust that after PostgreSQL's decade-long history all threats to my data are long gone, and I can always whip out an SQL query when I do need it. It's nice to see so many NoSQL solutions migrating back to an SQL-like API and gaining enough maturity to keep your data safe. In the near future, I expect them to be nothing more than "Architecture in a box" solutions for when you don't want to implement specific architectures in SQL. And I expect more and more "architecture plugins" to become available: with a library, turn a herd of SQL databases into a distributed architecture of type X.
- bitL 12y agoI am wondering if the prolonged latencies from most of the Google services I am experiencing in the past few years are somehow related to using Spanner instead of the previously conventional "compute in advance offline"-"serve immediately" approach. The Spanner paper mentioned they sacrificed a bit of latency (from nanoseconds in DHT lookup to milliseconds). I understand it's more convenient for Google developers to have more predictable thought framework instead of making every single project a piece of art while spending most of the time fighting inconsistencies arising from eventual consistency. The question is if it was worth it? I remember time when Google services were amazingly fast - that time is unfortunately gone :-(
- nevi-me 12y agoIn what context are you referring to the "compute in advance offline" part? Sometimes it's costly to compute things in advance and store them, when you might not need them. If storage space becomes a concern, then Google would have to write some fancy algorithms based on data access patterns that determines what to compute in advance, and what to compute JIT. Which complicates things because now they have to write and regularly run programs that tell them what to compute, which simplistically means that they are at least running 3 different tasks, instead of 1 on-the-fly computation.
- bitL 12y agoI meant the traditional NoSQL architecture (like what LinkedIn was using) where you had offline batch processing based on Map Reduce executed regularly during the day and the results were stored in a distributed hash table with super low latency and eventual consistency. This made access super fast but inflexible, i.e. only what was precomputed could have been accessed.
- room271 12y agoAn interesting article and it does seem like people rushed to embrace NoSQL and are now trying to force it into some kind of consistency after the event - not a bad thing, incidentally - there's a lot of interesting work here (Vector clocks, CRDTs, etc.). One thing that surprised me though was the lack of a key player: Amazon. Their Dynamo paper was hugely significant and as a company they use eventually consistent stores for a whole swathe of products at scale. Why mention Facebook and Google but omit this other major player, especially are their experiences tell a different story.
- bitL 12y agoAmazon estimated that each 1 ms of additional latency costs them a few million dollars a year as well as decreases rate of returning users. For low latency there is still nothing better than eventual consistency, so they may be driven by the bottom line. Also, from my experience of being a merchant on Amazon with hundreds of thousands of items in inventory that need to be updated almost realtime as prices/stock changes all the time, leading to a few millions updates a day (in bulk), I can't envision how Amazon would be able to achieve fast import from a horde of merchants like me on any SQL system.
- room271 12y agoYeah, the Dynamo paper is very clear about the rationale behind creating DynamoDB and its usage: - downtime (lack of availability) is very expensive to them, both directly and also reputationally - traditional ACID stores don't scale with writes whereas an eventually consistent store can achieve this So I completely agree with you in short: Amazon have great reasons to pursue approaches based on eventual consistency.
- guard-of-terra 12y agoIf you're doing accounting, sure you would want ACID compliant database. There you have a limited amount of kinds of data to store with strict and rarely changing constraints. You can keep it on one big expensive server (plus backup) and it's better to bring the system down to allow inconsistency. However, for most web development SQL is seriously not good. You end up with hundreds of loosely coupled tables which constantly change their structure for new features. Half dozen of joins on every request. Hundreds of lines of SQL. Constant pain. And it's not like you cared so much for the consistency - if once a year three comments disappear from your web site, so what? And it's painful to make SQL multi-master. For (the most of) web development document databases are so much better. MongoDB is pretty nice because it makes hundreds of lines of SQL with ten files of code redundant - all per one complex document.
- deleted 12y ago[deleted]
- dchuk 12y agoI'm far from an expert on this topic, but I am a developer, and I am currently working on a project that leverages both MySQL and ElasticSearch for storage. MySQL handles all of the "boring" data such as users, profiles, comments, etc. ElasticSearch is basically a giant product database, no real relational data, just a normalized structure and easy to query. ElasticSearch is serving as a primary data store, there is no backing in a database because ES is just that good. Whenever I see these "SQL vs NoSQL" arguments, I always have to wonder: Why one over the other? A lot of projects can benefit from both and there's no reason you absolutely HAVE to use one or another. It's perfectly reasonable (and probably ideal) to use more than one storage system in your projects. If you have a bunch of nails to hammer in and bolts to tighten, you don't choose just a hammer or a wrench to do that job...you grab both and use each for what they do best.
- jcroll 12y ago> ElasticSearch is serving as a primary data store, there is no backing in a database because ES is just that good. And you've been running this application in production how long?
- room271 12y agoI agree that using Elasticsearch as a primary data store seems risky. The situation is improving though; the v1 release introduces easy backups (for example, to s3) which from my experience work very well. Prior to this backing up was a messy process.