5 ms·
A Comprehensive Guide to Moving from SQL to RethinkDB
- baldfat 12y agoI still struggle for why someone would go RethinkDB over Hardoop. If you have the "Big Data" issue for using a Non-SQL DB (Another struggle of mine to find a clear case for its use). Why go RethinkDB when you limit what you can do with your data at the start. You can't do so things but the big one for me is you really can't do statistical analysis (Author pointed it out in the article). Why have a Non-SQL that really hinders the biggest selling point for the non-SQL DB? Wouldn't it just be better to use a SQL and expand the scheme horizontally??? Serious question that I don't understand.
- mglukhovsky 12y agoRethinkDB and Hadoop solve fundamentally different use cases. RethinkDB is designed for building scalable, realtime apps -- the query language, realtime push features (e.g. subscribing to queries), and clustering are purpose-built to help developers build and scale realtime apps. Hadoop is designed for distributed processing of large data sets, and is generally used for analytics workloads and data processing. NoSQL really just means "not using SQL". A lot of projects fit under that umbrella, including graph databases, key-value stores, document databases, and analytics databases, and they all address very different use cases.
- ma2rten 12y agoFirst, this is about scalability of your application and not your data analysis. If your application itself has a lot of writes, it may require to shard your SQL database. That may be easier using NoSQL. Also, NoSQL has advantages other than scalability. RethinkDB has a more pleasant query language (arguably), shamelessness (meaning it is easy to add a column), build-in ui, ...
- mberning 12y agoI love me some document based databases. But I find it ridiculous how much FUD there is out there about them. After being severely burned by some "It's not fully ACID" types in my organization I am seriously reluctant to even use them on projects where they are ideally suited. Sorry for the rant. I think they are wonderful tools, but I also think people should know about the political implications of using them in organizations where 'old fashioned' or conservative thinking might come in and ruin their day.
- ppj606 12y agoI get your point of view. My hypothesis for this behaviour, or part of the reason for the behaviour, is that I think the education system has traditionally been geared to teaching about SQL systems over their NoSQL counter-parts. This creates developers with strong opinions of one, without necessarily having a sound understanding of the other(s). The people who cause pain in your organisation are almost definitely committing a logical fallacy - see [https://yourlogicalfallacyis.com/ https://yourlogicalfallacyis.com/]
- matthewmacleod 12y agoI'm not sure the level of FUD is unjustified. I've seen far, far more cases of document databases being used in situations that are entirely inappropriate, than in situations where they are ideally suited. I'd argue that the situations in which it's a good choice to use a document database are pretty rare, and that you're almost always better off using a bog-standard relational database.
- mberning 12y agoI think many problems could be solved equally well using either technology IF you understand the strengths and weaknesses of each technology and implement them appropriately. It just so happens that knowledge of how to use a RDBMS appropriately is MUCH more common than document databases.
- ppj606 12y agoWith the rise of realtime applications I'd argue that the situations for using NoSQL for the common project are growing. I think the author of the post has made a good point, that many seem to forget - it is always a trade-off whichever tool you use.
- moe 12y agoWith the rise of realtime applications What does "realtime" have to do with the datastore?
- NDizzle 12y agoAfter spending a decade working with Lotus Domino, I think all of these people who are willingly going back to document style databases are insane. Being able to fiddle with your schema on the fly isn't necessarily a good thing. Brainstorm, plan, execute. Or just slap another field on the end of the document. Whichever sounds better for the long term health of your data!
- jchrisa 12y agoGood to see a complaint about NoSQL that is at least targeted in a reasonable direction. Personally I see a lot of success coming from flexible document schemas but obviously it's not perfect for every app. What did you think about Domino replication?
- NDizzle 12y agoI have used it extensively. It's brittle. It's brilliant. The way Domino works at a high level is that you have a "data" and "design" template applied to a database. A database would be a single file on disk, and it's something like a table (or a set of related tables) and various views attached to that data. We had an internal Domino server with a set of internally focused 'design' templates applied to them. We used the notes clients, and our analysts had their workflow applications there. The design of these databases was focused on serving content to the notes clients - not to a web browser. We replicated the data from those databases to our staging and development machines. Those staging and development machines had their own designs - which was the web-centric, customer focused design. There was a minimum amount of information available to you in the notes clients - data at this point was meant to be seen in a browser. Those staging machines pushed the data on to our final production cluster. Every 15 minutes something was pushed. We'd go from internal to dev/stage 15 after the hour. 30 minutes after the hour the data from stage would go to production. We used normal build/release tools control the flow of design template replicas making their way out of their playgrounds, but sometimes it happened. I say it's brittle because if you accidentally push design to the wrong place you could grenade a lot of stuff. More than once we accidentally flipped the switch to push design changes to development. Here's a link to a presentation of the last thing I was a part of in the Lotus community. I'm pretty proud of the site we put together. It had a pretty nice feature set for the time period. This was put together by a coworker. https://www.youtube.com/watch?list=PL6D93ED85F970BAE6&v=v9IUly7np0A https://www.youtube.com/watch?list=PL6D93ED85F970BAE6&v=v9IU...
- robodale 12y ago...and why the hell would you want to do this.
- jimmytucson 12y agoSome of these examples are not equivalent. For example, r .db('dragonball') .table('characters') .filter(function(row){ return row('maxStrength').gt(700000).and(row('species').contains('Saiyan')); }) .orderBy(r.desc('maxStrength')); presumably returns one row per character, whereas SELECT c.* FROM characters c INNER JOIN character_species cs ON c.id = cs.character_id INNER JOIN species s ON cs.species_id = s.id WHERE max_strength > 700000 AND s.name = 'Saiyan' ORDER BY max_strength DESC; returns at least one row per character. Rather, the equivalent SQL (depending on what syntax is available) might be: SELECT c.* FROM characters c WHERE max_strength > 700000 AND EXISTS (SELECT NULL FROM character_species cs INNER JOIN species s ON cs.species_id = s.id WHERE c.id = cs.character_id AND s.name = 'Saiyan') ORDER BY max_strength DESC; The above returns one row per character, no matter how many matching entries there are in `character_species` or `species`. Pursuing this a little further then, suppose you wanted to expand your search to include humans and androids. In SQL, only a slight modification is needed: SELECT c.* FROM characters c WHERE max_strength > 700000 AND EXISTS (SELECT NULL FROM character_species cs INNER JOIN species s ON cs.species_id = s.id WHERE c.id = cs.character_id AND s.name IN ('Saiyan', 'Human', 'Android')) ORDER BY max_strength DESC; I shudder to think what that would look like in functional form. Something like this? r .db('dragonball') .table('characters') .filter(function(row){ return row('maxStrength').gt(700000).and(row('species').contains('Saiyan')).or(row('species').contains('Human')).or(row('species').contains('Android')); }) .orderBy(r.desc('maxStrength')); But how would it know how to group the disjunctions? In my very limited experience, examples like these always look so tantalizing until you start taking them a little further and then you quickly realize there's no such thing as a free lunch. In this case, you pay for the flexibility of RethinkDB with the expressiveness SQL.
- atnnn 12y agoThe RethinkDB query language can be very expressive. What you shudder to think of could be expressed as row('species').contains(function(s){ r.expr(['Saiyan', 'Human', 'Android']).contains(s) }) Or also as: row('species').setIntersection(['Saiyan', 'Human', 'Android']).isEmpty().not() The succintness of ReQL depends on that of the host language. JavaScript's bulky syntax for functions and lack of operator overloading make queries a lot more cluttered.