6 ms·
Three words: JOIN, JOIN, JOIN. For any non-trivial domain, you'll need to join different entity types. For example, a collection of users is fine for user names
by boomzilla 13y ago
Three words: JOIN, JOIN, JOIN. For any non-trivial domain, you'll need to join different entity types. For example, a collection of users is fine for user names, (hashed) passwords, last logins, etc. A collection of document is fine for text, fonts, URL. Now what happens if you have multiple users collaborating on multiple documents? Are you going to denormalize the shared documents to user objects, or are you only keeping document IDs in them? If you choose the former, you'll run into inconsistency very quickly. If you choose the latter, you'll need join support, or you have to write the code to do the poor man's join which could be very inefficient.
RethinkDB supports join out of the box.
- octix 13y agoSo, if I have 30 nodes and my data is sharded between them, how is that better than having app/client side join, where I have more control what to fetch and what not? I could potentially cache results. I mean, you realize if data is distributed at large scale, it may take a while till it gets from all nodes the data and joins it...
- boomzilla 13y agoIf you let the DB do the joins, it could handle more efficiently. For example, it could distribute the joins to those 30 partitions of the main table, and then merge the results, so the heavy computation is distributed, and less bits to move around the network. Now in the cases where if you can optimize the joins, you still have the option of doing it in your code in RethinkDB/CouchDB. I've done that too, and it's usually when I know for sure that I can prune a big collection to a very small subset more efficiently than using an index. I would still argue that client app is not the right level of abstraction for data join though, unless it is a big performance gain for very little extra complication.
- coffeemug 13y agoSlava @ rethink here. I'd be curious to see a use case where doing the join on the client is more efficient than doing it on the server. I can't think of a single one off the top of my head.
- SamReidHughes 13y agoIf the join is "compute all f(s, t) for s in S and t in T" then you'll save bandwidth (having O(|S| + |T|) bandwidth over the network instead of O(|S||T|)) by doing it on the client. Of course you could just run `rethinkdb proxy` if you want to save ethernet bandwidth and run the query on RethinkDB while connecting to the local cluster node.
- coffeemug 13y agoAh, I see, that makes sense. Usually pulling both tables in full (or even a single table in full) to the client is not an option (and nobody does cross products in real-time systems). So people end up pulling a subset of table A, and then for each document in the subset issue a separate get to the db for table B (which is obviously worse than having the db do it).
- fourstar 13y agoI actually have been running into this a lot with my documents in my project. Thinking about making the switch to Rethink. The only thing I'm bummed about is having to migrate all my Mongo code over. Anyone here end up doing something similar?