3 ms·
If you let the DB do the joins, it could handle more efficiently. For example, it could distribute the joins to those 30 partitions of the main table, and then
by boomzilla 13y ago
If you let the DB do the joins, it could handle more efficiently. For example, it could distribute the joins to those 30 partitions of the main table, and then merge the results, so the heavy computation is distributed, and less bits to move around the network.
Now in the cases where if you can optimize the joins, you still have the option of doing it in your code in RethinkDB/CouchDB. I've done that too, and it's usually when I know for sure that I can prune a big collection to a very small subset more efficiently than using an index.
I would still argue that client app is not the right level of abstraction for data join though, unless it is a big performance gain for very little extra complication.
- coffeemug 13y agoSlava @ rethink here. I'd be curious to see a use case where doing the join on the client is more efficient than doing it on the server. I can't think of a single one off the top of my head.
- SamReidHughes 13y agoIf the join is "compute all f(s, t) for s in S and t in T" then you'll save bandwidth (having O(|S| + |T|) bandwidth over the network instead of O(|S||T|)) by doing it on the client. Of course you could just run `rethinkdb proxy` if you want to save ethernet bandwidth and run the query on RethinkDB while connecting to the local cluster node.
- coffeemug 13y agoAh, I see, that makes sense. Usually pulling both tables in full (or even a single table in full) to the client is not an option (and nobody does cross products in real-time systems). So people end up pulling a subset of table A, and then for each document in the subset issue a separate get to the db for table B (which is obviously worse than having the db do it).