3 ms·
Thanks for all the questions @eloff! I might try to separate responses for hopefully some threading :) The software in our architecture is actually unified:
by mfreed 6y ago
Thanks for all the questions @eloff! I might try to separate responses for hopefully some threading :)
The software in our architecture is actually unified: You spin up K instances of TimescaleDB, log into one of them and then call "add_data_node" to attach others and data nodes, and "create_distributed_hypertable" to make this node the access node _for that particular hypertable_.
In fact, you can explicitly specify different data nodes as the "backend" for different distributed hypertables, e.g., same access node, but data nodes {A, B} for distributed hypertable 1; data nodes {C, D, E, F} are for distributed hypertable 2. You could also have different nodes play the role of access nodes for different hypertables and vice-versa. Lots of flexibility here, although we imagine most users won't need it, so by default your distributed hypertable is created across all connected data nodes.
Now, because these nodes are just running TimescaleDB, if you were to create a regular table or even a regular (non-distributed) hypertable on that access node, it would be just stored locally. Then, any JOIN between the distributed hypertable data and regular table happens transparently on the access node.
Now, that's what in the current release, but not surprisingly, we plan to add support for better local JOINs. For regular tables, one approach will likely be to provide an option to replicate the table on all data nodes, so that you can perform the JOINs locally (and because TimescaleDB 2.0 already supports 2PC for individual rows across data nodes to support replication factors > 1, that machinery somewhat already exists). A second approach for regular tables that are also partitioned -- particularly if using the same "space" partitioning key as your distributed hypertable -- is to collocate partitions of the regular table across the data nodes, so that JOINs on the space partition keys (e.g., server_id, user_id, etc.) can again be local.
- eloff 6y agoOk, I think I understand. Each node is just Postgres with TimescaleDB, and acts as a separate Postgres server with it's own replication (if any) - at least at present. So your "fact" regular tables are created on the server you configure as the access node and then you can join with distributed hypertables in your queries. If one wanted the same fact tables on multiple access nodes you could either use postgres partitioning or duplicate them somehow yourself. I like the flexibility here in cluster configuration. It's good to have options when scaling a system.
- mfreed 6y agoYes, with the observation is that this is the base flexibility, while we will be adding much more transparent stuff in the future. So in 2.0, if you want a chunk of a distributed hypertable replicated, you don't do it manually, but just configure the distributed hypertable to use "replication_factor = 2", etc. Native replication (and the approaches I described for regular "fact" tables as well) are major focuses of the 2.x roadmap.