4 ms·
I have these use cases: 1. Syncing data in Postgres to ElasticSearch/ClickHouse (which handle search/analytics on the data I store in PG) 2. Invoking my own w
by iyn 2y ago
I have these use cases:
1. Syncing data in Postgres to ElasticSearch/ClickHouse (which handle search/analytics on the data I store in PG)
2. Invoking my own workflow engine — I have built a system that allows end-users to define triggers that start workflows when some data change/is created. To determine whether I need to start the workflow, I need to inspect every CRUD operation and check it against triggers defined by the users.
I'm currently doing that in a duck-tape like way by publishing to SNS from my controllers and having SQS subscribers (mapped to Lambdas) that are responsible for different parts of my "pipeline". I don't like this system as it's not fault-tolerant and I'd prefer to do this by processing WAL and explicitly acknowledging processed changes.
- shayonj 2y agoI think PeerDB is a smart choice if you are looking to sync the data to Clickhouse. They have nice integration support as well. re: 2. I have some plans around a control plane that allows users to define these transformation rules and routing config and then take further actions based on the outcomes. If you are interested in it, feel free to sign up on the homepage. Also, very happy for some quick chats too (shayonj at gmail). Thanks
- iyn 2y agoYeah, I've looked into PeerDB but in terms of self-hosting it's not really lightweight, as they depend on Temporal [0]. I'm currently optimizing for less complexity/budget, as I have just a few customers. [0] https://docs.peerdb.io/architecture#dependencies https://docs.peerdb.io/architecture#dependencies
- shayonj 2y agoYeah, totally fair. Are you ok with a NATS dependency ? Happy to work with you in supporting a new destination like ES. Also looking to make NATS optional for smaller/simpler setups (https://github.com/shayonj/pg_flo/issues/21 https://github.com/shayonj/pg_flo/issues/21)
- iyn 2y agoYes, I think NATS is reasonable — I don't have operational experience with it but based on my earlier reading it seems that it can be run on a smaller budget. Is this "regular" NATS or the Jetstream variant?
- shayonj 2y agoPerf! From testing on some of my staging workloads, the footprint isn't too high and I can get 5-6k messages/s. Esp. since there is only one worker instance involved (for strict ordering). Yes, it does use NATS JetStream.
- saisrirampur 2y agoSai from PeerDB here. Temporal has been very impactful for us and a major factor in our ability to build a production-grade product that supports large-scale workloads. At a high level, CDC is a complex state machine. Temporal helps building the state machine taking care of auto-retries/idempotency at different failure points and also aids in managing and observing it. This is very useful to identify root causes when issues arise. Managing Temporal shouldn’t be complex. They offer a well-maintained, mature Docker container. From a user standpoint, the software is intuitive and easy to understand. We package the temporal docker container in our own Docker setup and have it integrated into our Helm charts. We’ve quite a few users smoothly using Enterprise (that we open sourced recently) and standard OSS! https://github.com/PeerDB-io/peerdb/blob/main/docker-compose.yml https://github.com/PeerDB-io/peerdb/blob/main/docker-compose... https://github.com/PeerDB-io/peerdb-enterprise https://github.com/PeerDB-io/peerdb-enterprise Let me know if there are any questions!
- merb 2y ago1. i am not sure if the helm chart can be used for the oss version? 2. if a helm chart needs sh files, it’s already an absolut no-go since it won’t work with gitops that well.
- assaxor 2y agoHi, the helm chart uses the OSS PeerDB images. The sh files were created to bootstrap the values files for easier (and faster) POCs. You can append a `template` argument when running the script files which will lead to a set of values file being generated, which you can then modify accordingly. There is a production guide for the same as we have customers in production using GitOps (ArgoCD) with the provided charts (https://github.com/PeerDB-io/peerdb-enterprise/blob/main/PRODUCTION.md https://github.com/PeerDB-io/peerdb-enterprise/blob/main/PRO...)
- iyn 2y agoThanks for reaching out! Just to be clear, from what I can tell, both PeerDB and Temporal are great (and I’ve been hoping to learn Temporal for a while). At some point I considered self-hosting PeerDB but my impression was that it required multiple nodes to run properly and so it wasn’t budget friendly - this is also based on your pricing plans with $250 being the cheapest which suggests that it’s not cheap to host it (I’m trying to minimize costs until I have more customers). Please correct me if I’m wrong! Can you give me an example of a budget friendly deployment, e.g. how many EC2 instances for PeerDB would I need for one of the smaller RDS instances? Given the acquisition by ClickHouse (congrats!), what can we expect for the CDC for sinks other than CH? Do you plan to continue supporting different targets or should we expect only CH focus? Edit: also, any plans for supporting e.g. SNS/SQS/NATS or similar?