7 ms·
> but every time I try to use it, I get stuck on "how do I get data into it reliably" That's the same stage I get stuck every time. I have data emitters (in t
by pbowyer 2y ago
> but every time I try to use it, I get stuck on "how do I get data into it reliably"
That's the same stage I get stuck every time.
I have data emitters (in this example let's say my household IoT devices, feeding a MQTT broker then HomeAssistant).
I have where I want the data to end up (Clickhouse, Database, S3, whatever).
How do I get the data from A to B, so there are no duplicate rows (if the ACK for an upload isn't received when the upload succeeded), no missing rows (the data is retried if an upload fails), and some protection if the local system goes down (data isn't ephemeral)?
The easiest I've found is writing data locally to files (JSON, parquet, whatever), new file every 5 minutes and sync the older files to S3.
But then I'm stuck again. How do I continually load new files from S3 without any repetition or edge cases? And did I really need the intermediate files?
- wiredfool 2y agoEasiest way is to post csv/json/whatever through the http endpoint into a replacing merge tree table. Duplicates get merged out, and errors can be handles at the http level. (Admittedly, one bad row in a big batch post is a pain, but I don’t see that much)
- Narhem 2y agoHTTP errors aren’t the most readable, although traditional database errors aren’t too readable most of the time.
- wiredfool 2y agoWhat I meant is that you'll get an HTTP error code from the insert if it didn't work, so that can go through the error handling. This isn't really an "explore this thing", it's a "splat this data in, every minute/file/whatever". I've churned through TBs of CSVs this way, with a small preprocessor to fix some idiosyncratic formatting.
- maccard 2y agoThis is _exactly_ my problem, and where I've found myself.
- zbentley 2y agoThis isn't appropriate for all use-cases, but one way to address your and GP's problem is as follows: 1. Aggregate (in-memory or on cheap storage) events in the publisher application into batches. 2. Ship those batches to S3/alike, NFS that clickhouse can read, or equivalent (even a dead-simple HTTP server that just receives file POSTs and writes them to disk, running on storage that clickhouse can reach). The tool you use here needs to be idempotent (retries of failed/timed out uploads don't mangle data), and atomic to readers (partially-received data is never readable). 3. In ClickHouse, run a scheduled refresh of a materialized view pointed at the uploaded data (either "SELECT ... INFILE" for local/NFS files, or "INSERT INTO ... SELECT s3(...)" for an S3/alike): https://clickhouse.com/docs/en/materialized-view/refreshable-materialized-view https://clickhouse.com/docs/en/materialized-view/refreshable... This is only a valid solution given specific constraints; if you don't match these, it may not work for you: 1. You have to be OK with the "experimental" status of refreshable materialized views. My and other users' experience with the feature seems generally positive at this point, and it has been out for awhile. 2. Given your emitter data rates, there must exist a batch size of data which appropriately balances keeping up with uploads to your blob store and the potential of data loss if an emitter crashes before a batch is shipped. If you're sending e.g. financial transaction source-of-record data, then this will not work for you: you really do need a Kafka/alike in that case (if you end up here, consider WarpStream: an extremely affordable and low-infrastructure Kafka clone backed by batching accumulators in front of S3: https://www.warpstream.com/ https://www.warpstream.com/ If their status as a SaaS or recent Confluent acquisition turns you off, fair enough.) 3. Data staleness of up to emitter-flush-interval + worst-case-upload-time + materialized-view-refresh-interval must be acceptable to you. 4. Reliability wise, the staging area for shipped batches (S3, NFS, scratch directory on a clickhouse server) must be sufficiently reliable for your use case, as data will not be replicated by clickhouse while it's staged. 5. All uniqueness/transformations must be things you can express in your materialized view's query + engine settings.
- 2y ago
- masterj 2y agoCloudflare workers combined with their queues product https://developers.cloudflare.com/queues/ https://developers.cloudflare.com/queues/ might be a cheap and easy way of solving this problem