3 ms·
I'm not the OP, but I can report on how we evolved our BI analytics at Appcues. We started out writing events into a Postgres table, but BI queries were slow a
by gamache 7y ago
I'm not the OP, but I can report on how we evolved our BI analytics at Appcues.
We started out writing events into a Postgres table, but BI queries were slow and Postgres was an expensive place to put the events.
So then we started writing the events into S3 in batches of 10,000. That number was chosen semi-arbitrarily, intending that any Lambda function would be able to process an entire batch within the 5 minute execution limit. We started also keeping aggregate stats on this data by updating counters and HyperLogLog estimators in a Redis store, updated as each batch hits S3. Athena became our BI tool.
About a year after that, we'd evolved the system so that we were splitting events by customer (instead of being an arbitrary time-slice of a day's traffic). This became a suitable backend for a customer CSV export system, as well as making BI cheaper by allowing us to zero in on the customer(s) we were curious about.
And in the last year, we've begun to use the batched event data in S3 to feed a Snowflake DB (via Snowpipe), which we use for both offline BI and online analytics as part of our product. Snowflake is not free and requires some sophistication, but it supports the leading analytics tools and visualizers, and it's part of a direct evolution from keeping JSON files on S3.