5 ms·
AFAIK Druid is for time series and it's a columnar database. Their format has dimensions and then pre-aggregated fields on those dimensions. Afterwards you can
by dialtone 10y ago
AFAIK Druid is for time series and it's a columnar database. Their format has dimensions and then pre-aggregated fields on those dimensions. Afterwards you can run SQL queries on that data format to get full aggregations in return.
This is a library to read and write a data format that is optimized to give access to granular events and actors within an event stream. For example this could be used to trail all of the events generated by one entity (credit card, cookie, email, account and so on) over a dataset. At that point you can choose what you want to do with it: extract features for ML, train ML directly on raw data, run arbitrary queries for outliers and anomaly detection and what have you.