2 ms·
At the physical (storage) level it is in-memory and organized as a column store, that is, internally it stores a list of tables (with no data) and a list of col
by asavinov 9y ago
At the physical (storage) level it is in-memory and organized as a column store, that is, internally it stores a list of tables (with no data) and a list of columns (each being a Java array).
It is definitely interesting and important to implement persistence (for example using Parquet or Arrow) as well as other mechanisms like sharding or replication (for big data processing, fault tolerance etc.) Yet, currently this direction has lower priority because the next task I want to focus on is in-stream analytics (an alternative to Kafka Streams).
In general, the whole approach is focused on the logical level of data modeling and processing, that is, the goal is to increase performance and simplicity of development. The general idea (and hypothesis) is that defining how data is being processed using column operations is easier, more intuitive, less error-prone and easier to maintain than using purely set-operations.
In other words, at logical level, it is an alternative to map-reduce, SQL, pandas and other models and frameworks where set-operations are used to process data.
- jnordwick 9y ago> defining how data is being processed using column operations is easier, more intuitive, less error-prone and easier to maintain than using purely set-operations. And every APL programmer just nodded in agreement.
- PeCaN 9y agoAll 6 of us. :( Really though, I spend the time reading the readme thinking “This looks very cool, shame it doesn't have a nice language to go with it…”
- heavenlyblue 9y agoCan you define the column operations by reducing them to a set of analogical steps performed after the query planner has run in an SQL database?