5 ms·
Bistro is not about a physical model and columnar (physical) representation although it relies on it. It is about logical level of representation and processing
by asavinov 9y ago
Bistro is not about a physical model and columnar (physical) representation although it relies on it. It is about logical level of representation and processing.
There are at least three major questions when we want to introduce a new logical data model:
* How we define columns within one table. Conceptually, it is easy, e.g., SELECT x, y, c = a+b FROM T. Yet, even in this simple case we see a controversy: this statement will create a table but our goal is not to create a table - we want to create a column (function). Bistro uses calc operation for that purpose. But it is of course not new. In pandas, for example, one can use df.apply.
* How to connect several tables. RM, map-reduce and other set-oriented approaches use join which produces a new table. Here we have a similar controversy [1]: I do not want to produce a new table, it is not my goal. My goal is to link these two tables and this means creating a new column. Bistro introduces link operation for such columns.
* How to aggregate data. Here a typical operation is group-by. It is an eclectic operation which combines several other operations and such an approach also has some problems [2]. Bistro changes the way data is aggregated by introducing accumulate functions which get one input (not a subset) and return one output value. An accumulate function is called for each element of the group by updating the current aggregate instead of computing the aggregate by getting the whole group.
So linking and aggregation are what distinguishing Bistro from other approaches and frameworks to data processing including pandas, SQL and map-reduce.
[1] https://www.researchgate.net/publication/301764816_Joins_vs_Links_or_Relational_Join_Considered_Harmful https://www.researchgate.net/publication/301764816_Joins_vs_...
[2] https://www.researchgate.net/publication/316551218_From_Group-by_to_Accumulation_Data_Aggregation_Revisited https://www.researchgate.net/publication/316551218_From_Grou...
- agibsonccc 9y agoOk so you're defining new operations on top of existing primitives. Makes sense! The concepts look more interesting than the library right now (it doesn't move the needle for me production wise yet) but it has potential! Every project starts somewhere. I'm glad you wrote this in java at least. There's a ton of things I'd be missing to start looking at this. 1. Backend agnostic: let me run on different backends like flink/spark 2. Give me off heap memory please. Let me play dangerous and use pointers directly to optimize interactions with transforms. We wrote our own GC among other things for our tensor lib due to the GC bottlenecks and copying and the like. That's personally what I kind of like about tablesaw. A library for 1 off adhoc analysis in memory isn't a bad start though, especially since most folks don't actually have that large of problems.
- asavinov 9y agoYou are right - it is an MVP, and the goal is to choose a direction. In fact, I am still not sure in which direction to go: * JavaScript data processing framework for in-browser data processing * Python framework like pandas * Big data processing framework like Spark * Database management system * Data integration system like typical ETL and BI * Stream analytics like Kafka Streams * IoT (light weigh) stream processing engine * Something else? I would be very thankful for any suggestion from people who know the market and (acute) needs of the customers. What is the best niche for this kind of technology?
- agibsonccc 9y agoThe arrow integration and persistence and connecting this to jdbc would be an MVP to me that would at least allow folks to imitate pandas. While you do have new ideas here if folks can express these ideas in terms of ways they're familiar with that would likely help. Maybe you could use calcite[1] as an engine or as a base kind of like arrow. If you do python, maybe look in pyjnius[2]. There's a lot of things that already have connectors. I would continue along your MVP route allowing folks to do basic things with your framework first then you can improve it as you go. SQL databases aren't a bad initial target. Most folks can do SQL. [1]: https://calcite.apache.org/ https://calcite.apache.org/ [2]: https://github.com/kivy/pyjnius https://github.com/kivy/pyjnius
- munro 9y agoI'm working on a columnar Python DSL right now, I think of it like SQLAlchemy for Pandas/Spark/Flink. My goal is to create a language that makes the cumbersome parts of the PySpark API much easier to express. I started off with the intent of replicating the R's dataframe API because it feels more fluid--but what you're doing feels eerie, because I came to the same conclusions as you around focusing on a language for columnar manipulations, and letting the "linking" become implicit. Then I want to transition to rethinking the ML pipeline, I like Spark's more than sklearn's, but there are still cumbersome parts that my intuition says can be solved by a columnar API. So if I were you, I would do Python DSL, because it addresses my immediate needs of increasing productivity. :)