4 ms·
The core looks close enough to dataframes that I'd be curious to know how you compare to tablesaw: https://github.com/jtablesaw/tablesaw https://github.com/jtab
by agibsonccc 9y ago
The core looks close enough to dataframes that I'd be curious to know how you compare to tablesaw:
https://github.com/jtablesaw/tablesaw https://github.com/jtablesaw/tablesaw
This looks neat but I'm not sure why I would care about this. There's a ton of solutions out there in the ecosystem out there already with a columnar like interface.
Granted, we wrote our own as well[1] that uses the builder pattern that you then toss to an executor (our main backend is spark for this). One reason we wrote this is for persistence purposes. Being able to encode and persist a series of transforms that you can then load remotely has been very helpful for us in machine learning.
We've since migrated this project to the eclipse foundation and intend on doing a rewrite of the interface as well as integrate our baked in tensor library[2] in to certain parts of the pipeline for speed purposes and handling things like computer vision workloads.
In general, I always like seeing new takes on the columnar format processing approach but I'm just not seeing anything novel here. Clarification of intent would be great!
[1]: https://github.com/deeplearning4j/DataVec https://github.com/deeplearning4j/DataVec
[2]: https://github.com/deeplearning4j/nd4j https://github.com/deeplearning4j/nd4j
- asavinov 9y agoBistro is not about a physical model and columnar (physical) representation although it relies on it. It is about logical level of representation and processing. There are at least three major questions when we want to introduce a new logical data model: * How we define columns within one table. Conceptually, it is easy, e.g., SELECT x, y, c = a+b FROM T. Yet, even in this simple case we see a controversy: this statement will create a table but our goal is not to create a table - we want to create a column (function). Bistro uses calc operation for that purpose. But it is of course not new. In pandas, for example, one can use df.apply. * How to connect several tables. RM, map-reduce and other set-oriented approaches use join which produces a new table. Here we have a similar controversy [1]: I do not want to produce a new table, it is not my goal. My goal is to link these two tables and this means creating a new column. Bistro introduces link operation for such columns. * How to aggregate data. Here a typical operation is group-by. It is an eclectic operation which combines several other operations and such an approach also has some problems [2]. Bistro changes the way data is aggregated by introducing accumulate functions which get one input (not a subset) and return one output value. An accumulate function is called for each element of the group by updating the current aggregate instead of computing the aggregate by getting the whole group. So linking and aggregation are what distinguishing Bistro from other approaches and frameworks to data processing including pandas, SQL and map-reduce. [1] https://www.researchgate.net/publication/301764816_Joins_vs_Links_or_Relational_Join_Considered_Harmful https://www.researchgate.net/publication/301764816_Joins_vs_... [2] https://www.researchgate.net/publication/316551218_From_Group-by_to_Accumulation_Data_Aggregation_Revisited https://www.researchgate.net/publication/316551218_From_Grou...
- agibsonccc 9y agoOk so you're defining new operations on top of existing primitives. Makes sense! The concepts look more interesting than the library right now (it doesn't move the needle for me production wise yet) but it has potential! Every project starts somewhere. I'm glad you wrote this in java at least. There's a ton of things I'd be missing to start looking at this. 1. Backend agnostic: let me run on different backends like flink/spark 2. Give me off heap memory please. Let me play dangerous and use pointers directly to optimize interactions with transforms. We wrote our own GC among other things for our tensor lib due to the GC bottlenecks and copying and the like. That's personally what I kind of like about tablesaw. A library for 1 off adhoc analysis in memory isn't a bad start though, especially since most folks don't actually have that large of problems.
- asavinov 9y agoYou are right - it is an MVP, and the goal is to choose a direction. In fact, I am still not sure in which direction to go: * JavaScript data processing framework for in-browser data processing * Python framework like pandas * Big data processing framework like Spark * Database management system * Data integration system like typical ETL and BI * Stream analytics like Kafka Streams * IoT (light weigh) stream processing engine * Something else? I would be very thankful for any suggestion from people who know the market and (acute) needs of the customers. What is the best niche for this kind of technology?
- agibsonccc 9y agoThe arrow integration and persistence and connecting this to jdbc would be an MVP to me that would at least allow folks to imitate pandas. While you do have new ideas here if folks can express these ideas in terms of ways they're familiar with that would likely help. Maybe you could use calcite[1] as an engine or as a base kind of like arrow. If you do python, maybe look in pyjnius[2]. There's a lot of things that already have connectors. I would continue along your MVP route allowing folks to do basic things with your framework first then you can improve it as you go. SQL databases aren't a bad initial target. Most folks can do SQL. [1]: https://calcite.apache.org/ https://calcite.apache.org/ [2]: https://github.com/kivy/pyjnius https://github.com/kivy/pyjnius