4 ms·
Ok so you're defining new operations on top of existing primitives. Makes sense! The concepts look more interesting than the library right now (it doesn't mov
by agibsonccc 9y ago
Ok so you're defining new operations on top of existing primitives. Makes sense! The concepts look more interesting than the library right now (it doesn't move the needle for me production wise yet) but it has potential! Every project starts somewhere. I'm glad you wrote this in java at least.
There's a ton of things I'd be missing to start looking at this.
1. Backend agnostic: let me run on different backends like flink/spark
2. Give me off heap memory please. Let me play dangerous and use pointers directly to optimize interactions with transforms.
We wrote our own GC among other things for our tensor lib due to the GC bottlenecks and copying and the like.
That's personally what I kind of like about tablesaw.
A library for 1 off adhoc analysis in memory isn't a bad start though, especially since most folks don't actually have that large of problems.
- asavinov 9y agoYou are right - it is an MVP, and the goal is to choose a direction. In fact, I am still not sure in which direction to go: * JavaScript data processing framework for in-browser data processing * Python framework like pandas * Big data processing framework like Spark * Database management system * Data integration system like typical ETL and BI * Stream analytics like Kafka Streams * IoT (light weigh) stream processing engine * Something else? I would be very thankful for any suggestion from people who know the market and (acute) needs of the customers. What is the best niche for this kind of technology?
- agibsonccc 9y agoThe arrow integration and persistence and connecting this to jdbc would be an MVP to me that would at least allow folks to imitate pandas. While you do have new ideas here if folks can express these ideas in terms of ways they're familiar with that would likely help. Maybe you could use calcite[1] as an engine or as a base kind of like arrow. If you do python, maybe look in pyjnius[2]. There's a lot of things that already have connectors. I would continue along your MVP route allowing folks to do basic things with your framework first then you can improve it as you go. SQL databases aren't a bad initial target. Most folks can do SQL. [1]: https://calcite.apache.org/ https://calcite.apache.org/ [2]: https://github.com/kivy/pyjnius https://github.com/kivy/pyjnius
- munro 9y agoI'm working on a columnar Python DSL right now, I think of it like SQLAlchemy for Pandas/Spark/Flink. My goal is to create a language that makes the cumbersome parts of the PySpark API much easier to express. I started off with the intent of replicating the R's dataframe API because it feels more fluid--but what you're doing feels eerie, because I came to the same conclusions as you around focusing on a language for columnar manipulations, and letting the "linking" become implicit. Then I want to transition to rethinking the ML pipeline, I like Spark's more than sklearn's, but there are still cumbersome parts that my intuition says can be solved by a columnar API. So if I were you, I would do Python DSL, because it addresses my immediate needs of increasing productivity. :)