3 ms·
You are right - it is an MVP, and the goal is to choose a direction. In fact, I am still not sure in which direction to go: * JavaScript data processing framew
by asavinov 9y ago
You are right - it is an MVP, and the goal is to choose a direction. In fact, I am still not sure in which direction to go:
* JavaScript data processing framework for in-browser data processing
* Python framework like pandas
* Big data processing framework like Spark
* Database management system
* Data integration system like typical ETL and BI
* Stream analytics like Kafka Streams
* IoT (light weigh) stream processing engine
* Something else?
I would be very thankful for any suggestion from people who know the market and (acute) needs of the customers. What is the best niche for this kind of technology?
- agibsonccc 9y agoThe arrow integration and persistence and connecting this to jdbc would be an MVP to me that would at least allow folks to imitate pandas. While you do have new ideas here if folks can express these ideas in terms of ways they're familiar with that would likely help. Maybe you could use calcite[1] as an engine or as a base kind of like arrow. If you do python, maybe look in pyjnius[2]. There's a lot of things that already have connectors. I would continue along your MVP route allowing folks to do basic things with your framework first then you can improve it as you go. SQL databases aren't a bad initial target. Most folks can do SQL. [1]: https://calcite.apache.org/ https://calcite.apache.org/ [2]: https://github.com/kivy/pyjnius https://github.com/kivy/pyjnius
- munro 9y agoI'm working on a columnar Python DSL right now, I think of it like SQLAlchemy for Pandas/Spark/Flink. My goal is to create a language that makes the cumbersome parts of the PySpark API much easier to express. I started off with the intent of replicating the R's dataframe API because it feels more fluid--but what you're doing feels eerie, because I came to the same conclusions as you around focusing on a language for columnar manipulations, and letting the "linking" become implicit. Then I want to transition to rethinking the ML pipeline, I like Spark's more than sklearn's, but there are still cumbersome parts that my intuition says can be solved by a columnar API. So if I were you, I would do Python DSL, because it addresses my immediate needs of increasing productivity. :)