6 ms·
Great job, XTDB's utilization of Apache Arrow for in-memory columnar data structures is highly appreciated. Polars dataframe library in my benchmark is also bas
by athanassios 3y ago
Great job, XTDB's utilization of Apache Arrow for in-memory columnar data structures is highly appreciated. Polars dataframe library in my benchmark is also based on Apache arrow. Thank you for the reference on XTDB benchmark datasets. How can I import `tsv` and/or n-triplets file into v2 XTDB so I can run a couple of queries on WatDiv ? It worth mentioning here that there is a growing interest on extracting subsets of data from large knowledge triplets bases and there are specialized tools for this (e.g. KGTK). I guess v2 XTDB with datalog will make this process a lot easier.
Well, regarding the "one size fits all" database engine, I would rephrase this as "one size fits the most" :-). Your reference about SingleStore (it has a remarkable performance similar to ClickHouse), Rowstore and Columnstore engines in one DBMS, fits the most if not all the cases. But in my opinion, the ideal solution to the DBMS-database challenge resides at the application level and within the realm of database modeling. I would contend that in today's landscape, it's the (V)ariety dimension of big data, that poses the most significant challenge – specifically, the seamless integration of data from various sources. And of course if you manage to lift the heavy load (ETL), as Stonebraker said, you run into the graph/traversal problem.
If you take into account the extensive and continuous development efforts spanning many years, along with the evolving trends in the realm of database query languages, all aimed at enhancing the often troublesome SQL and its numerous alternatives like SPARQL, graph query languages, noSQL, and newSQL, you'll begin to grasp that there is a fundamental issue underlying all of these approaches.
Primarily, it's the lack of a distinct separation between the logical and computational layer and the storage (physical layer). Only recently, this crucial aspect has begun to take center stage in the development of DBMS systems featuring multiple storage engines. Was it Datomic that played a pioneering role in this shift ? The second equally import aspect that it is still heavily underestimated nowadays is the principle of compositionality both at the level of the query language and the data types. Datalog, especially if it is used in a logic programming environment like Clojure, is in the right direction. I will argue, that several Datalog systems, which exclusively adhere to the EAV/RDF data model, lack a crucial element. They are forcing users to create queries and data models through the cumbersome process of breaking down and building n-ary relations from triplets. I believe that the remedy lies in developing robust transformation functions capable of handling input and output across triplets, tabular, or nested data structures.
PS: About Logica: The only columnar DBMS that is supported is BigQuery and it is available as Google proprietary cloud service. As it concerns in-memory processing I could not find any reference on what kind of data structures it uses.
- refset 3y ago> How can I import `tsv` and/or n-triplets file into v2 XTDB so I can run a couple of queries on WatDiv ? I can't offer easy instructions for that right now as we've not adapted the benchmark for the 2.x branch yet (and I would not expect performance to match 1.x right now anyway), although all the code one might need is in the repo for 1.x already and should be straightforward enough to convert to the new APIs given some familiarity with Clojure: https://github.com/xtdb/xtdb/blob/master/bench/src/xtdb/bench/watdiv_xtdb.clj https://github.com/xtdb/xtdb/blob/master/bench/src/xtdb/benc... > Was it Datomic that played a pioneering role in this shift ? I know Datomic has inspired a lot of people - certainly our entire team - but I'm not sure how much it's influenced the existence of things like Cozo or the wider database ecosystem. I get the feeling that Datomic's impact on application developer communities and library builders has been far greater (also amplified further by DataScript) than the impact on established database vendors & research communities. It has at least been good to see more systems start treating data immutably and exposing time-travel capabilities, even if Datomic was merely ahead of the curve rather than a direct inspiration. > They are forcing users to create queries and data models through the cumbersome process of breaking down and building n-ary relations from triplets. I believe that the remedy lies in developing robust transformation functions capable of handling input and output across triplets, tabular, or nested data structures. Despite originally being attracted to triple-oriented modeling I have come around to agreeing with this point of view. Processing with n-ary relations is unavoidable so the idea of not persisting them also feels inherently limiting. I am also quite drawn to "Object Role Modeling" (though I've not yet applied it in anger) which really demands n-ary tuples, and the RelationalAI folks seem to have reached a similar conclusion: https://docs.relational.ai/rel/concepts/relational-knowledge-graphs/schema-visualization#schema-visualization-orm-and-lpg-diagrams https://docs.relational.ai/rel/concepts/relational-knowledge... > in my opinion, the ideal solution to the DBMS-database challenge resides at the application level and within the realm of database modeling I generally agree with this too. All I know for certain is that almost everything in software should be more relational and declarative, and then machine learning stands the best chance of figuring out how to make things fast. Good discussion :)