5 ms·
MillenniumDB: Property graph and RDF engine, still in development
- deleted 2y ago[deleted]
- UltraSane 2y agoWhat is a domain graph?
- leetrout 2y agoWeird title here. The repo says "Property Graph and RDF engine, still in development" with no mention of domain.
- dang 2y agoWe've changed the title to that of the page. (Submitted title was "MillenniumDB: A graph database engine using domain graphs")
- throwaway867500 2y ago"Domain Graph" [1] was renamed to "Multilayer Graphs" [2]. The "Multilayer Graphs Model" was aimed to address the limitations found in prior Graph Models (e.g. RDF, Property Graphs) in representing higher-arity graphs without having to resort to reification or reserved words/vocab. Skipping over the formal math definitions, a Multi-Layer Graph, in practical terms, is represented by statements of "quads": "{edge id, source, label, target}" -- similar on the surface to RDF-named graphs, but not. The source and target may also refer other edges in addition to referring to the "entities" of real-world-concepts; targets may also be data types (e.g. strings, ints, etc)-- I believe that MilleniumDB puts edges, sources, targets and some simple values under the same packed int64 namespace. Useful for MilleniumDB under-the-hood design/architecture, and probably for Wikidata where their data model are qualified facts-- a qualifying statement (e.g. "valid from 2020 to 2024") about a factual statement ("Alice Lived in America"). But this is just me, a non-expert, trying to cut to the core points after discovering and reading the papers just recently since knowledge graphs are hot topics right now. [1] https://arxiv.org/abs/2111.01540 https://arxiv.org/abs/2111.01540 [2] https://users.dcc.uchile.cl/~ahogan/docs/mutlilayer_graphs.pdf https://users.dcc.uchile.cl/~ahogan/docs/mutlilayer_graphs.p...
- UltraSane 2y agoallowing edges to have edges is something that RDF* allows. Property graphs DBs like Neo4j don't support it but you can do it by using a node as a relationship. This is called a metanode or a hypernode. The need for this is mitigated somewhat by the fact that property graphs allow edges to have properties themselves. so you would use (Alice)-[:LIVES_IN{valid_from:2020,valid_to:2024}]-(USA) Edit: Read the paper and it is actually an attempt to unify the RDF, RDF* and property graph data models which is VERY interesting.
- jitl 2y agoIs it any good?
- WhatIsDukkha 2y agoHere is a bug with some back and forth between millenniumdb and qlever in starting a benchmarking attempt but I don't see results, though they managed to build and import. https://github.com/MillenniumDB/MillenniumDB/issues/10 https://github.com/MillenniumDB/MillenniumDB/issues/10 https://github.com/ad-freiburg/qlever https://github.com/ad-freiburg/qlever
- deleted 2y ago[deleted]
- smarx007 2y agoI think if someone is just trying out RDF, it is better to start with Apache Jena/Fuseki or Eclipse RDF4J. Maybe https://github.com/oxigraph/oxigraph https://github.com/oxigraph/oxigraph if you like to live dangerously (i.e. to use pre-1.0 DBMSs). Use of other systems involves factoring tradeoffs and considerations that are probably not the best for the newcomers. For example, qLever mentioned here is good in query performance and relative disk use but once the import is done, it's essentially a read-only DB and completely unsuitable for a typical OLTP scenario. Having said that, the Chilean research group that is driving the development of MilleniumDB is very well-regarded in the RDF/semantic web querying space.
- FjordWarden 2y agoIf you expect Jena to be more battle-tested because it is older, forget it, if the process is killed by a unexpected shutdown or some other reason it results in data corruption. At least this was my experience a few years ago. I found graph databases a beguiling idea when I first learned about them, and this is a welcome addition, but I've since temperated my excitement. They are not as flexible and universal a modal as is often promised. Everything is a graph, sure but the result of your SPARQL query not necessarily. I found classical DBMS based on sets/multisets to be much easier to compose from a querying point of view. A table is a set/multiset and a result of a query is also a set/multiset, SPARQL guarantees no such composability. Maybe, if you want to start mucking around with inference engines, but you'll either run into problems of undecidability.
- zozbot234 2y ago> SPARQL guarantees no such composability. SPARQL has a CONSTRUCT clause which gives you RDF as your query output. Isn't that compositional enough?
- FjordWarden 2y agoOk, that is true, but how do I tell my graph database that the result of the construct query is some other graph in my DB?
- jerven 2y agoMilleniumDB is an interesting engine, as is Qlever mentioned in other comments. I think both are good candidates at making RDF graphs one or two orders of magnitude cheaper to host as sparql endpoints. Both seem to have arrived at the stage of transitioning from research to production code. Very exiting for those of us providing our data in RDF and exposing Sparql. AWs Neptune analytics is also very interesting, allowing Cypher on RDF graphs. Even the Oracle inbuilt RDF+Sparql seems to have improved greatly in 23ai.
- UltraSane 2y agoIt seems like writing Cypher to query RDF would be hard.
- j-pb 2y agoThese guys write really great papers! We implemented a simplified version of their ring index for our data space (https://github.com/triblespace/tribles-rust/blob/master/src/blob/schemas/succinctarchive.rs https://github.com/triblespace/tribles-rust/blob/master/src/...), and it's a really simple and cool idea once you wrap your head around it. Funnily enough, we build this even before the paper was officially published, because we found a preprint on one of the authors blogs. The idea itself was published by them before but their new paper made this a lot easier to understand. (burrows wheeler transforms vs. stable column sorting). It's really too bad that the whole linked-data space is completely gunked up with RDF. Ps: If anyone plans on implementing their ring index, using 0 based offsets makes the formulas much more streamlined, their paper uses 1 based indexing and they have to +/-1 all over the place.
- bawolff 2y agoThe whole ring index thing is one of the more fascinating ideas i've read about (i didnt realize milleniumDB was same authors). Sent me down a whole rabbit hole of learning about succinct data structures and burrows-wheeler transform. Sometimes you encounter a computer science idea that just sounds like pure magic.
- sunshine-o 2y agoI got very interested in RDF about 20-25 years ago. Obviously it did not really succeeded but it seems some industries invested a lot into the tech and it is still around. Especially since AWS built a service around it. I am really curious, what are the top use cases for it today?
- bawolff 2y agoI think wikidata (https://query.wikidata.org https://query.wikidata.org) is one of the more well known ones.