4 ms·
The locality that you want for performance is available as row values in sqlite. Sqlite seems to have the parts needed to build a graph db.
by oever 5y ago
The locality that you want for performance is available as row values in sqlite.
Sqlite seems to have the parts needed to build a graph db.
- deleted 5y ago[deleted]
- WJW 5y agoWhy do you keep making the same argument again? Yes, SQLite has the parts needed to build a (very poor) graph db. But it will never be as performant as something dedicated, because the data structures used in SQLite have terrible characteristics for the algorithms used in graph operations. I like SQLite a lot, but it targets a niche and graph operations is simply not that niche.
- oever 5y agoUsing Sqlite for semantic web is the topic of the post. And while Sqlite is not suited for billions of triples, it can be adequate for the small graphs. The advantage is that Sqlite is widely deployed.
- wheels 5y agoYou're not understanding what I mean by locality in this context. I mean that the on-disk structure should approximate an array of indexes, e.g.: const unsigned *edge_indexes = (const unsigned *) mmaped_structure_on_disk[disk_offsets[index]]; That's not the same as just being able to access those via an API; locality in this case means that you shouldn't need extra seeks for every single value, nor have to make a bunch of round trips through SQL. That is the primary difference between traditional relational databases and column-oriented databases. Normal relational databases have row-based locality; column-oriented databases have column-based locality.
- oever 5y agoSqlite does indeed not have an array type for columns. The overhead of (de)serializing binary values or json values would have to be offset against the advantage of cache coherence. Or one could implement a virtual table that stores arrays in integers.
- fauigerzigerk 5y agoIf your typical query visits all adjacent edges of a few nodes then the locality you're talking about is great. If your queries typically aggregate over a few types of edges across many nodes then this locality is the worst case. So the question is, do we actually want to run graph algorithms or is the data graph structured for other reasons? You're implying that choosing a graph representation means we want to perform graph analyses. I disagree with that if we're still talking about the semantic web. RDF is a general purpose knowledge representation model. The triple structure lends itself well to combining data from different sources with little coordination. It happens to form a graph, but running graph algorithms is just one of many special purpose problems.
- wheels 5y agoYes, you're mostly correct. (The one caveat being that if you know your access patterns in advance, you can choose if you store the edge type along with the edge, or if you store separate edge types in separate columns.) I'm not even arguing against storing RDF in SQLite. There are times that would make sense, and times that it wouldn't. I'm primarily replying to: > What is a graph database? A miserable little pile of joins. > Though to be serious: what do you expect a graph database to provide that sqlite cannot / does not do efficiently? That seemed like a general questions of, "What are graph databases for, and why would someone use them?" And I'm trying to answer that question.
- Groxx 5y agoIt was intended as both a "what significant benefits would it bring for this problem" (for implied exposed DBs with relatively small datasets... though "relatively" is relative of course), and "I don't see how using SQLite would somehow prevent this from happening". If there was some kind of "normal" query on thousands-to-millions of items that was prohibitively terrible on SQLite but not on graph-database-X, yea - I'm interested :) And I totally buy that graph DBs are better at graph queries in general. I just have yet to hit these kinds of limits in my use of SQLite (a fair number of instances with tens of gigabytes, a few with billions of rows) - a sprinkling of reasonable database design addresses almost all issues. The main one I can see is that, with longer-term use, SQLite's lack of any way to force locality would be fairly crippling. You'd need to make a reasonable sort order and periodically re-insert data in that order to optimize / vacuum. That's... technically achievable, but is a big downside compared to something that can dynamically organize / optimize it based on [some heuristic].