10 ms·
Graph Databases 101
- valine 11y agoQuestion as someone new to graph databases: Are there any open source graph databases worth looking into?
- timClicks 11y agoNeo4j is a very good option.
- iod 11y agoArangoDB is free open source multi model no-sql db that has decent¹ perfomance with graph support: https://www.arangodb.com https://www.arangodb.com ¹ https://www.arangodb.com/2015/10/benchmark-postgresql-mongodb-arangodb/ https://www.arangodb.com/2015/10/benchmark-postgresql-mongod...
- rspeer 11y ago"The performance will suffer if the dataset is much bigger than the memory." That is a huge drawback when compared to relational databases. A good follow-up question would be: which open-source graph databases can reasonably import and store graph data that's not small -- that is, more data than fits in than RAM? Without proprietary extensions?
- jerven 11y agoBlazegraph Virtuoso Jena tdb Can easily load large to very large graphs.
- reactor 11y agoLooks like they have persistent indexes in the roadmap (ver 3.0) https://www.arangodb.com/roadmap/ https://www.arangodb.com/roadmap/ which might help.
- LyndsySimon 11y agoHow much data are you dealing with? In my (admittedly limited) experience, it's usually cheaper to throw money at hardware than other options.
- rspeer 11y agoThrowing money on RAM because most graph databases haven't figured out how to use the disk effectively is not a good use of money. What's cheaper than buying all the RAM in the universe is figuring out a different system besides a graph database that does the job. A good example would be the graph of Wikipedia links. About 100 million edges among 5 million nodes, last I checked. The nodes have large differences in degree. The raw data for this is not the slightest bit large. We're only talking about gigabytes. But it would absolutely destroy Neo4J to even try to import it, to say nothing of running an interesting algorithm that justifies using a graph database on it, and Neo4J seems to be everyone's favorite open-source graph database for some reason.
- neunhoef 11y agoOne of the developers of ArangoDB here. Let me explain this quotation. When your graph data (including indices) do no longer fit into the RAM of a single server, you can either live with the higher latency of loading data from disk or you can use sharding, which will lead to communication and therefore slower traversals. That does not mean that things stop working, but performance will be less good, you can no longer visit tens of millions of nodes per second in a traversal as in RAM on a single server. If you actually only traverse a much smaller hot subgraph, I would probably go for the disk based single server approach. If your graph has a natural known clustering, then an optimized sharding solution with fine tuned sharding keys us probably your best bet. You can do all this with ArangoDB. However, graph traversals vary greatly in many respects, and your mileage may vary accordingly, with any approach. I would love to chat in more detail about your use case.
- kinow 11y agoDepends on what kind of data and graph you are going to store/use. Neo4j is quite popular, cypher isn't very hard to learn, and it has lots of examples. Might be a good choice for a beginner. https://en.wikipedia.org/wiki/Graph_database#List_of_graph_databases https://en.wikipedia.org/wiki/Graph_database#List_of_graph_d...
- valine 11y ago> cypher isn't very hard to learn Oh but I love a challenge. Are there reasons to choose cypher besides a gentle learning curve?
- kinow 11y agoNot really, if you learn Cypher you should be fine learning the basics of Gremlin, SPARQL, or other languages to operate on graphs in a few hours. There was some post about enabling SPARQL in Neo4J, but when you install Neo4J it comes with cypher by default (not sure if it supports anything else). I use Apache Jena + SPARQL, but had to use Neo4J to help in a master thesis. Took me a few hours of "How the heck can I do that same thing I'd do in SPARQL that way?", plus some reading of the tutorials. Edit: some old post with example of Cypher, Gremlin and SPARQL: http://kinoshita.eti.br/2014/09/09/cypher-gremlin-and-sparql-graph-dialects.html http://kinoshita.eti.br/2014/09/09/cypher-gremlin-and-sparql...
- kevinschumacher 11y agoYou can definitely run gremlin queries against Neo4j by a couple of methods. One example: https://github.com/thinkaurelius/neo4j-gremlin-plugin https://github.com/thinkaurelius/neo4j-gremlin-plugin Also can use the Tinkerpop3 or Blueprints APIs to access your graph with Gremlin.
- jexp 11y agoCypher expressses the graph patterns that you're looking for in an ASCII-art-syntax, so you don't loose sight of the core of your question. On top of that you get filtering, projection, aggregation, pagination. The most fun things are in-query dataflow which allows you to pass information from one query part to the next (projected, aggregated, ordered etc). And the really cool collection and map functions, so you save a lot of roundtrips between client and server. See: http://neo4j.com/developer/guide-sql-to-cypher/ http://neo4j.com/developer/guide-sql-to-cypher/
- kawera 11y agoCayley is a good option; we use it in production. https://github.com/google/cayley https://github.com/google/cayley
- jalfresi 11y agoMight I ask what sort of dataset size, servers etc you are using? I'm looking for an graph database and Cayley seems the best fit, though I'm not sure what sort of limits on the data there would be in the real world.
- owen11 11y agoCayley can store 130 million quads (2 nodes + connecting edge) on 20 GB harddrive. Join #cayley on freenode and https://groups.google.com/forum/#!forum/cayley-users https://groups.google.com/forum/#!forum/cayley-users and be part of our community!
- jerven 11y agoVirtuoso, does 2 billion in 50GB. Just so you know. And hard graphs like the complete UniProt data of 20 billion+ in 800GB.
- kawera 11y agoCayley has been stable for us so far but I can't vault for it's scalability as our database is very small, less than 6M quads, so an 8GB machine is more than enough.
- jedc 11y agoIs Cayley mature enough for production? I thought it was still relatively new. Would love to know a little more about how you're using it.
- owen11 11y agoThere is a similar question on Cayley's google group - https://groups.google.com/forum/#!topic/cayley-users/nirhCbqwnfU https://groups.google.com/forum/#!topic/cayley-users/nirhCbq... Also, I know it's being used by some companies. you can ask directly the people who uses it on IRC - #cayley (freenode).
- emehrkay 11y agoYes. I love the Tinkerpop stack (http://tinkerpop.incubator.apache.org http://tinkerpop.incubator.apache.org). I am currently writing a Python library as I develop an application around it called gizmo (https://github.com/emehrkay/gizmo https://github.com/emehrkay/gizmo).
- jerven 11y agoIn the rdf space there are a whole bunch. Graph as in sparkling gives you : Virtuoso Blazegraph Jena Rdf4j Ruby.rdf There are more but these are opensource and I know them. And money more commercial ones.
- rusabd 11y agoVirtuoso
- SanderMak 11y agoWe use http://orientdb.com/orientdb/ http://orientdb.com/orientdb/, seems decent so far.
- whazor 11y agoThere are multiple systems out there, however I have my doubts. It is important that your data does not get corrupted, and that your transactions will not get lost. Furthermore, speedups are possible with certain indices. That is why I personally would want to see some more safety/speed analysis and comparisons between the different systems.
- marknadal 11y ago(Full disclosure: I'm the author, we are VC backed) https://github.com/amark/gun https://github.com/amark/gun is an Open Source graph database with Firebase like realtime synchronization.
- phpnode 11y agoWe've had this discussion before but that product is not a graph database, it has no graph traversal features
- marknadal 11y agoDijkstra's algorithm is wonderful, but by no means is a requirement for being a graph database. A graph database is that, a database composed of nodes that can interconnect into a graph. GUN supports this and allows for traversing the graph. We haven't implemented Dijkstra's algorithm, which is what your "discussion" refers to.
- phpnode 11y agoWe've never talked about Dijkstra as far as I recall. Your product doesn't support graphs any more than, say, mongodb does because ultimately all you are doing is loading an object from JSON and then sending it to the consumer. This is in contrast to true graph databases whose main selling point is their ability to efficiently traverse, filter and aggregate large graphs to find the answer to some question, and then send the answer to the consumer. Gun doesn't meet anyone's definition of "graph database" other than your own. If I load some JSON from a URL, and use lodash to pluck some data out of it, is it a graph database?
- marknadal 11y agoDijkstra was your complaint about shortest path. GUN can do efficient traversal and filtering, and this is going to be even better in our 0.5.x release with lexical cursor support. By "anyone" do you mean Wikipedia's? https://en.m.wikipedia.org/wiki/Graph_database https://en.m.wikipedia.org/wiki/Graph_database , because GUN does match its definition. Although we haven't implemented Dijkstra's. I'm out in France right now and just boarded a plane to Slovenia, so I won't be able to reply again. Have a good one.
- rail2rail 11y agoWe're using TitanDB. One of the main benefits for us is that AWS has provided backend integration with DynamoDB. This affords you practically infinite and painless scaling on a pay-as-you-go model. Love it. https://aws.amazon.com/blogs/aws/new-store-and-process-graph-data-using-the-dynamodb-storage-backend-for-titan/ https://aws.amazon.com/blogs/aws/new-store-and-process-graph...
- deleted 11y ago[deleted]
- espeed 11y agoLook at Blazegraph, an open-source GPU-accelerated distributed graph database. See previous discussion: https://news.ycombinator.com/item?id=11197880 https://news.ycombinator.com/item?id=11197880
- karussell 11y agoHere is an (old) overview: https://docs.google.com/spreadsheets/d/1XGapLHpSd2Ta8019VlwYX-K-n7QqgrVkmxvU4Ws8lHc/edit?pref=2&pli=1 https://docs.google.com/spreadsheets/d/1XGapLHpSd2Ta8019VlwY...
- d0ne 11y agohttp://stingergraph.com/ http://stingergraph.com/ - From Georgia Tech
- TimPrice 11y ago1-Would it be more efficient to store objects that contain its relations if you only do (simple) read operations? (e.g. JSON database) 2-Instead, do graph DB engines try to break through bottlenecks for big data and analytics scenarios?
- thesz 11y agoIt introduces false dichotomy "graph vs relational". In fact, most (if not all) graph algorithms can be expressed using linear algebra (with specific addition and multiplication). And matrix multiplication is a select from two matrices, related with "where i=j" and aggregation over identical result coordinates. The selection of multiplication and addition operations can account for different "data stored in links and nodes". So there is no such dichotomy "graph vs relational".
- grandalf 11y agoTrue but highly irrelevant to why anyone might choose to use a graph database (or choose not to use one in favor of a relational database)...
- sklogic 11y agoOf course you can express anything on top of a relational model. But for graphs such a representation would have been awfully inefficient. For this reason, CADs never even tried to switch to a relational data storage once that fancy new relational databases appeared, most of the professional CADs are still using good old graph databases.
- thesz 11y agoI beg to disagree. I am part of the team developing Russian CAD system [0]. It uses what one can consider a hypergraph db (relation includes many objects), but that DBMS system has queries on par with SQL. And they prove themselves very useful in development of CAD. What you describe can be explained with development inertia. Most CADs are C/C++ and these languages are not very well suited for changes that go through all code base (change of storage engine and data model). I also made experiments during a dev of graph analytics engine in one part of my experience. The relational model (actually, linear algebra model) has proven itself very competitive. It allows for easy distribution of data, the operations over distributed data are close to optimal, etc, etc. [0] http://dd.ru/ http://dd.ru/
- sklogic 11y ago
- AdamN 11y agoEverybody's focused on graph databases here but let's talk about Cray! One of the most forward-thinking computer technology companies ever to exist is starting to get out there again. If they got a few hundred million dollars from an outside investor, they could do friggin' incredible things. They already do incredible things but not out there in the way it so easily could be.
- fulafel 11y agoCray is a brand name that has been passed around between half a dozen companies (including Sun and SGI) dotted by various kinds of product reboots and commercial failures. Cool stuff but supercomputing isn't the most financially sound business it seems. The current name holder is the company previously called Tera, originally famous for making an aggressively multithreaded HPC computer.
- mitchty 11y agoI'm not clear how Sun ever owned Cray care to explain? The provenance was Cray Research -> SGI -> Tera/Cray according to those that have been around since the Cray Research days. Source: err, I work here and asked a couple people a few cubes over. :) The Sun deal was apparently more SGI wouldn't be caught dead with a supercomputer that ran on sparc so it got sold off to Sun.
- pinewurst 11y ago"an aggressively multithreaded HPC computer" - still sold as Urika-GD, marketed as a dedicated graph appliance. That's where Cray Graph Engine originated.
- SloopJon 11y agoThe author's next post describes RDF and SPARQL in the context of the Cray Graph Engine: http://www.cray.com/blog/how-cray-graph-engine-manages-graph-databases/ http://www.cray.com/blog/how-cray-graph-engine-manages-graph...
- gtrubetskoy 11y agoI have spent a lot of time figuring out how to deal with a large graph a couple of years ago. My conclusion - there will never be such a thing as a "graph database". There are many efforts in this area, someone here already mentioned SPARQL and RDF, you can google for "triple stores", etc. There are also large-scale graph processing tools on top of Hadoop such as Giraph or Graphx for Spark. For the particular project we ended up using Redis and storing the graph as an adjacency list in a machine with 128GB of RAM. The reason I don't think there ever will be a "graph database" is because there are so many different ways you can store a graph, so many things you might want to do with one. It's trivial to build a "graph database" in a few lines of any programming language - graph traversal is (hopefully) taught in any decent CS course. Also - the latest versions of PostgreSQL have all the features to support graph storage. It's ironic how PostgreSQL is becoming a SQL database that is gradually taking over the "NoSQL" problem space.
- yeukhon 11y agoYears ago PostgreSQL already support recursive query, and in Oracle you have CONNECT BY. I have only used the recursive with once and it was just a quick demo, but my understanding is update is extremely expensive.
- amirouche 11y agoIn my point of view, the fact that you can add an expert index very easily to a graph database written in a modern language (say no C/C++) makes it even easier to customize an existing graph database to suit your direct need. In turn, storage and runtime can be tunned more easily. Making so easy to have the performance you need. But at the end of the day not dealing with algreba is the best.
- valhalla 11y agoIf anyone's curious about Network Science/Graph Theory in general here's a free online textbook used by a grad student friend of mine http://barabasilab.neu.edu/networksciencebook/downlPDF.html http://barabasilab.neu.edu/networksciencebook/downlPDF.html
- rfreytag 11y agoBlasted down arrow is so close to the up-arrow I clicked down when I meant up. Someone please cancel out my mistake.
- GFK_of_xmaspast 11y agoBe careful about Barabasi: https://liorpachter.wordpress.com/2014/02/10/the-network-nonsense-of-albert-laszlo-barabasi/ https://liorpachter.wordpress.com/2014/02/10/the-network-non... (FWIW, I had previously read some Barabasi papers and had come away seriously unimpressed, see also https://news.ycombinator.com/item?id=9555547 https://news.ycombinator.com/item?id=9555547)
- gilleain 11y agoI misremembered Albert-László Barabási for Laszlo Babai - I was wondering why you were unimpressed! Yes, scale-free networks (and so on, and so on), are oversold. Is his work really that bad though?
- amirouche 11y agoI am huge fan a graph-y stuff. I did several iteration over a graph database written -- in Python -- using files, bsddb and right now wiredtiger. I also use Gremlin for querying. Have a look at the code https://github.com/amirouche/ajgudb https://github.com/amirouche/ajgudb. Also, I made an hypergraphdb, atom-centered instead of hyperedge focused in Scheme https://github.com/amirouche/Culturia/blob/master/culturia/culture.md https://github.com/amirouche/Culturia/blob/master/culturia/c.... Did you know that Gremlin, is only srfi-41 aka. stream API with a few graph centric helpers. edit: it's srfi 41, http://srfi.schemers.org/srfi-41/srfi-41.html http://srfi.schemers.org/srfi-41/srfi-41.html
- lobster_johnson 11y agoI've seen people using graph databases as a general-purpose backing store for webapps/microservices. What are people's opinions about this? My feeling is that graph databases are not suitable/ready for — for lack of a better term — the kind of document-like entity relationship graphs we typically use in webapps. Typical data models don't represent data as vertices and edges, but as entities with relationships ("foreign keys" in RDBMS nomenclature) embedded in the entities themselves. This coincidentally applies to the relational model, in its most pure, formal, normal form, but the web development community has long established conventions of ORMing their way around this. The thing is, you shouldn't need an ORM with a graph database.
- Xyik 11y agoOne of the biggest challenges in databases is handling concurrency and sharding, wish this would have talked a bit more about how that changes between a graph database and a relational database.
- jmartins 11y agoAnybody know dgraph.io? it's a Scalable, Distributed, Low Latency, High Throughput Graph Database over terabytes of structured data. DGraph supports facebook GraphQL as query language, and responds in JSON and the storage engine is facebook rocksdb a very fast database. see more in https://github.com/dgraph-io/dgraph https://github.com/dgraph-io/dgraph