3 ms·
Why can't you do this in a relational DBMS? Recent research says it that would probably be faster. The Case Against Specialized Graph Analytics Engines http:/
by assface 11y ago
Why can't you do this in a relational DBMS? Recent research says it that would probably be faster.
The Case Against Specialized Graph Analytics Engines
http://www.cidrdb.org/cidr2015/Papers/CIDR15_Paper20.pdf http://www.cidrdb.org/cidr2015/Papers/CIDR15_Paper20.pdf
- mey 11y agoIt can and is done. Place I last worked at built an in house data model and lookup engine on MS SQL Server to find relations between arbitrary data points on edges. I know that it is currently used live in production and handles real time fraud evaluation at a large scale (it's a private company and not a liberty to disclose actual payment volume). With proper indexing and careful design it is scalable, but as with everything, you need to dedicate time to performance analysis and design work.
- ryguyrg 11y ago(disclosure: i work for neo4j) This paper is a bit unrelated to Neo4j. Neo4j is an ACID-compliant native graph database. It is not a "specialized graph analytic engine." Neo4j stores the graph data on disk (and caches in memory) as nodes and relationships. After index lookups to find the start points in the graph, all traversals of relationships are done in constant time -- allowing it to scale with linear performance characteristics, regardless of the size of the graph. The referenced paper cites two main reasons that RDBMS would be better for the analytics use cases: (1) ability to express graph queries in SQL and (2) performance of executing those queries. (1) Cypher is, like SQL, a declarative language. However, it represents graph constructs in a much more natural way -- "ASCII art for graphs." There's significant praise from developers on the web of the benefits of Cypher for traversing graphs, which is why we decided to open up the language: http://www.opencypher.org/ http://www.opencypher.org/ (2) As Neo4j isn't really intended as an analytics engine, its performance characteristics are not included in this paper. However, (expensive) indexes do not need to be created and maintained for every relationship in Neo4j. Similarly, these indexes do not need to be accessed for traversal (also expensive).