12 ms·
Bullshit graph database performance benchmarks
- alexchantavy 4y agoThanks for digging and sharing, I enjoyed your snark. > They decided to provide the data not in a CSV file like a normal human being would, but instead in a giant cypher file performing individual transactions for each node and each relationship created. Not batches of transactions… but rather painful, individual, one at a time transactions one point 8 million times. So instead of the import taking 2 minutes, it takes hours. Yeahhh I noticed this too when I looked at the repo when their blog was posted a couple weeks back. Running a transaction for each object will of course be very slow and real production code will (hopefully) not do this. > Those are not “graphy” queries at all, why are they in a graph database benchmark? Ok, whatever. I’m definitely interested in seeing more realistic scenarios of actual “graphy” queries with batched transactions comparing the two. Oh, and comparing against Neptune would be cool too since that supposedly uses openCypher now (which I hear is kinda close to neo4j cypher?).
- mapleeman 4y agoThis is true regarding the transactions and cypherl. All data is cypherl transactions because memgraph can handle a large volume of transactions. mgbench was designed to run in-house CI/CD, and mgBench is still tightly coupled with Memgraph. That is the reason we are still running everything in transactions. We did open an issue where we plan to improve things, adding CSV support for faster imports being one of them. https://github.com/memgraph/memgraph/issues/689 https://github.com/memgraph/memgraph/issues/689 Feel free to suggest things, some things Max suggested we will add. Agree on the more complex queries, and different vendors.
- PreInternet01 4y agoWhile this seems to be a pretty egregious example of a vendor benchmark misleading through cherry-picked unrealistic results, I'm not sure I share the author's pessimism about how these kinds of stunts will hold back the graph database market. Why? Simple: pretty much any benchmark I've seen of anything, ever, was similar nonsense -- give people numbers to game and they'll do so, enthusiastically. Even supposedly gold-standard benchmarks like the TechEmpower framework benchmarks quickly devolve into "application server handling HTTP requests by responding with predefined strings", which is as fast as it's utterly useless in most people's version of the real world. The only way to get usable benchmark data is to run your own workloads in your own environment: everything else is pretty much noise.
- LLcolD 4y agoYes, running a benchmark on your data is the only way. I've taken a look at both benchmarks (the one from OP and the one from Memgraph). They seem like different types of benchmarks and different approaches. But I still find it interesting that although the numbers in OP's are not so much in favor of Memgraph it turns out that Memgrpah is faster than Neo4j in large number of benchmark queries. So yes, it all comes down to type of benchmark and data that you use. I've also noticed (from OPs tweet https://twitter.com/maxdemarzi/status/1613075177704677376 https://twitter.com/maxdemarzi/status/1613075177704677376) that he used Enterprise version of Neo4j, but it doesn't say which Memgrpah version was used. I don't have experience with this two databases, but usually ENT versions are somewhat better than community ones. [EDIT]: I fixed few typos.
- dtomicevic 4y agoMemgraph compared the freely available open source editions of both databases. Neo4j Enterprise seems to have more performance optimizations compared to the community Edition.
- gymbeaux 4y agoI mean it should come as no surprise that an in-memory graph DB outperforms one that stores data on a hard disk, even an NVMe SSD. I would also add that the primary sell for Memgraph seems to be “fast enough that it can process data as it comes in via a stream, and present it to the user in a reasonable timeframe”. Anyone facing this use-case would want to use Memgraph regardless of how much faster it is than Neo4j.
- anonbystander 4y agoThat's their claim, but who knows. The article shows: * Memgraph's benchmark only show SQL ~where clauses, not graph ones * (nor streaming ones) * The existing memgraph numbers are questionable, and if the competitor tuned, who knows * The memgraph team refuses to use community-defined graph benchmarks for these articles.. so we won't know * Memgraph uses weird patterns like doing bulk loads as a query stream of atomic singleton creations vs batching (csv, arrow, ...), so even if it was graph/streaming, a proper benchmark would show tools going way faster b/c the relevant task would instead be for csv/arrow/etc bulk loaders or some other form of micro/macro batching It's not just this article but the others too. It's frustrating to watch the memgraph leaders take their VC money and dump it into a big negative campaign lying about basically anyone in the community. They even spend money punching down at academics doing OSS. I haven't been this annoyed at a seemingly real tech company in a long time.
- taubek 4y agoFor more context: - blog post that sparked the discussion - https://memgraph.com/blog/memgraph-vs-neo4j-performance-benchmark-comparison https://memgraph.com/blog/memgraph-vs-neo4j-performance-benc... - earlier discussion about this Memgraph benchmark HackerNews - https://news.ycombinator.com/item?id=33813781 https://news.ycombinator.com/item?id=33813781 - the benchmark results - https://memgraph.com/benchgraph/ https://memgraph.com/benchgraph/ - benchmark repo and methodology - https://github.com/memgraph/memgraph/tree/master/tests/mgbench https://github.com/memgraph/memgraph/tree/master/tests/mgben...
- deleted 4y ago[deleted]
- zwaps 4y agoAh yes, memgraph. I argued with them here about another bullshit benchmark https://news.ycombinator.com/item?id=33717766 https://news.ycombinator.com/item?id=33717766 they did reply tho
- szarnyasg 4y agoA plug: if you are looking for TPC-style application-level benchmarks for database systems, check out the LDBC Social Network Benchmark [1]. It has workloads for both OLTP and OLAP systems. We designed both of these to prevent many of the common benchmarking mistakes. To ensure that implementations follow the specification and their results are reproducible, we have a rigorous auditing process (similarly to TPC's benchmarks) [2]. [1] https://ldbcouncil.org/docs/presentations/ldbc-snb-2022-11.pdf https://ldbcouncil.org/docs/presentations/ldbc-snb-2022-11.p... [2] https://ldbcouncil.org/benchmarks/snb/ https://ldbcouncil.org/benchmarks/snb/
- mapleeman 4y agoThis is indeed a good industry leading benchmark.
- mbuda 4y agoLDBC is really great, it is and it should be the standard non-biased benchmark to look at. The point with https://memgraph.com/benchgraph https://memgraph.com/benchgraph is to be more tilted towards some specific workloads, easy to extend and run, etc. Time will tell to which extend that could be achieved, but definitely a place to look for some interesting workloads and results. + the whole thing is super early and it's going to evolve! Again, definitely take a look at LDBC benchmarks they are in general the most relevant.
- alfiedotwtf 4y agoOn a tangent, what Graph Database would people recommend in 2023? In particular, I would like something that's linked in like SQLite rather than a full blown service like MySQL etc
- rjh29 4y agoFor simple cases, you can get pretty far storing relations 6 times in SQLite, or any old key/value store. (a-b-son, b-a-father, son-a-b, father-b-a, a-son-b, b-father-a)
- steve_gh 4y agoThis is interesting - can you expand a little or provide a link. I get a-b links, but where do father and son come in?
- rjh29 4y agoI think I read it in a tech company's blog post, but there are some Wikipedia articles on the subject: https://en.wikipedia.org/wiki/Triplestore https://en.wikipedia.org/wiki/Triplestore https://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80%93value_model https://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80... The father/son is the relation between the two nodes, to support multiple relation types.
- stevesimmons 4y agoKuzu looks very interesting: https://github.com/kuzudb/kuzu https://github.com/kuzudb/kuzu Discussed here yesterday: https://news.ycombinator.com/item?id=34358912 https://news.ycombinator.com/item?id=34358912
- alfiedotwtf 4y agoThanks for the link!
- chillfox 4y agoIf you only need a few graph queries then you could just use SQLite, it’s capable of doing it (I have done it before). But writing graph queries in SQL is painful, so I wouldn’t do it if you need more than a handful.
- LAC-Tech 4y agoWhat are people using graph databases for, and what do your queries look like? I've read about them briefly but I have to admit my imagination fails me as to how it would look in the real world.
- jiggawatts 4y agoIn a word: Facebook. A more technical use case that I liked was a system that can analyse the configuration of resources across and entire network and find a "path" from a normal user account to a full admin privilege. Something like: "Helpdesk user A can reset the password of a service account that can write to a file share that contains a script that is run on logon by every user including the full admin, allowing user A to trigger an action in the context of an admin B, making them equivalent to an admin." You map out "things" on the network like file shares, security groups, accounts, etc... with links between them, and then ask for the shortest path from A to B.
- randomdata 4y ago> In a word: Facebook. Which, funnily enough, uses a relational database.
- bryanrasmussen 4y agoI'm betting Facebook uses a lot of different types of databases.
- randomdata 4y agoNo doubt, but the core product known for being graph-y is based on MySQL. Indeed, there is a graph data store (TAO) built on top of that base, but as we're talking about databases...
- jiggawatts 4y agoMany graph databases are relational "under the hood". The graph part is often just a specialised index.
- mhio 4y agoI get that this is trying to point out that neo4j shouldn't be that far behind, but why are the i7/gatling test numbers being directly compared to memgraphs g6 test results? The conclusion is a bit premature without the other half of the test... What performance does memgraph have on the newer, single socket hardware?
- junon 4y agoYeah that was strange, it's my understanding that you can't compare benchmarks between different machines, especially if they're not 1:1 identical hardware. If you're referring to this line, then it struct me as very odd. > Instead of 112 queries per second, I get 531q/s. Instead of a p99 latency of 94.49ms, I get 28ms with a min, mean, p50, p75 and p95 of 14ms to 18ms. Alright, what about query 2? Same story. Otherwise, the article holds up.
- robertlagrant 4y agoAuthor is just stating the differences between the benchmarketing hardware and his own. Not comparing new hardware and one DB with old hardware and other DB.
- roetlich 4y agoSeems like he does in his conclusion: > It looks like Neo4j is faster than Memgraph in the Aggregate queries by about 3 times.
- robertlagrant 4y agoHaving re-read it, I now can't decide. It would be a little silly if the author is completely different devices, so I'm going to stick to that interpretation.
- mhio 4y agoThe bottom table contains the memgraph mgBench (G6 2x Xeon X5650) "hot run, medium, isolated" throughput results: https://memgraph.com/benchgraph/base?condition=hot&datasetSize=medium&workloadType=mixed&querySelection=%5B%7B%22index%22:0,%22queries%22:%5B0,1,2,3%5D%7D,%7B%22index%22:1,%22queries%22:%5B0,1,2,3,4,5,6,7,8,9,10,11,12,13,14%5D%7D,%7B%22index%22:2,%22queries%22:%5B0%5D%7D%5D&tab=Global%20Results https://memgraph.com/benchgraph/base?condition=hot&datasetSi... The "By" columns compares those results to the new test suite on the ~10y newer cpu.
- mastermedo 4y agoI wouldn’t say the benchmarks put out by graph databases are bullshit. But there is a need for a standardisation of how they’re produced. The main problem is that when you’re comparing two products you’re bound to be comparing apples to oranges. Every product solves a slightly or majorly different challenge. So when you run n tests on two different products some tests are bound to perform better on one product and some on the other. Misleading marketing comes into the picture if you only publish the ones that went your way or just partial results. But that’s why if you believe in your own product and want benchmarks you hire a reputable third party to do them on their own accord.
- th3sly 4y agoit even has a name: benchmarketing :D
- beastman82 4y agoWhich graph database is actually fast and doesn't use deceptive marketing?
- dwroberts 4y agoThe author works on RageDB (https://ragedb.com/ https://ragedb.com/) and this doesn't seem to be disclosed in the article
- alias_neo 4y agoI'm not sure why it would need to be "disclosed", other than to suggest the author knows what they're talking about due to "domain knowledge".
- chaps 4y agoIt's because he has reason to destroy other competitors. Nice read otherwise.
- LLcolD 4y agoHe does "disclose" in related blog post [1]: "I don’t work for Neo4j anymore, why am I here defending them? Well… that and the fact that I still have a dinghy load of vested shares I have to sell so I can buy a place in the Villages and begin a new life as a golf cart driving day drinker." This seems like a series of post on benchmarking results from different vendors so if he "disclosed" it once I don't think that there is need for another one. [1] https://maxdemarzi.com/2022/12/06/khop-baby-one-more-time/ https://maxdemarzi.com/2022/12/06/khop-baby-one-more-time/
- ramraj07 4y agoMaybe because I don’t trust someone who allegedly writes databases but is proud about not knowing python.
- oxfordmale 4y agoBenchmarks are generally useless unless they test real world scenarios. The DataBricks data warehouse record costed $5,190,345 USD to run over a period of 3 years. If I spend that amount of money, I will get fired. Such benchmarks also ignore the engineering expertise an organisation has. Do you need to be an expert to fine tune 6000 parameters or can you tune the system to an acceptable standard by reading a few blogs. Some people pointed out the actual query only coated $242. My counter argument is that this appears to be based on buying reserved instances from AWS for 3 years. In real life this query would also run daily, or at least you would need several iterations to get the results you want. The costs also include a super low budget laptop ($279). It is more than fine for running the query, however, you wouldn't use it a development machine. This shows these results have been heavily massaged.
- dgb23 4y agoNot to mention that engineering expertise is just the potential. You then also need the time and the willingness to actually do that kind of tedious and potentially slow moving work instead of all the other things on your list. And as we all know, the list of things that can be improved in any system typically grows over time. The 'out of the box' or naive and un-optimized performance of something is the baseline. And with something as huge and self-contained as a database you want the happy path to be fine in terms of performance.
- salt-shaker 4y agoI was curious about it, so I tried to figure out where you got this number from. It looks like your source is https://www.tpc.org/results/individual_results/databricks/databricks~tpcds~100000~databricks_sql_8.3~es~2021-11-02~v01.pdf https://www.tpc.org/results/individual_results/databricks/da..., but you interpreted it wrong. The number you quoted is the projected 3-year ownership of the system configuration that was used to run the test, so the actual cost is a small fraction of the number you quoted.
- oxfordmale 4y agoIt is worth noting the compute cost appears to be based on purchasing reserved instances from AWS. The price of on demand instances is much higher. The laptop is also very low budget. I am sure it is fine to run the final query, however, you would unlikely to be able to use that as a development machine.
- ashvardanian 4y agoI am really stunned by this story. It made me check the MemGraph benchmarks section. Don't get me wrong, it may be 10-100x faster than Neo4J in even the most basic operations. Moreover, given the quality of Neo4J, it is hard not to be that much quicker. Even Postgres and MySQL are better at storing graphs than Neo4J. --- Disclosure: I have worked on Graph Algorithms, Graph Databases, and Database Engines for years, and we are now preparing a commercial solution based on UKV [1]. I don't know anyone at MemGraph or Neo4J. Never used the first. As for the second, I am not a fan. --- Aside from licensing, there are 3 primary complaints. I will address them individually, and I am open to a discussion. A. Using Python for Benchmarks instead of Gatling. I don't entirely agree with this. Python still has the fastest-growing programming community while already being one of the 2 most popular languages. Gatling, however, never heard of it. Choosing between the two, I would pick Python. But neither works if you want to design a High-Performance benchmark for a fast system. Without automatic memory management and expensive runtimes, you can only implement those in C, C++, Rust, or another systems-programming language. We have faced that too many times that the benchmark itself works worse than the system it is trying to evaluate [2]. B. Using hardware from 2010 [3], weird datasets [4]. This shocked me. When I looked at the charts [5] and the benchmarking section, it seemed highly professional and good-looking. I wouldn't expect less from a startup with $20M VC funding. But the devil is in the details. I would have never expected anyone benchmarking a new DBMS to use now 13-year-old CPUs and an unknown dataset. Assuming current developer salaries, hiring people to design a DBMS doesn't make sense if you will be evaluating on a $1000 machine is just financially irresponsible. We buy expensive servers, they cost like sports cars or even apartments in poorer countries. It is hard to maintain, but they are essential to quality work. It is sad to see companies taking such shortcuts. But to be a devil's advocate, there is no 1 graph benchmark or dataset that everyone agrees on. So I imagine people experimenting with multiple real datasets of different sizes or generating them systemically using one of the Random Generator algorithms. In UKV, we have used Twitter data to construct both document and graph collections. In the past, we have also used `ci-patent`, `bio-mouse-gene`, `human-Jung2015-M87102575`, and hundreds of other public datasets from the Network Repository and SNAP [6]. There are datasets of every shape and size, reaching around 1 Billion edges, in case someone is searching for data. For us the next step is the reconstruction of the Web from the 300 TB CommonCrawl dataset [7]. There is no such Graph benchmark in existence, but it is the biggest public dataset we could find. C. Running query different number of times for various engines. This can be justified, and it is how current benchmarks are done. You are tracking not just the mean execution time but also variability, so if at some point results converge, you abrupt before hitting the expected iterations number to save time. --- LDBC [8] seems like a good contestant for a potential industry standard, but it needs to be completed. Its "Business Intelligence workload" and "Interactive workload" categories exclude any real "Graph Analytics". Running an All-Pairs-Shortest-Paths algorithm on a large external memory graph could have been a much more interesting integrated benchmark. Similarly, one can make large-scale community detection or personalized recommendations based on Graphs and evaluate the overall cost/performance. It, however, poses another big challenge. Almost all algorithm implementations for those problems are vertex-centric. They scale poorly with large sparse graphs that demand edge-centric algorithms, so a new implementation has to be written from scratch. We will try to allocate more resources towards that in 2023 and invite anyone curious to join. --- [1] https://github.com/unum-cloud/ukv https://github.com/unum-cloud/ukv [2] https://unum.cloud/post/2022-03-22-ucsb https://unum.cloud/post/2022-03-22-ucsb [3] https://github.com/memgraph/memgraph/tree/master/tests/mgbench#intel---hp https://github.com/memgraph/memgraph/tree/master/tests/mgben... [4] https://github.com/memgraph/memgraph/tree/master/tests/mgbench#pokec https://github.com/memgraph/memgraph/tree/master/tests/mgben... [5] https://memgraph.com/benchgraph/base https://memgraph.com/benchgraph/base [6] https://snap.stanford.edu/data https://snap.stanford.edu/data [7] https://commoncrawl.org https://commoncrawl.org [8] https://ldbcouncil.org/benchmarks/snb https://ldbcouncil.org/benchmarks/snb
- Annatar 4y ago[dead]
- college_physics 4y agomy feeling is that graph databases face an uphill battle for mass adoption not because their architects or vendors doing anything wrong but some intrinsic aspects of information exchange in most current situations and use cases * information tends to be private and/or commercially sensitive, this severs the links that graph dbs are good at representing (and made the "node focused" SQL approach the ubiquitous model that it currently is) * objects in typical schemas have many more attributes that relations. while you could model things RDF style where everything is a relation, it is not the most intuitive for people * when the previous constraint does not apply (e.g. data from a centralized social network), it is typically not too hard to emulate an adequate graph structure on a Pareto 20/80 basis using an RDBMS so graph dbs end up being optimal only for a niche of situations and probably not the impact that the people / investors involved in their development would be happy with on the other hand the ginie is out of the bottle and the decades-long SQL monoculture seems to be coming to an end. but maybe what results is a relational database+ type thingy [0] rather than two disconnected paradigms [0] https://postgresconf.org/conferences/2020/program/proposals/postgres-as-a-graph-database https://postgresconf.org/conferences/2020/program/proposals/...
- apavlo 4y ago> but maybe what results is a relational database+ type thingy This is called already called "object-relational" model. It was invented by Postgres in the 1980s. The relational model / SQL absorbs the best part of alternative systems and get better over time. SQL:2023 is adding support for graph queries (SQL/PCG). Graph DBMSs are a passing fade.
- mbuda 4y agoIt takes time, but szarnyasg and I will convince you otherwise :)
- college_physics 4y agoI find projects like apache/age [0] are very promising in this direction. But I wouldn't call graph db's a passing fade. A more appropriate description might be "too important to be left alone, yet not important enough to form a second type of mass market database engine". [0] https://github.com/apache/age https://github.com/apache/age
- bjornsing 4y ago> Why would they do this? Because it’s a bullshit benchmark and they don’t actually want anybody looking too deeply at it. Very unlikely I would say. Most likely ”the clowns” never considered how the license terms would impact use of benchmarks. “Never attribute to malice why can be sufficiently explained by incompetence” as the saying goes.
- 0xbadcafebee 4y agoCompletely useless tangent: the word "benchmark" comes from a mark that surveyors would make in rock so that they could place a leveling rod for surveying. Benchmarks are made relative to other benchmarks so that surveying can be done relative to the height of one known fundamental benchmark. It could be argued that it isn't really a benchmark unless you can accurately calculate the result based off of a common fundamental benchmark.
- DougBTX 4y agoCool, I always assumed that a benchmark was a mark on a bench, but it turns out that it is the bench which goes on the mark. https://en.wikipedia.org/wiki/Benchmark_(surveying) https://en.wikipedia.org/wiki/Benchmark_(surveying)
- chairmanwow1 4y agoThis is a great article. Usually authors have snark and no substance but this author was able to back up his takes with excellent notes.
- lbriner 4y agoThere are lots of what I would call "grey" marketing/sales type articles like this across virtually every saas business, it's how they get people onto their site. Unfortuantely, an article that overstates benefits without any caveats is not illegal so it will carry on. Many of us would like disclaimer e.g. "I work for the company" but also a much more bounded discussion, "this performance test works for this particular scenario" and perhaps "Please note, your scenario might be very different" and especially "Please contact me if you think I have missed something out". I have worked for a business where you felt compelled to amplify the good and not talk about the bad but the world keeps spinning...
- jerf 4y agoOne of my favorite things is "the thing that sounds obvious when I say it, but you didn't think of it before". Here's one related to benchmarking: For A to be 120x better than B in a comparable task, that has to mean that B is leaving that much performance on the table in the first place. Now, let's combine this with one of the persistent tendencies of developers to take one specific benchmark as indicative of the overall performance, which is often preyed on by benchmarkers trying to sell things. Is it really plausible that neo4j takes 120x longer than it needs to on all operations? A dedicated graph database that has been tuned and optimized for that task for quite a while now? I'm not quite going to rate that a 0 probability, but it's definitely a very big claim. While the probability is not 0, it is comfortably below "someone's gaming the numbers" and "the benchmark is not as comparable as claimed". There's a faint chance the latter may match a production use case; for instance, certain comparisons of NoSQL DBs and SQL DBs are "not fair" in that they won't be doing remotely the same things for the queries and the performance landscape is very complicated, with one side winning handily for some tasks and the other side handily winning for others, but if your use case falls into one of those big wins you may not care about the "fairness". But it's still a pretty big chunk of probability mass that it's just plain not comparable; how many times have we seen a ludicrous benchmarking claim of relative superiority just for the losing side to pop up and say something to the effect of "Hey, did you consider adding the correct index to the data, oh look if you do that we win by a factor of 4." Tell me you're 1.2x or 1.5x faster or something, or that your clever compression means I can remove 1/3rd of my systems or something. Keep it in the range of plausible. While I'm sure this won't affect the marketing of this company any, ludicrously large claims of 10x+ speed improvements actually turn me off, not attract me. You'd better have some sort of super compelling reason why you somehow managed to be that fast over your competitor, like, "we're the first to successfully leverage GPUs" or something like that. Otherwise I'm going to guess "Actually, you have an O(log n log log n) algorithm over their O(log n log n) algorithm and you cranked the data set up to the ludicrous sizes it takes to get an arbitrarily large X factor improvement over your competition" or something like that. (Always gotta love people comparing two completely different O(...) algorithms against each other and declaring one is X times faster than the other. This is another major source of "10,000x faster!"... yeah, O(n log n) is "10,000 faster!" than O(n^2), sure. It's also 100,000 times faster, 10 times faster, and a billionkajillion times faster, all at the same time.)
- maxdemarzi 4y agoAuthor here to clear up a few questions: I did not run any benchmarks for Memgraph, just Neo4j on my machine and compared them to their numbers on their machine. My 8 faster cores to their 12 slower cores, so not apples to apples, but close enough to make the point that Memgraph is not 120x times faster than Neo4j. I used to work at Neo4j, then at AWS for Neptune, I work on my own graph database http://ragedb.com/ http://ragedb.com/, and work for another database company https://relational.ai/ https://relational.ai/ If you want to be my hero, find a way to fix this problem: https://maxdemarzi.com/2023/01/09/death-star-queries-in-graph-databases/ https://maxdemarzi.com/2023/01/09/death-star-queries-in-grap...
- redbar0n 4y ago> If you want to be my hero, find a way to fix this problem: https://maxdemarzi.com/2023/01/09/death-star-queries-in-graph-databases/ https://maxdemarzi.com/2023/01/09/death-star-queries-in-grap... Let me (try to) be your hero, Marzi. (Insert favorite reference to famous cheezy pop music song, if you like.) Couldn't you use GraphBLAS algorithms, like they do in RedisGraph (which supports Cypher, btw) to fix that problem with "death star" queries? Those algorithms are based on linear algebra and matrix operations on sparse matrices (which are like compressed bitmaps on speed, re: https://github.com/RoaringBitmap/RoaringBitmap https://github.com/RoaringBitmap/RoaringBitmap ). The insight is that the adjacency list of a property-graph is actually a matrix, and then you can use linear algebra on it. But it may require the DB is built bottom up with matrices in mind from the start (instead of linked lists like Neo4j does). Maybe your double array approach in RageDB could be made to fit.. I think you'll find this presentation on GraphBLAS positively mind-blowing, especially from this moment: https://youtu.be/xnez6tloNSQ?t=1531 https://youtu.be/xnez6tloNSQ?t=1531 Such math-based algorithms seem perfect to optimally answer unbounded (death) star queries like “How are you connected to your neighbors and what are they?” That way, for such queries one doesn't have to traverse the graph database as a discovery process through what each node "knows about", but could view and operate on the database from a God-like perspective, similar to table operations in relational databases. Further reading: https://graphblas.org/ https://graphblas.org/
- deleted 4y ago[deleted]
- AtNightWeCode 4y agoNeo4J is bad at aggregated queries but that is not what to use a graphdb for in the first place.
- manv1 4y agoThe real problem with these kinds of "benchmarks" is that either the company doesn't have anyone on staff that's calling "bullshit" on it or the marketing people don't care that it's bullshit. Either one is a bad sign if they're going to be a vendor. At that point how can you trust their SLAs and/or their presales team?
- taubek 4y agoIn this thread I've seen comments from Memgraph CTO and founder (mbuda), CEO and founder (dtomicevic), and one of the developers that worked on mbench (mapleeman). To me it seems that they are addressing all of the questions in comments section.
- AtlasBarfed 4y agoBwahahaha, it's been against the terms of use of Oracle to benchmark forever. Lies, damn lies, and benchmarks people, take them all with a huge grain of salt.