2 ms·
Hi everyone! I’m one of the co-authors of BenchGraph[1], a platform for Graph Database Performance Benchmarks. Our platform shows the results of running benchma
by mapleeman 4y ago
Hi everyone! I’m one of the co-authors of BenchGraph[1], a platform for Graph Database Performance Benchmarks. Our platform shows the results of running benchmark tests (via mgBench) on supported vendors. It shows the overall performance of each system relative to others.
Inspiration came from ClickBench, a Benchmark For Analytical DBMS.
We previously developed mgBench as in-house testing infrastructure to benchmark Memgraph, and now we are adapting it to support other graph database vendors. In order to test graph database performance, mgBench executes Cypher queries on a given dataset. Queries are general and represent a typical workload that would be used to analyse any graph dataset. Running this benchmark is automated, and the code used to run benchmarks is publicly available. You can run mgBench yourself to validate the results on the BenchGraph platform. The methodology is explained in detail on GitHub repo [2]
As you can see, at the moment, we have two vendors on the platform. We would like to add more vendors to our platform. If you want, feel free to contribute.
Let me know if you have any questions or suggestions.
[1] https://memgraph.com/benchgraph https://memgraph.com/benchgraph
[2] https://github.com/memgraph/memgraph/tree/master/tests/mgbench#fire-mgbench-benchmark-for-graph-databases https://github.com/memgraph/memgraph/tree/master/tests/mgben...
- anualvis 4y agoHey mapleeman, why is the benchmark infrastructure not language agnostic? Shouldn't it be more desirable to keep the infra as a scheduler and verifier which can take an input in any language and process it while collecting the useful metrics. Restricting it to cypher leaves a bunch of DBs unable to run this.
- mapleeman 4y agoHey, thanks for the comment, you are 100% right, this is just the initial version since we are compatible with Neo4j, so it was a least effort to do it. It is just the initial setup, making it language agnostic will take a bit time. If you peak at the methodology and future part: https://github.com/memgraph/memgraph/tree/master/tests/mgbench#future-of-mgbench https://github.com/memgraph/memgraph/tree/master/tests/mgben... You will see that we have the plan to add more database vendors + make it language-agnostic. We are also keeping track of all comments regarding this, I have opened an issue: https://github.com/memgraph/memgraph/issues/689 https://github.com/memgraph/memgraph/issues/689, there is a language agnostic note in there. If you have any other input, it would be highly appreciated.
- samsquire 4y agoThat's a lot of work and would be a large project in itself to separate the test driver to different query formats and have it drivable from multiple programming languages. I implemented a toy Cypher database (samsquire/hash-db) and I just use a python test script. I am yet to benchmark, the performance is probably poor. I tried running standardised SQL benchmarks against MySQL but the benchmark code fell behind the MySQL client and it's work to maintain it. I inherited a Jepsen suite to test ActiveMQ and it wasn't easy to understand Testing can be a full time job!
- mapleeman 4y agoYes, testing and benchmarking is a full-time job! The extra issue here is making it work under a single client to minimise latency penalties across measurements, then, there is also a protocol issue. Nice project with hash-db, I guess it is quite the learning experience(distributed-multimodal)?
- menaerus 4y agoI agree. Designing a generic test runner running workloads by reading them from raw arbitrary SQL files and then executing them against XYZ database backend would have been much easier for adding new workloads and extending supported database backends. I wrote few such frameworks so I for sure know it's not a big deal. Sysbench does this through Lua which is also ok if you need more advanced scripting capabilities in the workload itself.
- jandrewrogers 4y agoThis is nice but I have a few comments: Measuring peak memory will be nonsensical for some implementations. Some databases do minimal dynamic allocation for performance reasons. Some will be paging to storage, which can work well for graph databases with an appropriate I/O scheduler design. This benchmark seems to assume all graph databases are in-memory and doing dynamic memory allocation. The test data models are tiny. Even the “large” test data model falls below the noise floor of scalable graph database architectures. This has the implication that good results overfit for graph databases that scale poorly. You need something closer to a billion edges to really exercise and differentiate the performance characteristics of graph databases, and in realistic applications that still isn’t a particularly large graph (maybe medium-sized?). It would also be useful to benchmark how long it takes to load and prepare the data. This is important operationally and, for many graph databases, unreasonably slow. Graph databases tend to skip over this part when talking about performance.
- mapleeman 4y agoThanks for all the comments and inputs. I have added the suggestions that we plan to implement: https://github.com/memgraph/memgraph/issues/689 https://github.com/memgraph/memgraph/issues/689. Both on Memory usage tracking and precise data on load/input. Regarding scale, we are aware of the issue, listed in limitations: https://github.com/memgraph/memgraph/tree/master/tests/mgbench#limitations https://github.com/memgraph/memgraph/tree/master/tests/mgben.... Next versions will probably have a billion nodes/relationships. Actually, Neo4j is particularly slow on writes, import/load times were 50x faster on Memgraph, but we didn't show it. Will do it in the next version for all vendors.