3 ms·
***Benchmark summary (TuringDB vs Neo4j)*** We benchmarked TuringDB against Neo4j using the Reactome biological knowledge graph, which is a real-world, highly
by remy_boutonnet 8mo ago
***Benchmark summary (TuringDB vs Neo4j)***
We benchmarked TuringDB against Neo4j using the Reactome biological knowledge graph, which is a real-world, highly connected dataset (millions of nodes/edges, deep traversal patterns).
- Both systems were run out of the box, cold start.
- No manual indexing, tuning, or query rewriting on either side.
- Same logical queries, same dataset.
- Benchmarks are reproducible via an open-source runner.
Multi-hop traversal from a small seed set (15 nodes):
- 1–4 hops: ~110×–130× faster
- 7 hops: ~300× faster
- 8 hops: ~200× faster
Example:
- Neo4j: ~98s for an 8-hop traversal
- TuringDB: ~0.48s for the same query
Label scans and label-constrained traversals:
- Simple label scan (`match (n:Drug)`): ~1200× faster
- Multi-label scan: up to ~4000× faster
More complex bidirectional traversals:
- Speedups range from ~6× to ~600× depending on query shape and result materialisation.
Why the difference exists:
The speedups are not from query tricks, but from architectural choices:
- *Column-oriented execution*: vectorized, SIMD-friendly scans instead of record-at-a-time execution.
- *Streaming traversal engine*: processes node/edge chunks in batches rather than pointer chasing one node at a time.
- *Immutable snapshots*: no read locks, no coordination overhead during deep analytical queries.
- *Graph-native storage*: traversal is the primary access path, not an emergent property of joins.
Neo4j’s record-oriented model performs reasonably for short traversals but degrades sharply on long paths and broad scans, especially on cold starts.
Scope and caveats:
- These benchmarks focus on *analytical read-heavy workloads*, not high-write OLTP scenarios.
- We’re not claiming universal superiority, different workloads want different systems.
- The full benchmark code and instructions are public and reproducible:
https://github.com/turing-db/turing-bench https://github.com/turing-db/turing-bench
This is not meant to replace transactional graph databases or general-purpose OLTP systems. It’s designed for large, read-heavy analytical graphs where latency, explainability, and control matter more than write throughput. The exact applications that blocked us when doing biomedical digital twins & large knowledge graphs.
Feedback always welcome.