9 ms·
While this seems to be a pretty egregious example of a vendor benchmark misleading through cherry-picked unrealistic results, I'm not sure I share the author's
by PreInternet01 4y ago
While this seems to be a pretty egregious example of a vendor benchmark misleading through cherry-picked unrealistic results, I'm not sure I share the author's pessimism about how these kinds of stunts will hold back the graph database market.
Why? Simple: pretty much any benchmark I've seen of anything, ever, was similar nonsense -- give people numbers to game and they'll do so, enthusiastically. Even supposedly gold-standard benchmarks like the TechEmpower framework benchmarks quickly devolve into "application server handling HTTP requests by responding with predefined strings", which is as fast as it's utterly useless in most people's version of the real world.
The only way to get usable benchmark data is to run your own workloads in your own environment: everything else is pretty much noise.
- LLcolD 4y agoYes, running a benchmark on your data is the only way. I've taken a look at both benchmarks (the one from OP and the one from Memgraph). They seem like different types of benchmarks and different approaches. But I still find it interesting that although the numbers in OP's are not so much in favor of Memgraph it turns out that Memgrpah is faster than Neo4j in large number of benchmark queries. So yes, it all comes down to type of benchmark and data that you use. I've also noticed (from OPs tweet https://twitter.com/maxdemarzi/status/1613075177704677376 https://twitter.com/maxdemarzi/status/1613075177704677376) that he used Enterprise version of Neo4j, but it doesn't say which Memgrpah version was used. I don't have experience with this two databases, but usually ENT versions are somewhat better than community ones. [EDIT]: I fixed few typos.
- dtomicevic 4y agoMemgraph compared the freely available open source editions of both databases. Neo4j Enterprise seems to have more performance optimizations compared to the community Edition.
- gymbeaux 4y agoI mean it should come as no surprise that an in-memory graph DB outperforms one that stores data on a hard disk, even an NVMe SSD. I would also add that the primary sell for Memgraph seems to be “fast enough that it can process data as it comes in via a stream, and present it to the user in a reasonable timeframe”. Anyone facing this use-case would want to use Memgraph regardless of how much faster it is than Neo4j.
- anonbystander 4y agoThat's their claim, but who knows. The article shows: * Memgraph's benchmark only show SQL ~where clauses, not graph ones * (nor streaming ones) * The existing memgraph numbers are questionable, and if the competitor tuned, who knows * The memgraph team refuses to use community-defined graph benchmarks for these articles.. so we won't know * Memgraph uses weird patterns like doing bulk loads as a query stream of atomic singleton creations vs batching (csv, arrow, ...), so even if it was graph/streaming, a proper benchmark would show tools going way faster b/c the relevant task would instead be for csv/arrow/etc bulk loaders or some other form of micro/macro batching It's not just this article but the others too. It's frustrating to watch the memgraph leaders take their VC money and dump it into a big negative campaign lying about basically anyone in the community. They even spend money punching down at academics doing OSS. I haven't been this annoyed at a seemingly real tech company in a long time.
- gymbeaux 4y agoFair enough. I didn’t realize they were being so shady with the benchmarks. Isn’t Neo4j written in Java and Memgraph written in C++ (with lots of Python extensibility)? By that alone I would think Memgraph would be more performant most of the time, unless Memgraph is poorly-written/optimized vs Neo4j, which is very possible. I work on the “R&D” team for my company so we spend a lot of time researching and building PoC apps. I did one with Memgraph a few months ago after concluding it ought to outperform Neo4j, however I did not build the app with Neo4j to do a side by side comparison of performance. Both support Cypher so I wasn’t attached to one or the other, but I’ve always liked the idea of using in-memory stuff (like RAMDisk) to achieve extreme performance, and I figured at worst Memgraph would be “as fast” as Neo4j… that is 100% an assumption though and assumes that Memgraph is well-written. It sounds like it’s not though.
- anonbystander 4y agoTotally agree with doing your own benchmark, and when performance matters, work with someone who knows the systems I'm not a neo4j expert, and am not paid to write this. That said, their GDS subengine from the last couple of years appears to be distributed in-memory, essentially a view, and their year-over-year improvements there have been substantial. There might be no difference at the checkbox level. Likewise, when we did billion-scale work here with a variety of common queries, we found that the existence of basic features like indexes quickly changed what was fast vs slow. Historically, C++ vs Java is often < 2X of a difference, so when we're talking parallel & distributed hardware with tricky query planners & data representations... I have many questions beyond the language. If they were targeting something like FPGAs, I might feel differently.
- makeitdouble 4y agoHe didn't run the Memgraph one, only took their numbers https://news.ycombinator.com/item?id=34368362 https://news.ycombinator.com/item?id=34368362 I'd assume it's what the disclaimer at the top was about: that the code is in Python which he's not familiar enough with. The license bit on not having the right to integrate it if you're building a competing product might not be decisive but a good enough reason to not invest more time into running the code.
- mbuda 4y ago"egregious example" is unfair because there was and will be a huge effort in comparing different options. Ofc it's biased. Every single benchmark is biased towards something, but take a look at the specifications and what has actually being compared. 100% agree that everyone should run their own benchmark, that's not possible to do from the vendor perspective, it's just not possible for every single usage. Public benchmarks are something to look into if you want to get cheap info on how different vendors might do for you. Huge but is that it's just a guide, not the only piece of info you should base your decision on.
- 908B64B197 4y ago> Even supposedly gold-standard benchmarks like the TechEmpower framework benchmarks quickly devolve into "application server handling HTTP requests by responding with predefined strings", which is as fast as it's utterly useless in most people's version of the real world. It sets an upper bound on a server's performance given that page generation completes instantly. Sure it won't reflect real world performance, but in this case the benchmark should be read as "higher requests per second = lower resource footprint for the server". Engineering is about being able to understand what a benchmark or measure truly means, and what useable information it contains.