6 ms·
Believe it or not, Dgraph has far and away the worst FTS of any graph or multi-model DB I evaluated. After speaking with Dgraph's engineers about it on their Sl
by staticautomatic 7y ago
Believe it or not, Dgraph has far and away the worst FTS of any graph or multi-model DB I evaluated. After speaking with Dgraph's engineers about it on their Slack channel-- for reasons left unsaid and which I'm pretty sure I would charitably describe as ridiculous-- that is apparently by design.
Like Couchbase, Dgraph has decided to use the Go library Bleve for FTS. Bleve is good at what it does, but it just doesn't do very much. The number and kinds of analyzers that Bleve has absolutely pale in comparison to Lucene. So for starters, it's just not all that great. Never mind that Bleve is pretty easy to extend. I don't want to reinvent the wheel writing analyzers that are freely available in Lucene, and I certainly don't want to have to deal with incorporating them back into my code base and testing them every time there's a new release.
But it gets worse. Unlike Couchbase, Dgraph doesn't even fully use Bleve. Rather than tracking Bleve releases and inheriting its analyzers, Dgraph has made the completely baffling decision to implement only a subset of them. It's already the case with Bleve, for example, that pretty much the only sentence tokenizer available is the unicode tokenizer. I don't want to use the unicode tokenizer anyway, but even if I wanted to do the tokenization myself, there's no straightforward way for me to get them into Dgraph because Bleve's "single token" analyzer (which just accepts a stream of individual tokens) is not one of the two or three analyzers Dgraph elected to incorporate.
As far as I'm concerned, that is some bullllll shit.
- mrjn 7y ago(Author of Dgraph) This is the first complaint I’ve heard about the full text index of Dgraph — are there examples where Dgraph’s FTS isn’t as good compared to others? If so, could you please file an issue on our GitHub repo, so we can investigate and bring it at par with what Elastic and others have to offer. Also, Dgraph allows custom indexers. So, you could build a custom FTS which would better fit the task at hand.
- staticautomatic 7y agoI appreciate you weighing in here, but with all due respect this is homework you would have already done if you were serious about your FTS being on par with Lucene, and that I shouldn't have to do for you. It's totally obvious just from comparing Elastic's documentation with Bleve's that Elastic has way more tokenizers and filters than Bleve. And it's also totally obvious from comparing Bleve's code to Dgraph's that Dgraph implements a subset of Bleve's. Whoever was responding to me on Slack sounded like they didn't even know what I was talking about. When I asked whether Dgraph planned to implement all of Bleve's tokenizers, the response I got was "Dgraph uses Bleve to generate the full-text tokens for the full-text index." When I pointed out that currently only some of them have been implemented and reiterated my question, the answer I got was "Dgraph's product decisions are independent of Bleve's features." Bleve would be a fine choice for FTS if you were planning on implementing the whole thing, writing additional analyzers to reach parity with Elastic, and preferably up-streaming them to Bleve. But if you're asking me to open a GitHub issue saying "will you please consider implementing the rest of Bleve?" the answer is no, thanks. I'll just revisit Dgraph some other time.