39 ms·
That's just because there's no a lucene equivalent C library with the same level of attention? however, there are increasingly such written in C++ (pisa) and
by JetSetWilly 5y ago
That's just because there's no a lucene equivalent C library with the same level of attention?
however, there are increasingly such written in C++ (pisa) and rust (tantivy). They handily beat lucene in benchmark suites [1] - so it seems like lucene does suffer from a java penalty - despite getting even more developer attention than pisa and tantivy I would think.
1: https://tantivy-search.github.io/bench/ https://tantivy-search.github.io/bench/
- jillesvangurp 5y agoNo, because a lot of the work would end up being the same kind of work with no inherent advantages of one over the other. Easy to predict because there used to be a C implementation of Lucene. It couldn't keep up in terms of features or performance so work on that stopped a long time ago. Most of the libraries you mention don't come close to even implementing a tiny portion of Lucene. Classic case of apples and oranges. Also, good examples of the niche things I was talking about. Benchmarks like you mention are kind of self serving like that. They measure something but not everything and probably very selectively. They are faster at what exactly? Under what circumstances? Why? As soon as you answer those questions, what inherent limitations do the Lucene developers have replicating that? A lot here boils down to how the underlying search engine implements tokenization, stemming, language analyzers, fuzzy matching, and a few other things that you'd need to build a search engine that doesn't suck. The benchmark conveniently does not specify any of that; presumably because it lacks many of those features or has extremely naive implementations of those things. That would be the kind of corner cutting I was talking about. What kind of relevance ranking is being used here? How good is it? Was that even evaluated or considered? Hint, Lucene gives you many options here. Search quality and performance are the big trade off here. Anything that trades off quality over performance is going to look good until you look at the quality. Also things like the index size and document volume are not specified. Or the hardware this ran on. Or the JVM configuration, compiler flags, etc. It's like benchmarking a formula 1 car by how quick it is at parallel parking. Yes, a T-Ford is going to be faster at that maybe. But is that even meaningful to look at?
- JetSetWilly 5y agoSure - there's no doubt that lucene is a much bigger and more "enterprise" solution that is used professionally today by many big companies from Elasticsearch to Mongo to Solr. But for core text search and indexing - tantivy does indeed support tokenization, stemming, fuzzy matching, and so on. For the core use case of performing a text search on a large corpus of text a lot of the functionality one would need is there - as you can see if you look at the list of queries in the benchmark. From the list of queries there's little I would miss from our "enterprise" use of lucene today - speaking for myself. One might argue "oh but if you add some obscure features that 5% of people use like a real enterprise solution then it will slow down by 2x" - but I doubt it. To me, tantivy and the likes are a technology demonstrator - they show that lucene could be significantly faster if it wasn't in java.
- jillesvangurp 5y agoThe whole point of a search engine is getting you good results. Not alright result, or some random results in the wrong order, but the best possible results. Bench marking relevance is much more important than bench marking performance for search. Lucene is a Swiss army knife for building you something to get you the best results possible. Those aren't enterprise features if you actually care about what your users do with search results. For example because your sales are directly correlated to that. The trick with Lucene is to do as much as you can possibly get away with as opposed to doing as little as you can do to call it a day that seems to be the strategy here for proving X is faster than Y with this particular benchmark. It's faster because it does less useful work. The proof of concept here would be building something that is as good and as fast. By not even trying to benchmark how good things are, you kind of make the point that you either don't know or don't care enough to even bother to measure it.
- JetSetWilly 5y agoWhy did you offer lucene as a counterpoint, if you deny that there’s any equivalent implementation elsewhere that isn’t lucene? That just means lucene cannot be a counterpoint if it just so happens that the only implementation is in java.
- dig1 5y ago> there's no a lucene equivalent C library with the same level of attention? Actually, there is [1]. Clucene has been around for years, it is a re-implementation of Lucene in C++, and yet the only serious usage I've seen was it as a file indexing engine in KDE many years ago (although I might miss other projects). Later it was dropped from KDE. I've had (indirect) relations with search/indexing stuff for years, and no one ever complained about Lucene's speed or had any problems with it. [1] http://clucene.sourceforge.net/ http://clucene.sourceforge.net/
- chx 5y agoFor C++ wouldn't Manticore, the Sphinx Search fork be a better example?
- spaetzleesser 5y agoThere used to be a C Lucene version. I used it for a desktop software. It was way faster than the Java version but wasn't maintained well.