30 ms·
Sure - there's no doubt that lucene is a much bigger and more "enterprise" solution that is used professionally today by many big companies from Elasticsearch t
by JetSetWilly 5y ago
Sure - there's no doubt that lucene is a much bigger and more "enterprise" solution that is used professionally today by many big companies from Elasticsearch to Mongo to Solr.
But for core text search and indexing - tantivy does indeed support tokenization, stemming, fuzzy matching, and so on. For the core use case of performing a text search on a large corpus of text a lot of the functionality one would need is there - as you can see if you look at the list of queries in the benchmark. From the list of queries there's little I would miss from our "enterprise" use of lucene today - speaking for myself.
One might argue "oh but if you add some obscure features that 5% of people use like a real enterprise solution then it will slow down by 2x" - but I doubt it.
To me, tantivy and the likes are a technology demonstrator - they show that lucene could be significantly faster if it wasn't in java.
- jillesvangurp 5y agoThe whole point of a search engine is getting you good results. Not alright result, or some random results in the wrong order, but the best possible results. Bench marking relevance is much more important than bench marking performance for search. Lucene is a Swiss army knife for building you something to get you the best results possible. Those aren't enterprise features if you actually care about what your users do with search results. For example because your sales are directly correlated to that. The trick with Lucene is to do as much as you can possibly get away with as opposed to doing as little as you can do to call it a day that seems to be the strategy here for proving X is faster than Y with this particular benchmark. It's faster because it does less useful work. The proof of concept here would be building something that is as good and as fast. By not even trying to benchmark how good things are, you kind of make the point that you either don't know or don't care enough to even bother to measure it.
- JetSetWilly 5y agoWhy did you offer lucene as a counterpoint, if you deny that there’s any equivalent implementation elsewhere that isn’t lucene? That just means lucene cannot be a counterpoint if it just so happens that the only implementation is in java.
- jillesvangurp 5y agoBecause people have been trying pretty hard to replicate what Lucene does in other languages specifically because they thought they could do a better job. The reason the Java implementation continues to dominate is that it being implemented in Java has repeatedly proven to be not as much of an issue as people assume it to be. At least not enough to matter. And since the article is called "why C is faster than Java" and the main argument is literally "here's this thing implemented in C that is fast", it's an excellent counterpoint to go "here's this other thing in Java that is fast".
- JetSetWilly 5y agoSure, I would agree that lucene is “fast enough”. But that doesn’t mean it couldn’t be faster if implemented in another language, and I believe that tantivy demonstrates that it could be. A historical accident isn’t a counterpoint.
- zelphirkalt 5y agoThis however, throws out all productivity, that a language like Java gives developers in contrast to C, plus memory management. The C equivalent most likely would be a buggy version with the typical memory management issues, or consume much more time to create. I am not a big Java fan. We are not talking a Haskell or Rust or similarly safe language here. However, Java is still far ahead of C in that regard.
- JetSetWilly 5y agoFrom what I’ve seen working in the financial industry, hardly anybody uses lucene directly. It is typically behind elasticsearch or solr or whatever - which then takes care of load balancing and replication and all that. For these use cases, you could swap out lucene for something faster or more “low level” in implementation without affecting the typical application using it much, as all interaction happens across some rest API.
- zelphirkalt 5y agoThe point I am making is, that you would have to write that lower level or somehow different thing first. If you do it in C, then how do you make it as safe, while still having the same functionality? In huge projects this has mostly not been achieved. Large C projects (and C++ project) suffer from memory safety issues. This seems to be the general rule. Even the best make mistakes, when working with big C code bases, or there are so few of those "best", that they cannot possibly create a huge project all by themselves. If you replace a more or less safe implementation, with a buggy, memory unsafe C version of it, you are not going to make anyone happy about the performance improvement. Always assuming, that you put in "only" the same amount of time and number of people, not more, than the original implementation. The amount of things you need to keep in mind with all the memory management and other details when writing C code can also make you blind to other issues like proper input validation and implementation of invariants of application settings. You have to put a lot of energy into making sure you are not leaking memory and avoiding memory safety issues, so you lack that energy when it comes to higher level issues. This is directly affecting developer productivity.