13 ms·
Xapiand: A fast, simple, modern search and storage engine
- merlincorey 8y agoStill beta, last release[0] was 0.8 on 2018 November 18. It is also not clear to me what, if anything, already integrates with this, and therefore how much code I need to write to try it out and compare against ElasticSearch. [0] https://kronuz.io/Xapiand/news/ https://kronuz.io/Xapiand/news/
- ploxiln 8y agofwiw also 11 patch releases (meaning 0.8.z) since then https://github.com/Kronuz/Xapiand/releases https://github.com/Kronuz/Xapiand/releases
- deleted 8y ago[deleted]
- drenvuk 8y agoIt's nice to see new search servers, especially low level ones. I'm going to give this a few tests.
- hardwaresofton 8y agoNot sure if you know about tantivy but it's cool too: https://github.com/tantivy-search/tantivy https://github.com/tantivy-search/tantivy
- cetra3 8y agoAlso worth mentioning is Toshi: https://github.com/toshi-search/Toshi https://github.com/toshi-search/Toshi Toshi is to ElasticSearch as Tantivy is to Lucene if that makes sense. Obviously as they are new they are not at feature parity, but Tantivy does win at some benchmarks: https://tantivy-search.github.io/bench/ https://tantivy-search.github.io/bench/
- mdnormy 8y agoNon-native English speaker. But isn't easier to understand like so, "Tantivity to Toshi, is as Lucene to Elasticsearch"
- hardwaresofton 8y agoAs a native english speaker, the earlier phrase ("tantivy is to toshi as lucene is to elastic search") is easier for me to understand. I find your phrase a bit harder to understand, but it looks like just the kind of reorganization other languages do structure wise -- I don't know how to express it in proper grammatical terms, but the way the prepositions are swapped around makes it seem like native english words but with a non-english structure. It might have to do with the use of Analogy questions in the SAT (a standardized test all but required for high school students wanting to attend good colleges in America), though it looks like they've been removed?[0]. "_____ is to ___ as ____ is to ______" was the verbatim format of those test questions. [0]: https://blog.prepscholar.com/sat-analogies-and-comparisons-why-removed-what-replaced-them https://blog.prepscholar.com/sat-analogies-and-comparisons-w...
- bmichel 8y agoThere is also Blast (golang), built on top of Bleve. - https://github.com/mosuka/blast https://github.com/mosuka/blast - http://blevesearch.com/ http://blevesearch.com/
- ddorian43 8y agoYeah but it's golang, so it's kinda like java, so I see no pros in it TBH.
- rakoo 8y agogolang isn't even close to using the same amount of memory as java, so at least there's that.
- hardwaresofton 8y ago
- nopacience 8y agoThere is apache lucy in C https://attic.apache.org/projects/lucy.html https://attic.apache.org/projects/lucy.html
- catmanjan 8y agoI'm interested, can anyone give a quick overview of why you'd use this over Elasticsearch?
- francislavoie 8y agoWell for one, Java is ridiculously memory hungry. The resource costs of Elasticsearch is the #1 reason I'm not using it. I've seen a few projects which had the aim of reimplementing the Elasticsearch backend in Rust, but were incomplete. That would be my ideal solution, personally.
- bheesham 8y agoI don't know how true that is anymore, that "Java is ridiculously memory hungry". It does power billions of devices, after all. :-P
- paradoxparalax 8y agoI was genuinely wondering recently why the Hackernews site search tool didn't show a very recent article with very obvious keywords, that google, for example, had in the first place in the results, when adding "Hn" to those 2 keywords in the search field. Is it a matter of the indexing, it means the article was too recent and it wasn't yet in the Algolia's(I believe) based HN's search tool memory; and in this case google copied it to memory faster; Or is this purely a matter of the Algorithms themselves? The algorithms for sure sound to be the matter, when the case is that of a search for an old article, that should be in memory already. It seems unnatural. Algolia has a free tier for open-source projects, what is very nice of them and thanks. I genuinely wonder if those algorithms are indeed so complex to justify those comparatively weaker behaviors seen at HN's internal search.
- latch 8y agoGoogle search with site:news.ycombinator.com (and optionally a time limit, which I wish wasn't limited to past hour/day/week/month/year) seems consistently superior to what Algolia provides. Algolia is YC company, so I assume that's the main reason it's being used. But that it does such an awful job with such a simply structured site isn't compelling.
- redox_ 8y agoHey latch, I've been working on the Algolia-based HN search and would love to improve it to provide you with a better search experience. Do you think about any specific improvements? Would you mind sharing with us some non-working queries? We can follow-up here and you can also open issues on https://github.com/algolia/hn-search https://github.com/algolia/hn-search
- latch 8y agoWell, if I do "activemq vs rabbitmq" I then switch to "comments", I get 2 results. The 2nd hit is reasonable: "ActiveMQ: Not ready for prime time" Google gives many more results, and a few on the first page seem quite relevant. Most notably: https://news.ycombinator.com/item?id=5531192 https://news.ycombinator.com/item?id=5531192 but also: https://news.ycombinator.com/item?id=1657574 https://news.ycombinator.com/item?id=1657574
- michelpp 8y agoXapian has a long history starting in the early 80s: https://xapian.org/history https://xapian.org/history I've used Xapian extensively, but not this new Xapiand tool, so I can only speak to the actual library. Xapian is a C++ library that accesses index data files directly on disk. There are bindings for various languages, say Python, let's you do 'import xapian' and get FFI bindings to the library, then you basically open your on disk index files and issue queries. Xapian supports many concurrent readers, but only one writer. It's not a server, there are no protocols. Maybe that's what this Xapiand tool adds. In general the overhead is very, very light, just enough ram to hold the library code, the OS takes care of all the filesystem level caching. Many of the very same concepts that are in Lucene, Documents, Terms, weights, flavors of BM25 relevance ranking, query parsing trees, relevancy operators, etc, all apply to Xapian as well.
- _wmd 8y agoI love Xapian, the quality of its recall is excellent and indexing performance very hard to find fault with. There's just a tiny problem - it's stuck with the GPL, despite a long effort to relicence the code going back years.
- Kronuz 8y agoMaybe xapian library just needs a little push from a larger community to make relicensing faster, Xapiand could help towards that end by helping brining in more people which can help. Xapiand source code is itself licensed as MIT (before compiling), and xapian community is already taking big steps towards relicensing.
- vkazanov 8y agoRemind me what's wrong with GPL, here? You can't repackage and resell it?
- ddorian43 8y agoYou can't include it in your code and sell your software without distributing the source.
- nopacience 8y agoThis will be interesting. Xapian is well known. Elasticsearch is lucene on the backend
- the_other_guy 8y agoI really hope this is the long awaited thing. From the number of commits, this project looks huge, Elasticsearch is really a big liability for resource limited deployments, I have seen some smaller projects made in Rust and Go but can't compete with Elasticseatch at any level, but this looks different and I hope it does.
- patelh 8y agoHaven't heard of Vespa? https://vespa.ai https://vespa.ai
- emmelaich 8y agoCan you tell us why you like it?
- patelh 8y agoHas been in production far longer than any other open source solution. Runs at scale across Yahoo, powering even Ad systems, with live configurations pushes. Everything you need for highly available product. It also has capabilities to be used for more complex uses cases around AI.
- manigandham 8y agoNobody has. There's no visibility or community around it which is a constant problem with Yahoo's open source projects. The only thing that really took off was Hadoop but there was very little back then. Vespa is also far more heavy and complex than any other search systems mentioned here.
- coleifer 8y agoSomewhat unrelated but I've written a restful search server that's powered by Sqlite's full-text search. Extremely lightweight python (flask) app. Nice for blogs or small projects, if I say so myself! https://github.com/coleifer/scout https://github.com/coleifer/scout
- nikolay 8y agoI am always surprised when I find out that developers with bold claims to fame have not heard of Sphinx [0] and of Xapian [1]! [0]: http://sphinxsearch.com/ http://sphinxsearch.com/ [1]: https://xapian.org/ https://xapian.org/
- jayalpha 8y agoapt-get install recoll Consider your problems solved...
- lodestone 8y agoI checked out Xapiand several months ago after I stumbled across it during a fit of Github browsing. It certainly seems fast and very easy to add documents but there is so little documentation that I was unable to test it out in any significant way. I'm very interested to see where the project goes, especially if Xapian itself switches away from the GPL.
- karterk 8y agoIf you are looking for an easy to run/manage, typo tolerant search engine, I've been working on this: https://github.com/typesense/typesense https://github.com/typesense/typesense
- amelius 8y agoFrom the features list: > Ranked search (so the most relevant documents are more likely to come near the top of the results list) with built-in support for multiple models from the Probabilistic, Divergence from Randomness, and Language Modelling families of weighting models. Custom user-supplied weighting models are also supported. Could someone explain in a little more detail what these terms mean?
- sciurus 8y agotl;dr is that those are different approaches to weighting documents in order to return the most relevant ones for a query. For an intro to the problem space, see https://opensourceconnections.com/blog/2014/06/10/what-is-search-relevancy/ https://opensourceconnections.com/blog/2014/06/10/what-is-se... If you want a lot more detail, check out the book Relevant Search. https://www.manning.com/books/relevant-search https://www.manning.com/books/relevant-search
- ngrilly 8y agoXapiand depends on xapian-core which is licensed under GPL (not LGPL). Makes me think that Xapiand should be licensed under GPL, instead of MIT?
- gdamjan1 8y agothe author can license his own source code however he likes. only, when distributing the compiled binaries, you're required to provide the whole sources under libre (gpl compatible) terms.