4 ms·
Searching with Riak Search
- boundlessdreamz 16y agoIf search performance and accuracy is the criteria of choosing the data store, how does riak+riak search compare with mysql+sphinx ?
- sjs 16y agoThe Sphinx website mentions scaling but just throws numbers out there without mentioning scaling out. Assuming it is distributed you'd have to run benchmarks with your data and queries to really know how it'll perform against Riak Search for any given number of nodes. If it's not distributed then Riak search can outperform it with enough nodes, without question. The question then becomes how many nodes. I'd love to see some benchmarks on this sort of thing if anyone is up to it.
- rb2k_ 16y agoIt is WAY easier to scale (up AND down) over several nodes. So if you've got BIG amounts of text that you have to fulltext-search, Riak might be the better option.
- bravura 16y agoAnd, for good measure, could you compare RiakSearch to horizontally scaling Lucene? And ElasticSearch, if you are familiar with it?
- samratjp 16y agoIMO, this is one of the better discussed comparisons on the whole Lucene sharding business: http://mail-archives.apache.org/mod_mbox/hbase-user/201006.mbox/%3C149150.78881.qm@web50304.mail.re2.yahoo.com%3E http://mail-archives.apache.org/mod_mbox/hbase-user/201006.m... EDIT: Also, keep an eye on Twitter's Lucene branch - http://engineering.twitter.com/2010/10/twitters-new-search-architecture.html http://engineering.twitter.com/2010/10/twitters-new-search-a...
- bravura 16y agoCould someone clarify on the RiakSearch scoring function for evaluating retrieved results? It appears that RiakSearch is modeled after Lucene in a variety of ways (https://wiki.basho.com/display/RIAK/Riak+Search https://wiki.basho.com/display/RIAK/Riak+Search): "At index time, Riak Search tokenizes a document into an inverted index using standard Lucene Analyzers. (For improved performance, the team re-implemented some of these in Erlang to reduce hops between Erlang and Java.)" "Search queries use the same syntax as Lucene, and support most Lucene operators including term searches, field searches, boolean operators, grouping, lexicographical range queries, and wildcards (at the end of a word only)." However, there is difference in the scoring function (https://wiki.basho.com/display/RIAK/Riak+Search+-+Querying https://wiki.basho.com/display/RIAK/Riak+Search+-+Querying): "Documents are scored using roughly the same formulas described here: http://lucene.apache.org/java/3_0_2/api/core/org/apache/lucene/search/Similarity.html http://lucene.apache.org/java/3_0_2/api/core/org/apache/luce... The key difference is in how Riak Search calculates the Inverse Document Frequency. The equations described on the /Similarity/ page require knowledge of the total number of documents in a collection. Riak Search does not maintain this information for a collection, so instead uses the count of the total number of documents associated with each term in the query." I am confused by this statement that they don't know "the total number of documents in a collection". If they were to say: "We don't use the document frequency (# of documents containing this term / total # documents) because we cannot compute (# of documents containing this term) over the entire corpus. Instead we estimate the document frequency using the term frequency (# of occurrences of this term in the corpus / total # terms)." that might be sensible. But I am not really clear what the current scoring function is.