5 ms·
I have spent many days searching for the best JavaScript implementation of full text search that handles typos (substitutions) well. Implementing a good indexin
by molf 4y ago
I have spent many days searching for the best JavaScript implementation of full text search that handles typos (substitutions) well. Implementing a good indexing algorithm for this is not easy. In particular if you are indexing large amounts of text (documents) instead of short strings.
I settled on MiniSearch. [0] It is fast & small enough and fairly feature complete.
Afterwards I made a few contributions to improve performance and implement a better scoring algorithm. So I'm probably a bit biased now. Take my recommendation with a grain of salt.
Personally I think that OP's library does not perform searches, fuzzy or otherwise. It's much more similar to 'grep'. Try searching for "mario adventures". It won't actually find the most obvious results, because the order of the keywords in the search string must match the order of the keywords in the indexed text.
[0]: https://github.com/lucaong/minisearch https://github.com/lucaong/minisearch
[1]: https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uFuzzy,MiniSearch&search=mario%20adventures https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uF...
- leeoniya 4y ago> It won't actually find the most obvious results, because the order of the keywords in the search string must match the order of the keywords in the indexed text. i'm really confused why people dont bother actually reading anything i spent so much time distilling (both in the readme and the short instructions in this submission itself), which explicitly spell out how to adjust the necessary options precisely for what you're asking. simply toggling outOfOrder returns perfectly good results: https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uFuzzy,MiniSearch&outOfOrder&search=mario%20adventures https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uF...
- molf 4y agoA blunt answer to your question is: I didn't take the time to read everything. Sorry for missing that feature. I stand by my overall point though. There is more to search than finding matches. Handling real world typing errors is one (try searching for "mario avdentures" or "mario adventutes"). Ordering results by relevance is not trivial either. And it does not appear this library intends to tackle that problem completely. Relevance scoring should take into account the term frequency and the document/field length, as well as average length. I don't mean to discount your project though. There are many distinct use cases related to searching and filtering, and power to you for solving it in a way that works for you (and I'm sure for others as well). I just wanted to share my experience in exploring the space of full text search libraries that handle long form text and typos gracefully.
- leeoniya 4y ago> A blunt answer to your question is: I didn't take the time to read everything. s/everything/anything this is by no means meant to replace fulltext search, with term omission tolerance, typos, stemming, etc. fwiw, i wrote this to replace a frontend search strategy that was originally based on MiniSearch and [apparently] gave underwhelming results and/or performance. since you seem to be familiar with MiniSearch, perhaps you can help improve the result ordering for this: https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uFuzzy,MiniSearch&search=super%20ma https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uF... this is my fundamental issue with fulltext searches. even if the result order could be made sane, how do you find the "junk cutoff" (there clearly is one). at least it does seem to place the best results on top, though ordered haphazardly, which is also true of FlexSearch. boiling things down to a single relevance score, and making the user tweak various boosting knobs to nudge things mostly [but not perfectly] into place seems like a dogma that most of these engines suffer from.
- molf 4y agoMany search engines default to returning hits when at least 1 keyword match. Not returning any results just because not every single keyword could be matched leads to a poor user experience. However, good relevance scoring is key here. Hits that match most keywords should be somewhere at the top. The reasoning is that a user will try to find the result they are looking for in the top hits. In the case of MiniSearch you can change this default and configure it to only return hits if all keywords match [0]. Searching for "super ma" seems like an autocomplete query, for which most users have subtly different expectations than regular search. MiniSearch has a separate method for that, essentially baking in different default settings. [1] [0] https://lucaong.github.io/minisearch/classes/_minisearch_.minisearch.html#combine-with-and https://lucaong.github.io/minisearch/classes/_minisearch_.mi... [1] https://lucaong.github.io/minisearch/classes/_minisearch_.minisearch.html#autosuggest https://lucaong.github.io/minisearch/classes/_minisearch_.mi... Edit: > s/everything/anything No need for the sneer.
- leeoniya 4y agook, thanks for the advice. i'll see if the MiniSearch settings can be improved for this demo without drastically affecting the quality of the "search" case.