4 ms·
Can you do fuzzy search with it?
by spa5k 2y ago
Can you do fuzzy search with it?
- jbaiter 2y agoNot by default, no. But you could kind of implement it by providing a custom tokenizer that emits multiple terms for the same position in the document, with different variants of the same token. This would not be "proper" fuzzy search, but might be enough depending on the use case. See https://www.sqlite.org/fts5.html#synonym_support https://www.sqlite.org/fts5.html#synonym_support for more details on the different approaches for implementing synonyms in custom tokenizers.
- spa5k 2y agoI don't think that will work for me, since I needed something that can handle mistakes in the words, like Du'ha to duha etc, Rahman to rehman, basically whatever looks closest.
- jbaiter 2y agoOne thing you could do: FTS5 has the `fts5vocab` virtual table [1] that has all the terms. You could provide a user-defined function that computes the levenshtein distance between your query terms and the terms in that table, obtain candidate terms that way and build a big query that searches for all those lexically close terms. [1] https://www.sqlite.org/fts5.html#the_fts5vocab_virtual_table_module https://www.sqlite.org/fts5.html#the_fts5vocab_virtual_table...
- jerrygenser 2y agoI can confirm an approach like this works in practice. Although instead of levenshtein I use spellfix (maybe it uses that under the covers? not sure). If there is no match from the first search, I use the sqlite spellfix extension [0] to find matches. Then feed those candidates into the terms. https://www.sqlite.org/spellfix1.html https://www.sqlite.org/spellfix1.html
- jbaiter 2y agoAh, that's great, didn't know about that :-) It seems to use a similar edit-distance algorithm under the hood, and the docs explicitely mention integration with the FTS extension, so this is probably the way to go!
- deleted 2y ago[deleted]