3 ms·
Does that mean that you can't query substrings or do fuzzy searches?
by SwiftyBug 1y ago
Does that mean that you can't query substrings or do fuzzy searches?
- panic 1y agoIf you want to adapt the technique to full-text search, you can index trigrams instead of full keywords.
- dangoodmanUT 1y agohaha you beat me to it! yes tokenize with trigrams is a very simple way to get this functionality. That's how systems like postgres has historically done it
- dangoodmanUT 1y agoThe BloomSearchEngine takes a TokenizerFunc so you can determine how JSON values are tokenized (that's why each path always returns an array of strings). The default tokenizer is a a whitespace one: https://github.com/danthegoodman1/bloomsearch/blob/148a7996737b586001c140c507884011b9415940/tokenizer.go#L87-L98 https://github.com/danthegoodman1/bloomsearch/blob/148a79967... So {"name": "John Smith"} is tokenized to [{Path: "name", Values: ["john", "smith"]}], and the bloom filters will store: - field: "name" - token: "john" - token: "smith" - fieldtoken: "name:john" - fieldtoken: "name:smith" The same tokenizer must be used at query time too. Fuzzy searches and sub-word searches could be supported with custom tokenizers (eg trigrams, stemming), but it's more generally targeting the "I know some exact subset of the record, I need all that have this exactly" searches
- bonobocop 1y agoNot OP, but to me, this reads fairly similar to how ClickHouse can be set up, with Bloom filters, MinMax indexes, etc. A way to “handle” partial substrings is to break up your input data into tokens (like substrings split in spaces or dashes) and then you can break up your search string up in the same way.