4 ms·
Around April 2020 I identified a project need for an "inverted search engine" - a system that would accept documents (recipe ingredient lines, like "three large
by jka 4y ago
Around April 2020 I identified a project need for an "inverted search engine" - a system that would accept documents (recipe ingredient lines, like "three large onions") as input, and would match those against a dataset of terms (ingredient names, like "tofu" or "tomato").
That would have been possible with a feature like percolation[1] in Elasticsearch, but I felt that the overhead of maintaining state (percolator queries) by using a network service would be excessive and that building an in-process alternative would be feasible.
The result is hashedixsearch[2], a pure-Python search engine library with support for stemming, synonyms and a few other features[3] to support the use-case.
It builds upon inverted index support provided by the hashedindex[4] library.
[1] - https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-percolate-query.html https://www.elastic.co/guide/en/elasticsearch/reference/curr...
[2] - https://pypi.org/project/hashedixsearch/ https://pypi.org/project/hashedixsearch/
[3] - https://github.com/openculinary/hashedixsearch/blob/6980ee633fd04f5b1354d542ce006ab9ed3bf149/tests/test_search.py#L19-L359 https://github.com/openculinary/hashedixsearch/blob/6980ee63...
[4] - https://github.com/michaelaquilina/hashedindex/ https://github.com/michaelaquilina/hashedindex/