2 ms·
Interesting DB for feature storage and LSH is good choice I believe. I'm wondering why the tight link to pytorch C++ tensors (under refactoring actually), bit I
by pilooch 8y ago
Interesting DB for feature storage and LSH is good choice I believe. I'm wondering why the tight link to pytorch C++ tensors (under refactoring actually), bit I haven't looked at the euclidendb code yet. Thanks for sharing !
Those interested can also find an open source integration of lmdb + annoy here: https://github.com/jolibrain/deepdetect/blob/master/src/simsearch.cc#L188 https://github.com/jolibrain/deepdetect/blob/master/src/sims...
This the underlying support for similarity search based on embeddings, including images and object similarity search, see https://github.com/jolibrain/deepdetect/tree/master/demo/objsearch https://github.com/jolibrain/deepdetect/tree/master/demo/obj...
This is running for apps such as a Shazam for art, faster annotation tooling and text similarity search.
Annoy only supports indexing once, while hnwlib supports incremental indexing, something I'm looking at.
- dtjohnnyb 8y agohttps://github.com/nmslib/hnswlib https://github.com/nmslib/hnswlib for anybody else googling this library
- perone 8y agoWe'll be integrating other indexing in near future (such as faiss), Annoy is just one option for indexing that was implemented. Each indexing method will have their pros/cons, so you'll be able to select the search engine backend according to your restrictions. There are many reasons why we depart from other libraries, many of them, for instance, uses JSON+base64 (http/1) for serialization, while we use protobuf+gGRPC (http/2).