4 ms·
Regarding under the hood details, getting haystack up and running in the browser was quite a challenge, but here's what I ended up using: - Storage: IndexDB
by _vxw6 4y ago
Regarding under the hood details,
getting haystack up and running in the browser was quite a challenge, but here's what I ended up using:
- Storage: IndexDB browser API for local browser storage, it stores read permission tokens to apps, the AI models and the document index.
- Indexing: fine-tuned TinyBERT-based bi-encoder for indexing documents, messages and emails.
- Searching: cosine similarity between query embedding and the built index, and then rerank using tuned TinyBERT cross-encoder.
- Building useful search resuts: search result building involved fine-tuning a t5-small model for summarization and text regression.
- Performance nodejs->browser js adaptations, wasm rewrites in rust for performance.
- mritchie712 4y agoNice! Do you only stored documents I visit in the browser (e.g. a specific slack thread / JIRA issue) or are you querying all the API's for each supported app?
- _vxw6 4y agoSo I'm specifically querying all the APIs of the applications you connected, and indexing all documents and threads in local browser storage.