8 ms·
Show HN: I built Haystack – your own google for scattered workplace knowledge
Hi all!
A few weeks ago I was scrolling through confluence pages trying to
find ssh connection details to our integration machine for 40 minutes
straight, later I discovered my co-worker slack'ed me the ssh
connection string two months ago.
So the same weekend I started working on haystack - a search engine
for workplace apps. that enables you to search slack, confluence,
jira, teams, sharepoint, github, and email in one place.
I wanted it to support natural language queries so a query like: "how
to connect to integ2 machine?" yields:
ssh -i private.pem ubuntu@ec2-integration2.eu-est-1.compute.amazonaws.com
I decided that user data should be stored locally, so all logic is
completely client-sided (including the NLP model) - I don’t want
access to your internal docs, thanks.
I rolled it out to my co-workers a week ago and they thought it's a
hit, so I'm planning on releasing it publicly on March 2023. But if you want to try it out before then it's
available here: https://haystack.it https://haystack.it. Thanks!
- mrmacha 4y agoAwesome idea and interesting approach! I am wondering how the search latentcy will be with your approach, especially for cases with more than a few hundreds documents. Do you have any insights about that?
- _vxw6 4y agoActually a few hundred documents is really no biggy, my current benchmarks is in the range of <250ms (instant feeling) for hundreds of thousands of paragraphs. I'm testing this on a large knowledge base.
- XCSme 4y agoWhy does your emoji (exploding head) on your landing page lead to an external page describing that emoji? https://emojipedia.org/exploding-head/ https://emojipedia.org/exploding-head/
- _vxw6 4y agoHaha I copied it from that page, hilliarious! I'll keep it haha
- XCSme 4y agoInteresting product, you should list all the integrations on the landing page.
- DarrenDev 4y agoKnowledge base for dev teams -- a problem we all have. I almost built a potential solution to this problem years ago but backed out. I'd love to see a solution that sticks, and to be wrong about this, but it feels very much like a Tarpit problem to me: https://www.youtube.com/watch?v=GMIawSAygO4 https://www.youtube.com/watch?v=GMIawSAygO4
- deleted 4y ago[deleted]
- amai 4y agoFind the needle in the haystacks: - https://haystacksearch.org/ https://haystacksearch.org/ - https://haystack.deepset.app/ https://haystack.deepset.app/ - https://www.haystackapp.io/ https://www.haystackapp.io/ - https://www.haystackteam.com/ https://www.haystackteam.com/ - https://thehaystackapp.com/ https://thehaystackapp.com/ - https://www.usehaystack.io/ https://www.usehaystack.io/
- _vxw6 4y agoyou forgot the most important one: haystack.it! I would argue that every known noun are the first domains to get registered. My goal is to associate workplace search engines with haystack.
- metaphor 4y agoThis[1] and this[2] also come to mind. [1] https://www.spglobal.com/engineering/en/products/haystack-gold.html https://www.spglobal.com/engineering/en/products/haystack-go... [2] https://www.plume.com/serviceproviders/platform/haystack/ https://www.plume.com/serviceproviders/platform/haystack/
- _vxw6 4y agoForgive me if I don't understand, but I don't think there's a problem with multiple companies using the same common noun as the base for their domain. Let the best product be remembered for the name.
- metaphor 4y agoCertainly not, presuming it doesn't escalate to the level of trademark/service mark infringement (to be fair, IANAL). Just a risk consideration...your product, your call. But I think there's value in at least recognizing that the namespace is quite crowded given the collisions that two interweb randoms were able to identify in short order.
- sz4kerto 4y agohttps://gethaystack.com/ https://gethaystack.com/ Same idea :)
- _vxw6 4y agoI’m adding some technical details! Haystack runs entirely client sided in the browser, so it has a unique tech-stack: Storage using IndexDB, haystack stores user indexes locally, + a compressed 90mb NLP model (t5-small) is stored. Indexing Locally in the browser, using a t5-small bi-encoder, and some parsing of documents happens in wasm. Search Query converted to embedding, then searched over index, atlast results are reranked with a t5-small based cross encoder, and top results go through a seq-to-seq transformer to produce a nice consise textual answer.
- mdmglr 4y agoDoes the app need to be open in the browser for the indexer to run?
- _vxw6 4y agoYes it needs to stay open. if that’s a problem, I thought of building an extension for continuous indexing.
- newman314 4y agoHere's a slightly different use case. I want something that can index all open tabs in the browser so that I do not have to leave hundreds of tabs open.
- _vxw6 4y agoThat’s extremely interesting, I would argue that the reason for keeping tabs open varies, but is something along the lines of: re-reaching the page in the tab is too slow
- newman314 4y agoMore specifically, I do a lot of context switching and still try to maintain a reasonable amount of open tabs. The problem that I have is that I can't remember days/months later if something I read is in an open tab or closed a while ago resulting in some frantic searches.
- d4rkp4ttern 4y agoTab grouping is one way to tame the madness. I’ve tried it with the default tab groups feature in Chrome but keep loosing the groupings whenever I restart chrome. Anyone know a better grouping extension?
- gamegoblin 4y agoThere are semantic embeddings libraries that are fast enough to run on entire webpages in ~hundreds of millis or low single digit seconds. I've been thinking a lot recently about making a browser plugin that simply does semantic embedding on a paragraph level of every web page I ever visit, and store it in a vector database. This would enable querying my little private search engine like "the HN story a few weeks ago that talked about ancient greek mining techniques" or "the reddit comment that had an analysis comparing Orwell's 1984 to the bible". For those not familiar, semantic embeddings take a chunk of text and embed it in a high dimensional vector space (~hundreds of dimensions) where semantically similar texts are closer together.
- raghavkhanna 4y agoDoes the index contain only info I have access to? How is authentication for all the knowledge sources handled?
- _vxw6 4y agoIn the setup process you sign in via SSO to all integrations, the token is saved in local storage. That token is used for indexing, and so if you don’t have access to info, the index doesn’t have access.
- pabue 4y agoInteresting idea and approach! Just a thought on the design: I really feel like the gold gradients make the whole site feel 'cheap', not trustworthy and not very 'professional'. Actually makes me not want to use it. Replacing them with a simple warm yellow improves this a lot. That might be just me, but maybe it's something for you to consider. Good luck!
- _vxw6 4y agoNot a designer, I appreciate this advice immensely, I’ll try!
- Winsaucerer 4y agoYep, specifically it looks to me like the kind of styling for a gambling or raffle website.
- MH15 4y agoHonestly it's not that bad. For someone who's not a designer it looks rather polished.
- Jeff_Brown 4y agoWhat determines relevance? It must do something other than page rank. Will it recognize synonyms and more subtle kinds of nearness in word-space?
- _vxw6 4y agoActually, the page rank is really based on semantic similarity and relevance of the matched paragraph. Which under the hood is based of a t5 encoder
- _vxw6 4y agoMakes me wonder, what kind of information sources do you use at work? Slack, teams, confluence or notion? airtable? jira?
- challenger-derp 4y agoCan see how this has potential. I've worked on a similar thing at work and there are some nuances as to the level the embedding is performed at (e.g. sentence level) and the kinds of queries that your search engine will be good at (i.e. good ranking of results). Also, depending on how heterogenous the data is, other factors (dated-ness, colleague whose writing/instructions/tutorials you prefer) can also be incorporated into the ranking algorithm.
- _vxw6 4y agoYou hit the nail on it’s head, some of these is something I’m dealing with right now (i.e dates)
- Game_Ender 4y agoHow do you compare to Glean? https://www.glean.com/ https://www.glean.com/ Not affiliated but just a happy user of their product it searches slack, confluence, jira, gmail, gdrive, github and source code all at once. With extras like Go links, verification, and some knowledge base features.
- _vxw6 4y agoopen & free for self hosted version, current alpha version is client side (runs in the browser).
- esperent 4y agoBy open do you mean open source?
- bluedevilzn 4y agoHow much does glean cost?
- thyrox 4y agoUnless there is some way to try it this post may be against Show HN guidelines specifically: > Show HN is for something you've made that other people can play with. HN users can try it out, give you feedback, and ask questions in the thread. > Off topic: Those can't be tried out, so can't be Show HNs. Make a regular submission instead. I'd suggest you change the title asap. (1) https://news.ycombinator.com/showhn.html https://news.ycombinator.com/showhn.html
- mdaniel 4y agoI seriously hope they don't submit this thing every 5 days until March https://news.ycombinator.com/item?id=34161085 https://news.ycombinator.com/item?id=34161085
- eru 4y agoWell, at least it got new text in the submission!
- _vxw6 4y agoHi, didn’t intend to repost this, it’s just that HN is hard to figure out. And most first posts don’t get the intended traction for various reasons (too long, bad wording, unclear). reposting and changing the post is totally allowed :)
- fzliu 4y agoAs someone who uses a multitude of workplace apps myself, this is amazing. What kind of model are you using, and do you have plans to provide a service based off this? Good stuff!
- wilg 4y agoSounds like a Searchable Library of All Corporate Knowledge
- jitl 4y agoThis is awesome! Is any of it open-source? I’d love to learn more about how the LLM works in the browser.
- _vxw6 4y agoIt will be very soon https://github.com/haystackoss/haystack https://github.com/haystackoss/haystack some rust code that compiles to WASM loads LLM from memory, and uses custom transformer.py like rust alternative we wrote.
- twobitshifter 4y agoHow long does the initial indexing take?
- VectorLock 4y agoSeems quite skinny on info about indexing. It nabs your login and stores tokens to the sites you want to index then... does it just spend CPU days running in your browser downloading and indexing all that data? How much storage do you need locally to index what can potentially be massive amounts of data in most corporate information sources?
- _vxw6 4y agogb's of storage, potentially 10+ for very large datasets. Minutes, not days. Very big data sets might take 30+ minutes (or even a couple of hours), but usefulness starts in the first few minutes (because of the priority algorithm)
- derekhingeveld 4y agoReally looking forward to trying out your app, it looks amazing.
- _vxw6 4y agoThanks
- anonymous344 4y agowhere does this exactly store the data? and is it manageable?
- _vxw6 4y agoIt stores it in browser local storage using IndexDB, if you have access to years of documents, it might take up to 10+gb of persistant storage.
- rco8786 4y ago> Browser based LLM Very cool - any more info on that?
- _vxw6 4y agoSame bundle of weights but is being run by some rust code that is compiled down to WebAssembly.