Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
_vxw6
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
_vxw6
4y ago
Haha I copied it from that page, hilliarious! I'll keep it haha
2.
▲
by
_vxw6
4y ago
It stores it in browser local storage using IndexDB, if you have access to years of documents, it might take up to 10+gb of persistant storage.
3.
▲
by
_vxw6
4y ago
gb's of storage, potentially 10+ for very large datasets. Minutes, not days. Very big data sets might take 30+ minutes (or even a couple of hours), but usefulness starts in the first few minutes (because of the priority algorithm)
4.
▲
by
_vxw6
4y ago
Actually a few hundred documents is really no biggy, my current benchmarks is in the range of <250ms (instant feeling) for hundreds of thousands of paragraphs. I'm testing this on a large knowledge base.
5.
▲
by
_vxw6
4y ago
Same bundle of weights but is being run by some rust code that is compiled down to WebAssembly.
6.
▲
by
_vxw6
4y ago
It will be very soon https://github.com/haystackoss/haystack some rust code that compiles to WASM loads LLM from memory, and uses custom transformer.py like rust alternative we wrote.
7.
▲
by
_vxw6
4y ago
Forgive me if I don't understand, but I don't think there's a problem with multiple companies using the same common noun as the base for their domain. Let the best product be remembered for the name.
8.
▲
by
_vxw6
4y ago
Not really, gethaystack searches your tabs
9.
▲
by
_vxw6
4y ago
Sorry didn’t assume otherwise, replied accidentally to the wrong thread
10.
▲
by
_vxw6
4y ago
Note taken lol wow
11.
▲
by
_vxw6
4y ago
Thanks
12.
▲
by
_vxw6
4y ago
I’m planning on releasing an open source version of this. But also have paid features that managers would like to use.
13.
▲
by
_vxw6
4y ago
I’m using a fine-tuned t5-small model, I fune-tuned it for two tasks, question answering from a paragraph, and highlighting relevant text of search results.
14.
▲
by
_vxw6
4y ago
The tricky part is to understand which parts are you indexing, paragraph vs. sentences vs. pages. But yeah
15.
▲
by
_vxw6
4y ago
I load the 90mb model into memory from IndexDB, and then some rust code compiled down to WASM does calculations with the model.
16.
▲
by
_vxw6
4y ago
Hi, didn’t intend to repost this, it’s just that HN is hard to figure out. And most first posts don’t get the intended traction for various reasons (too long, bad wording, unclear). reposting and changing the post is totally allowed :)
17.
▲
by
_vxw6
4y ago
open & free for self hosted version, current alpha version is client side (runs in the browser).
18.
▲
by
_vxw6
4y ago
You hit the nail on it’s head, some of these is something I’m dealing with right now (i.e dates)
19.
▲
by
_vxw6
4y ago
Makes me wonder, what kind of information sources do you use at work? Slack, teams, confluence or notion? airtable? jira?
20.
▲
by
_vxw6
4y ago
Very good explaination! hundred of millis is what I experienced.
21.
▲
by
_vxw6
4y ago
Actually, the page rank is really based on semantic similarity and relevance of the matched paragraph. Which under the hood is based of a t5 encoder
22.
▲
by
_vxw6
4y ago
Not a designer, I appreciate this advice immensely, I’ll try!
23.
▲
by
_vxw6
4y ago
In the setup process you sign in via SSO to all integrations, the token is saved in local storage. That token is used for indexing, and so if you don’t have access to info, the index doesn’t have access.
24.
▲
by
_vxw6
4y ago
That’s extremely interesting, I would argue that the reason for keeping tabs open varies, but is something along the lines of: re-reaching the page in the tab is too slow
25.
▲
by
_vxw6
4y ago
Yes it needs to stay open. if that’s a problem, I thought of building an extension for continuous indexing.
26.
▲
by
_vxw6
4y ago
I’m adding some technical details! Haystack runs entirely client sided in the browser, so it has a unique tech-stack: Storage using IndexDB, haystack stores user indexes locally, + a compressed 90mb NLP model (t5-small) is stored. Indexi
27.
▲
by
_vxw6
4y ago
you forgot the most important one: haystack.it! I would argue that every known noun are the first domains to get registered. My goal is to associate workplace search engines with haystack.
28.
▲
Show HN: I built Haystack – your own google for scattered workplace knowledge
(haystack.it)
96 points
by
_vxw6
4y ago
|
67 comments
29.
▲
by
_vxw6
4y ago
What was the name of your product hmu hey@haystack.it, would love to learn from your experience. It's self-hostable, doesn't store your data, more focused on developers and bottom up motions, I want you to be up and running within
30.
▲
by
_vxw6
4y ago
This is heart warming to hear! Happy new year!
More ›