Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
yuhongsun
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Lessons from building the best Deep Research and how you can build better agents
(onyx.app)
2 points
by
yuhongsun
8mo ago
|
1 comments
2.
▲
Hacking OpenAI's Internet Search
(onyx.app)
1 points
by
yuhongsun
1y ago
|
0 comments
3.
▲
Onyx (YC W24) – AI Assistants for Work Hiring Founding AE
(ycombinator.com)
1 points
by
yuhongsun
1y ago
4.
▲
Onyx (YC W24) Is Hiring for ML Engineer
(ycombinator.com)
1 points
by
yuhongsun
1y ago
5.
▲
Onyx (YC W24) Is Hiring
(ycombinator.com)
1 points
by
yuhongsun
2y ago
6.
▲
by
yuhongsun
2y ago
RAG is a tool for the deep research agent to use in finding all of the context it needs. Deep research can call the search many times and reflect on the results of the previous searches and then search for other things as needed. Deep resea
7.
▲
by
yuhongsun
2y ago
Yup, hopefully with Onyx, the folks who have these questions can just fire off a query with agent mode turned on and the LLM will research the relevant tree of knowledge and come back with an answer in the fraction of the time it would take
8.
▲
by
yuhongsun
2y ago
We have a dataset that we use internally to evaluate our search quality. It's more representative of our use case since it contains Slack messages, call transcripts, very technical design docs, company policies which is pretty differen
9.
▲
by
yuhongsun
2y ago
Assuming self-hosting, data is processed within the deployment with local deep learning models for embedding, identifying low information documents, etc. A hybrid keyword/vector index is built locally within the deployment as well. At
10.
▲
by
yuhongsun
2y ago
This is a large challenge in itself actually. Every external tool has it's own framework for permissions (necessarily so). For example, Google Drive docs have permissions like "global public", "domain public", "
11.
▲
by
yuhongsun
2y ago
It's like how OpenAI's deep research works by searching the internet, ours works by searching over our "RAG" system that indexes company documents.
12.
▲
by
yuhongsun
2y ago
Quite a lot to cover here! So in addition to the typical RAG pipeline, we have many other signals like learning from user feedback, time based weighting, metadata handling, weighting between title/content, and different custom deep lea
13.
▲
by
yuhongsun
2y ago
On privacy and security, we are the only option (as far as I know) that you can connect up to all your company internal docs and have it be all processed locally to the deployment and stored at rest within the deployment. So basically you c
14.
▲
by
yuhongsun
2y ago
Amazing to hear from a happy user! Thanks for the kind words!
15.
▲
by
yuhongsun
2y ago
Before sharing how it works, I want to highlight some of the challenges of a system like this. Unlike deep research over the internet, LLMs aren’t able to easily leverage the built in searches of these SaaS applications. They each have diff
16.
▲
Show HN: Open-source Deep Research across workplace applications
(github.com)
125 points
by
yuhongsun
2y ago
|
30 comments
17.
▲
Show HN: Danswer APIs – Open-source APIs for building RAG apps over company docs
(github.com)
3 points
by
yuhongsun
2y ago
|
0 comments
18.
▲
Ask HN: How to run a startup sponsored competition?
2 points
by
yuhongsun
2y ago
|
0 comments
19.
▲
by
yuhongsun
2y ago
Hey everyone, I’m one of the authors of the post. As an engineer at small companies, I’ve only ever heard about ERPs but never fully understood them. They seemed like clunky, expensive software that only large, non-technical companies use.
20.
▲
A Brief History of ERP
(danswer.ai)
3 points
by
yuhongsun
2y ago
|
2 comments
21.
▲
Danswer (YC W24) Is Hiring Founding Full Stack Engineer
(workatastartup.com)
1 points
by
yuhongsun
2y ago
22.
▲
by
yuhongsun
2y ago
Danswer AI | https://www.danswer.ai/ | Full Stack Engineer, Customer Engineer | On-site | $150,000 - $240,000 Danswer is an open-source GenAI assistant with organization specific knowledge. Danswer is focused on connecting
23.
▲
by
yuhongsun
3y ago
We have a connector interface and build guide for contributors: https://github.com/danswer-ai/danswer/blob/main/backend/dans... Should be not too bad to build one out! Fun fact, more than half the c
24.
▲
by
yuhongsun
3y ago
Ya, I haven't dug too deep into their project, to be honest. But we do think it's great more teams are going for open source. I tried to look up their RAG pipeline to see if they've invested effort into building a strong retr
25.
▲
by
yuhongsun
3y ago
There's a question by cpach further up on this page which is essentially asking this but also with some additional questions. Hopefully you find that thread useful! TLDR: There's the cloud version and also there are a set of nice
26.
▲
by
yuhongsun
3y ago
Are you referring to the approach of creating hypothetical questions for embedding along with the document. Or more along the lines of creating summaries of the documents during indexing and embedding those as well? Either case, the reason
27.
▲
by
yuhongsun
3y ago
Two other tidbits on this: 1. There's a difference between relevance and usefulness that cross-encoders cannot capture. Imagine a thread with a bunch of people complaining about an exception and each comment is another mention of the e
28.
▲
by
yuhongsun
3y ago
Hi, it may be an indexing issue. There's a precanned "Information not found" message that we show if document retrieval failed. A couple common causes for this are: - Not provisioning enough resources and processes are dying
29.
▲
by
yuhongsun
3y ago
Thanks for the kind words! There's nothing more motivating in the world than hearing that our users are loving what we've built!
30.
▲
by
yuhongsun
3y ago
Thanks for the kind words! Sorry for the delayed response, this post drew a lot more interest than anticipated and we've been swamped working with new folks coming in. Regarding: - PrivateGPT: they're for individual use where you
More ›