Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jeffchuber
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
jeffchuber
1y ago
https://x.com/jeffreyhuber/status/1732069197847687658
32.
▲
by
jeffchuber
1y ago
id be super curious to see if Chroma (written in Rust) can work better here!
33.
▲
by
jeffchuber
1y ago
This is still retrieval and RAG, just not vector search and indexing. it’s incredibly important to be clear about terms - and this article does not meet the mark.
34.
▲
Building a usage-based billing system
(trychroma.com)
2 points
by
jeffchuber
1y ago
|
0 comments
35.
▲
by
jeffchuber
1y ago
the table with comparable models is a really great way to show off things here
36.
▲
by
jeffchuber
1y ago
enterprise GTM has its own set of challenges and needs and warrants someone really focused on it
37.
▲
by
jeffchuber
1y ago
I’m Jeff, co-founder of Chroma. We build the most popular open-source AI vector database. When people use Chroma, the first question they ask is which embedding model to use. This choice affects how your RAG application will perform in prod
38.
▲
Show HN: Generative Benchmarking for RAG
(research.trychroma.com)
4 points
by
jeffchuber
1y ago
|
1 comments
39.
▲
by
jeffchuber
2y ago
This will work very poorly when your data is changing because the centroids degrade and you'll have very poor recall but likely not know it unless you are also monitoring recall. I didn't see this in the write-up, so adding it her
40.
▲
Chroma is now 4x faster, powered by Rust
(trychroma.com)
3 points
by
jeffchuber
2y ago
|
1 comments
41.
▲
by
jeffchuber
2y ago
chroma's api i think takes something like 20 seconds
42.
▲
by
jeffchuber
2y ago
Rest in peace Marshall
43.
▲
Designing a Query Execution Engine
(trychroma.com)
41 points
by
jeffchuber
2y ago
|
0 comments
44.
▲
by
jeffchuber
2y ago
i know it is the case in chroma this is supported out of the box with 0 lines of code. i’m pretty sure it’s supported everywhere else in no more than 3 lines of code.
45.
▲
by
jeffchuber
2y ago
> Vector databases treat embeddings as independent data, divorced from the source data from which embeddings are created With the exception of Pinecone: Chroma, Qdrant, Weaviate, Elastic, Mongo, and many others store the chunk/docum
46.
▲
by
jeffchuber
2y ago
pg_vector does post-filtering, not pre-filtering
47.
▲
by
jeffchuber
2y ago
its not
48.
▲
by
jeffchuber
2y ago
it’s fully open source, apache 2.0 and in the mono repo today. a distributed database is naturally has more complexity, but we’ve put a lot of effort in to make it as easy as possible to run.
49.
▲
by
jeffchuber
2y ago
Thanks ChatGPT! Yes - that's a great explanation.
50.
▲
by
jeffchuber
2y ago
Colbert is great! - Check out https://github.com/AnswerDotAI/RAGatouille by the excellent https://x.com/bclavie Relatedly ColPali ( https://arxiv.org/abs/2407.01449 ) is gaining a t
51.
▲
Retrieval powered by object storage: AMA
15 points
by
jeffchuber
2y ago
|
7 comments
52.
▲
by
jeffchuber
2y ago
most practitioners i have found are converging on the desire to have strong explainability and steer-ability of context - which means not YOLO dumping in 2M tokens - but we will see
53.
▲
Evaluating Chunking Strategies for Retrieval
(research.trychroma.com)
7 points
by
jeffchuber
2y ago
|
2 comments
54.
▲
by
jeffchuber
3y ago
chroma (single-node) doesn't use sqlite for vector - it's for metadata search.
55.
▲
by
jeffchuber
3y ago
https://pulley.com/ is great
56.
▲
by
jeffchuber
3y ago
Lem is wonderful and this is why we should call large models: LEMS
57.
▲
by
jeffchuber
3y ago
cool :) (jeff from chroma)
58.
▲
by
jeffchuber
3y ago
(disclaimer: i cofounded Chroma) if you are building locally and dont want to send your data anywhere - try the open-source alternative Chroma https://github.com/chroma-core/chroma
59.
▲
by
jeffchuber
3y ago
chroma can help here https://github.com/chroma-core/chroma
60.
▲
by
jeffchuber
3y ago
chroma is getting much faster on Monday. be cautious with pgvector - the recall can be extremely bad (% retrieved nearest neighbors vs ground truth)
More ›