6 ms·
Great point! (Disclaimer: I work for Elastic) Elasticsearch has recently added a data type called semantic_text, which automatically chunks text, calculates e
by morgango 2y ago
Great point!
(Disclaimer: I work for Elastic)
Elasticsearch has recently added a data type called semantic_text, which automatically chunks text, calculates embeddings, and stores the chunks with sensible defaults.
Queries are similarly simplified, where vectors are calculated and compared internally, which makes a lot less I/O and a lot simpler client code.
https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text https://www.elastic.co/search-labs/blog/semantic-search-simp...
- jdthedisciple 2y agoHow does their embedding model compare in terms of retrieval accuracy to, say `text-embedding-3-small` and `text-embedding-3-large`?
- binarymax 2y agoIt’s impossible to answer that question without knowing what content/query domain you are embedding. Checkout MTEB leaderboard, dig into the retrieval benchmark, and look for analogous datasets.
- splike 2y agoYou can use openai embeddings in elastic if you don't want to use their elser sparse embeddings
- pjot 2y agoI made something similar, but used duckDB as the vector store (and query engine)! It’s impressively fast https://github.com/patricktrainer/duckdb-embedding-search https://github.com/patricktrainer/duckdb-embedding-search
- ekianjo 2y agoThere is vector type data available in duckdb now?
- wild_egg 2y agoThey call it a fixed size array type but, yes. It was added earlier this year. Works really great https://duckdb.org/2024/05/03/vector-similarity-search-vss.html https://duckdb.org/2024/05/03/vector-similarity-search-vss.h...
- pjot 2y agoYep! It was added in v0.10.0 - which was released a month or two after I made this. This is using v0.9.1
- barrenko 2y agoAmy specific reason to use dDB? I've got a crapload of json q & a formatted discussions on a topic, and am trying to figure out if I just store it somewhere and query it, or do I also do vector embeddings, kinda lost with all the possible options.
- pjot 2y agoEmbeddings are what encode the “meaning” of a given text. Similarity search works by computing the angle between your query vector and the rest of the vectors already stored. DuckDB (and columnar stores in general) is great at aggregation. It’s particularly well suited because DuckDB is a single file. There’s no server to muck with.
- jackbravo 2y agoI love duckdb, but their concurrency model is very limiting: DuckDB has two configurable options for concurrency: 1. One process can both read and write to the database. 2. Multiple processes can read from the database, but no processes can write (access_mode = 'READ_ONLY'). https://duckdb.org/docs/connect/concurrency.html https://duckdb.org/docs/connect/concurrency.html