15 ms·
Embeddings are a good starting point for the AI curious app developer
- dullcrisp 2y agoIs there any easy way to run the embedding logic locally? Maybe even locally to the database? My understanding is that they’re hitting OpenAI’s API to get the embedding for each search query and then storing that in the database. I wouldn’t want my search function to be dependent on OpenAI if I could help it.
- dvt 2y agoYes, I use fastembed-rs[1] in a project I'm working on and it runs flawlessly. You can store the embeddings in any boring database (it's just an array of f32s at the end of the day). But for fast vector math (which you need for similarity search), a vector database is recommended, e.g. the pgvector[2] postgres extension. [1] https://github.com/Anush008/fastembed-rs https://github.com/Anush008/fastembed-rs [2] https://github.com/pgvector/pgvector https://github.com/pgvector/pgvector
- J_Shelby_J 2y agoFun timing! I literally just published my first crate: candle_embed[1] It uses Candle under the hood (the crate is more of a user friendly wrapper) and lets you use any model on HF like the new SoTA model from Snowflake[2]. [1] https://github.com/ShelbyJenkins/candle_embed https://github.com/ShelbyJenkins/candle_embed [2] https://huggingface.co/Snowflake/snowflake-arctic-embed-l https://huggingface.co/Snowflake/snowflake-arctic-embed-l
- deleted 2y ago[deleted]
- simonw 2y agoThere are a bunch of embedding models you can run on your own machine. My LLM tool had plugins for some of those: - https://llm.datasette.io/en/stable/plugins/directory.html#embedding-models https://llm.datasette.io/en/stable/plugins/directory.html#em... Here's how to use them: https://simonwillison.net/2023/Sep/4/llm-embeddings/ https://simonwillison.net/2023/Sep/4/llm-embeddings/
- notakash 2y agoIf you're building an iOS app, I've had success storing vectors in coredata and using a tiny coreml model that runs on device for embedding and then doing cosine similarity.
- ngalstyan4 2y agoWe provide this functionality in Lantern cloud via our Lantern Extras extension: <https://github.com/lanterndata/lantern_extras https://github.com/lanterndata/lantern_extras> You can generate CLIP embeddings locally on the DB server via: SELECT abstract, introduction, figure1, clip_text(abstract) AS abstract_ai, clip_text(introduction) AS introduction_ai, clip_image(figure1) AS figure1_ai INTO papers_augmented FROM papers; Then you can search for embeddings via: SELECT abstract, introduction FROM papers_augmented ORDER BY clip_text(query) <=> abstract_ai LIMIT 10; The approach significantly decreases search latency and results in cleaner code. As an added bonus, EXPLAIN ANALYZE can now tell percentage of time spent in embedding generation vs search. The linked library enables embedding generation for a dozen open source models and proprietary APIs (list here: <https://lantern.dev/docs/develop/generate https://lantern.dev/docs/develop/generate>, and adding new ones is really easy.
- charlieyuan 2y agoLantern seems really cool! Interestingly we did try CLIP (openclip) image embeddings but the results were poor for 24px by 24px icons. Any ideas? Charlie @ v0.app
- ngalstyan4 2y agoI have tried CLIP on my personal photo album collection and it worked really well there - I could write detailed scene descriptions of past road trips, and the photos I had in mind would pop up. Probably the model is better for everyday photos than for icons
- jmorgan 2y agoSupport for _some_ embedding models works in Ollama (and llama.cpp - Bert models specifically) ollama pull all-minilm curl http://localhost:11434/api/embeddings -d '{ "model": "all-minilm", "prompt": "Here is an article about llamas..." }' Embedding models run quite well even on CPU since they are smaller models. There are other implementations with a library form factor like transformers.js https://xenova.github.io/transformers.js/ https://xenova.github.io/transformers.js/ and sentence-transformers https://pypi.org/project/sentence-transformers/ https://pypi.org/project/sentence-transformers/
- laktek 2y agoIf you are building using Supabase stack (Postgres as DB with pgVector), we just released a built-in embedding generation API yesterday. This works both locally (in CPUs) and you can deploy it without any modifications. Check this video on building Semantic Search in Supabase: https://youtu.be/w4Rr_1whU-U https://youtu.be/w4Rr_1whU-U Also, the blog on announcement with links to text versions of the tutorials: https://supabase.com/blog/ai-inference-now-available-in-supabase-edge-functions https://supabase.com/blog/ai-inference-now-available-in-supa...
- jonplackett 2y agoSo handy! I already got some embeddings working with supabase pgvector and OpenAI and it worked great. What would the cost of running this be like compared to the OpenAI embedding api?
- laktek 2y agoThere are no extra costs other than the what we'd normally charge for Edge Function invocations (you get up to 500K in the free plan and 2M in the Pro plan)
- _bramses 2y agoneat! one thing i’d really love tooling for: supporting multi user apps where each has their own siloed data and embeddings. i find myself having to set up databases from scratch for all my clients, which results in a lot of repetitive work. i’d love to have the ability one day to easily add users to the same db and let them get to embedding without having to have any knowledge going in
- kiwicopple 2y agoThis is possible in supabase. You can store all the data in a table and restrict access with Row Level Security You also have various ways to separate the data for indexes/performance - use metadata filtering first (eg: filter by customer ID prior to running a semantic search). This is fast in postgres since its a relational DB - pgvector supports partial indexes - create one per customer based on a customer ID column - use table partitions - use Foreign Data Wrappers (more involved but scales horizontally)
- bryantwolf 2y agoThis is a good call out. OpenAI embeddings were simple to stand up, pretty good, cheap at this scale, and accessible to everyone. I think that makes them a good starting point for many people. That said, they're closed-source, and there are open-source embeddings you can run on your infrastructure to reduce external dependencies.
- jonnycoder 2y agoThe MTEB leaderboard has you covered. That is a goto for finding the leading embedding models and I believe many of them can run locally. https://huggingface.co/spaces/mteb/leaderboard https://huggingface.co/spaces/mteb/leaderboard
- deleted 2y ago[deleted]
- internet101010 2y agoOpen WebUI has langchain built-in and integrates perfectly with ollama. They have several variations of docker compose files on their github. https://github.com/open-webui/open-webui https://github.com/open-webui/open-webui
- thisiszilff 2y agoOne straightforward way to get started is to understand embedding without any AI/deep learning magic. Just pick a vocabulary of words (say, some 50k words), pick a unique index between 0 and 49,999 for each of the words, and then produce embedding by adding +1 to the given index for a given word each time it occurs in a text. Then normalize the embedding so it adds up to one. Presto -- embeddings! And you can use cosine similarity with them and all that good stuff and the results aren't totally terrible. The rest of "embeddings" builds on top of this basic strategy (smaller vectors, filtering out words/tokens that occur frequently enough that they don't signify similarity, handling synonyms or words that are related to one another, etc. etc.). But stripping out the deep learning bits really does make it easier to understand.
- pstorm 2y agoI'm trying to understand this approach. Maybe I am expecting too much out of this basic approach, but how does this create a similarity between words with indices close to each other? Wouldn't it just be a popularity contest - the more common words have higher indices and vice versa? For instance, "king" and "prince" wouldn't necessarily have similar indices, but they are semantically very similar.
- zachrose 2y agoMaybe the idea is to order your vocabulary into some kind of “semantic rainbow”? Like a one-dimensional embedding?
- svieira 2y agoYou are expecting too much out of this basic approach. The "simple" similarity search in word2vec (used in https://semantle.com/ https://semantle.com/ if you haven't seen it) is based on _multiple_ embeddings like this one (it's a simple neural network not a simple embedding).
- sdwr 2y agoIt doesn't even work as described for popularity - one word starts at 49,999 and one starts at 0.
- m1117 2y agoah pgvector is kind of annoying to start with, you have to set it up and maintain, and then it starts falling apart when you have more vectors
- sdesol 2y agoCan you elaborate more on the falling apart? I can see pgvector being intimidating for users with no experience standing up a DB, but I don't see how Postgres or pgvector would fall apart. Note, my reason for asking is I'm planning on going all in with Postgres, so pgvector makes sense for me.
- cargobuild 2y agohttps://www.pinecone.io/blog/pinecone-vs-pgvector/ https://www.pinecone.io/blog/pinecone-vs-pgvector/ check it out :)
- hackernoteng 2y agoWhat is "more vectors"? How many are we talking about? We've been using pgvector in production for more than 1 year without any issues. We dont have a ton of vectors, less than 100,000, and we filter queries by other fields so our total per cosine function is probably more like max of 5000. Performance is fine and no issues.
- ntry01 2y agoon the other hand, if you have postgres already, it may be easier to add pgvector than to add another dependency to your stack (especially if you are using something like supabase) another benefit is that you can easily filter your embeddings by other field, so everything is kept in one place and could help with perfomance it's a good place to start in those cases and if it is successful and you need extreme performance you can always move to other specialized tools like qdrant, pinecone or weaviate which were purpose-built for vectors
- chuckhend 2y agoTake a look at https://github.com/tembo-io/pg_vectorize https://github.com/tembo-io/pg_vectorize. It makes it a lot easier to get started. It runs on pgvector, but as a user, its completely abstracted from you. It also provides you with a way to auto-update embeddings as you add new data or update existing source data.
- patrick-fitz 2y agoNice project! I find it can be hard to think of a idea that is well suited to use AI. Using embeddings for search is definitely a good option to start with.
- ParanoidShroom 2y agoI made a reverse image search when I learned about embeddings. It's pretty fun to work with images https://medium.com/@christophe.smet1/finding-dirty-xtc-with-applied-machine-learning-61d1f9a2e857 https://medium.com/@christophe.smet1/finding-dirty-xtc-with-...
- LunaSea 2y agoDoes anyone have examples of word (ngram) disambiguation when doing Approximate Nearest Neighbour (ANN) on word vector embeddings?
- Imnimo 2y agoOne of the challenges here is handling homonyms. If I search in the app for "king", most of the top ten results are "ruler" icons - showing a measuring stick. Rodent returns mostly computer mice, etc. https://www.v0.app/search?q=king https://www.v0.app/search?q=king https://www.v0.app/search?q=rodent https://www.v0.app/search?q=rodent This isn't a criticism of the app - I'd rather get a few funny mismatches in exchange for being able to find related icons. But it's an interesting puzzle to think about.
- charlieyuan 2y agoGood call out! We think of this as a two part problem. 1. The intent of the user. Is it a description of the look of the icon or the utility of the icon? 2. How best to rank the results which is a combination of intent, CTR of past search queries, bootstrapping popularity via usage on open source projects etc. - Charlie of v0.app
- itronitron 2y ago>> If I search in the app for "king", most of the top ten results are "ruler" icons I believe that's the measure of a man.
- bryantwolf 2y agoYeah, these can be cute, but they're not ideal. I think the user feedback mechanism could help naturally align this over time, but it would also be gameable. It's all interesting stuff
- jonnycoder 2y agoAs the op, you can do both semantic search (embedding) and keyword search. Some RAG techniques call out using both for better results. Nice product by the way!
- bryantwolf 2y agoHybrid searches are great, though I'm not sure they would help here. Neither 'crown' nor 'ruler' would come back from a text search for 'king,' right? I bet if we put a better description into the embedding for 'ruler,' we'd avoid this. Something like "a straight strip or cylinder of plastic, wood, metal, or other rigid material, typically marked at regular intervals, to draw straight lines or measure distances." (stolen from a Google search). We might be able to ask a language model to look at the icon and give a good description we can put into the embedding.
- EcommerceFlow 2y agoEmbeddings have a special place in my heart since I learned about them 2 years ago. Working in SEO, it felt like everything finally "clicked" and I understood, on a lower level, how Google search actually works, how they're able to show specific content snippets directly on the search results page, etc. I never found any "SEO Guru" discussing this at all back then (maybe even now?), even though this was complete gold. It explains "topical authority" and gave you clues on how Google itself understands it.
- minimaxir 2y agoOne of my biggest annoyances with the modern AI tooling hype is that you need to use a vector store for just working with embeddings. You don't. The reason vector stores are important for production use-cases are mostly latency-related for larger sets of data (100k+ records), but if you're working on a toy project just learning how to use embeddings, you can compute cosine distance with a couple lines of numpy by doing a dot product of a normalized query vectors with a matrix of normalized records. Best of all, it gives you a reason to use Python's @ operator, which with numpy matrices does a dot product.
- twelfthnight 2y agoEven in production my guess is most teams would be better off just rolling their own embedding model (huggingface) + caching (redis/rocksdb) + FAISS (nearest neighbor) and be good to go. I suppose there is some expertise needed, but working with a vector database vendor has major drawbacks too.
- hackernoteng 2y agoUsing Postgres with pgvector is trivial and cheap. Its also available on AWS RDS.
- jonplackett 2y agoAlso on supabase!
- danielbln 2y agoOr you just shove it into Postgres + pg_vector and just use the DBMS you already use anyway.
- christiangenco 2y agoYup. I was just playing around with this in Javascript yesterday and with ChatGPT's help it was surprisingly simple to go from text => embedding (via. `openai.embeddings.create`) and then to compare the embedding similarity with the cosine distance (which ChatGPT wrote for me): https://gist.github.com/christiangenco/3e23925885e3127f2c1775871b8f52f1 https://gist.github.com/christiangenco/3e23925885e3127f2c177... Seems like the next standard feature in every app is going to be natural language search powered by embeddings.
- cargobuild 2y agoseeing comments about using pgvector... at pinecone, we spent some time understanding it's limitations and pain points. pinecone eliminates these pain points entirely and makes things simple at any scale. check it out: https://www.pinecone.io/blog/pinecone-vs-pgvector/ https://www.pinecone.io/blog/pinecone-vs-pgvector/
- gregorymichael 2y agoHas Pinecone gotten any cheaper? Last time I tried it was $75/month for the starter plan / single vector store.
- cargobuild 2y agoyep. pinecone serverless has reduced costs significantly for many workloads.
- dvaun 2y agoI’d love to build a suite of local tooling to play around with different embedding approaches. I’ve had great results using SentenceTransformers for quick one-off tasks at work for unique data asks. I’m curious about clustering within the embeddings and seeing what different approaches can yield and what applications they work best for.
- PaulHoule 2y agoIf I have 50,000 historical articles and 5,000 new articles I apply SBERT and then k-means with N=20 I get great results in terms of articles about Ukraine, sports, chemistry, and nerdcore from Lobsters ending up in distinct clusters. I’ve used DBSCAN for finding duplicate content, this is less successful. With the parameters I am using it is rare for there to be a false positives, but there aren’t that many true positives. I’m sure I could do do better if I tuned it up but I’m not sure if there is an operating point I’d really like.
- kaycebasques 2y agoI have been saying similar things to my fellow technical writers ever since the ChatGPT explosion. We now have a tool that makes semantic search on arbitrary, diverse input much easier. Improved semantic search could make a lot of common technical writing workflows much more efficient. E.g. speeding up the mandatory research that you must do before it's even possible to write an effective doc.
- gchadwick 2y agoFor an article extolling the benefits of embeddings for developers looking to dip their toe into the waters of AI it's odd they don't actually have an intro to embeddings or to vector databases. They just assume the reader already knows these concepts and dives on in to how they use them. Sure many do know these concepts already but they're probably not the people wondering about a 'good starting point for the AI curious app developer'.
- charlieyuan 2y agoApologies! Here's a good primer on embeddings from openai: https://platform.openai.com/docs/guides/embeddings https://platform.openai.com/docs/guides/embeddings
- simonw 2y agoI published this pretty comprehensive intro to embeddings last year: https://simonwillison.net/2023/Oct/23/embeddings/ https://simonwillison.net/2023/Oct/23/embeddings/
- nicbou 2y agoI found many of your other posts and they were the spark that finally made me "get it" and look deeper into LLMs. This post looks like another slam dunk. Keep up the good work!
- gk1 2y agoTo add to the other recommendations, here's a primer on vector DB's: https://www.pinecone.io/learn/vector-database/ https://www.pinecone.io/learn/vector-database/
- Samuel_w 2y ago[dead]
- hot_gril 2y agoThis is where I got started too. Glove embedding stored in Postgres. Pgvector is nice, and it's cool seeing quick tutorials using it. Back then, we only had cube, which didn't do cosine similarity indexing out of the box (you had to normalize vectors and use euclidean indexes) and only supported up to 100 dimensions. And there were maybe other inconveniences I don't remember, cause front page AI tutorials weren't using it.
- isoprophlex 2y agoPGvector is very nice indeed. And you get to store your vectors close to the rest of your data. I'm yet to understand the unique use case for dedicated vector dbs. It seems so annoying, having to query your vectors in a separate database without being able to easily join/filter based on the rest of your tables. I stored ~6 million hacker news posts, their metadata, and the vector embeddings in a cheap 20$/month vm running pgvector. Querying is very fast. Maybe there's some penalty to pay when you get to the billion+ row counts, but I'm happy so far.
- hot_gril 2y agoYou can also store vectors or matrices in a split-up fashion as separate rows in a table, which is particularly useful if they're sparse. I've handled huge sparse matrix expressions (add, subtract, multiply, transpose) that way, cause numpy couldn't deal with them.
- brianjking 2y agoAs I'm trying to work on some pricing info for PGVector - can you share some more info about the hacker news posts you've embedded? * Which embedding model? (or number of dimensions) * When you say 6 million posts - it's just the URL of the post, title, and author, or do you mean you've also embedded the linked URL (be it HN or elsewhere)? Cheers!
- thorum 2y agoCan embeddings be used to capture stylistic features of text, rather than semantic? Like writing style?
- levocardia 2y agoProbably, but you might need something more sophisticated than cosine distance. For example, you might take a dataset of business letters, diary entries, and fiction stories and train some classifier on top of the embeddings of each of the three types of text, then run (embeddings --> your classifier) on new text. But at that point you might just want to ask an LLM directly with a prompt like - "Classify the style of the following text as business, personal, or fiction: $YOUR TEXT$"
- vladimirzaytsev 2y agoYou may get way more accurate results from relatively small models as well as logits for each class if you ask one question per class instead.
- vladimirzaytsev 2y agoLikely not, embeddings are very crude. Embeddings of a text is just an average of "meanings" of words. As is embeddings lack a lot of tricks that made transformers so efficient.
- crowcroft 2y agoMy smooth brain might not understand this properly, but the idea is we generate embeddings, store them, then use retrieval each time we want to use them. For simple things we might not need to worry about storing much, we can generate the embeddings and just cache them or send them straight to retrieval as an array or something... The storing of embeddings seems the hard part, do I need a special database or PG extension? Is there any reason I can't store them as a blobs in SQlite if I don't have THAT much data, and I don't care too much about speed? Do embeddings generated ever 'expire'?
- H1Supreme 2y agoVector databases are used to store embeddings.
- crowcroft 2y agoBut why is that? I’m sure it’s the ‘best’ way to do things, but it also means more infrastructure which for simple apps isn’t worth the hassle. I should use redis for queues but often I’ll just use a table in a SQLite database. For small scale projects I find it works fine, I’m wondering what an equivalent simple option for embeddings would be.
- e0 2y agoPerhaps sqlite-vss? It adds vector searches to sqlite. https://github.com/asg017/sqlite-vss https://github.com/asg017/sqlite-vss
- chuckhend 2y agocheck out https://github.com/tembo-io/pg_vectorize https://github.com/tembo-io/pg_vectorize - we're taking it a little bit beyond just the storage and index. The project uses pgvector for the indices and distance operators, but also adds a simpler API, hooks into pre-trained embedding models, and helps you keep embeddings updated as data changes/grows
- deleted 2y ago[deleted]
- aidenn0 2y agoCan someone give a qualitative explanation of what the vector of a word with 2 unrelated meanings would look like compared to the vector of a synonym of each of those meanings?
- base698 2y agoIf you think about it like a point on a graph, and the vectors as just 2D points (x,y), then the synonyms would be close and the unrelated meanings would be further away.
- aidenn0 2y agoI'm guessing 2 dimensions isn't for this. Here's a concrete example: "bow" would need to be close to "ribbon" (as in a bow on a present) and also close to "gun" (as a weapon that shoots a projectile), but "ribbon" and "gun" would seem to need be far from each other. How does something like word2vec resolve this? Any transitive relationship would seem to fall afoul of this.
- clementmas 2y agoEmbeddings are indeed a good starting point. Next step is choosing the model and the database. The comments here have been taken over by database companies so I'm skeptical about the opinions. I wish MySQL had a cosine search feature built in
- bootsmann 2y agopg_vector has you covered
- mrkeen 2y agoGiven not because they’re sufficiently advanced technology indistinguishable from magic, but the opposite. Unlike LLMs, working with embeddings feels like regular deterministic code. <h3>Creating embeddings</h3> I was hoping for a bit more than: They’re a bit of a black box Next, we chose an embedding model. OpenAI’s embedding models will probably work just fine.
- akoboldfrying 2y agoI agree. The article was useful insofar as it detailed the steps they took to solve their problem clearly, and it's easy to see that many common problems are similar and could therefore be solved similarly, but I went in expecting more insight. How are the strings turned into arrays of numbers? Why does turning them into numbers that way lead to these nice properties?
- Aachen 2y agoSame here. I was saving the article for when I have a few hours to really dive into it, build upon it, learn from seeing and doing. Imagine my disappointment when I had the evening cleared, started reading, and discover all they're showing is how to concatenate a string, download someone else's black box model which outputs the similarity between the user's query and the concatenated info about each object, and then write queries on the output It's good to know you can do this performantly on your own system, but if the article had started out with "look, this model can output similarity between two texts and we can make a search engine with that", that'd be much more up front about what to expect to learn from it Edit: another comment mentioned you can't even run it yourself, you need to ask ClosedAI for every search query a user does on your website. WTF is this article, at that point you might as well pipe the query into general-purpose chatgpt which everyone already knows and let that sort it out
- mehulashah 2y agoI think he is saying: embeddings are deterministic, so they are more predictable in production. They’re still magic, with little explain ability or adaptability when they don’t work.
- benreesman 2y agoWithout getting into any big debates about whether or not RAG is medium-term interesting or whatever, you can ‘pip install sentence-transformers faiss’ and just immediately start having fun. I recommend using straightforward cosine similarity to just crush the NYT’s recommender as a fun project for two reasons: there’s an API and plenty of corpus, and it’s like, whoa, that’s better than the New York Times. He’s trying to sell a SaaS product (Pinecone), but he’s doing it the right way: it’s ok to be an influencer if you know what you’re taking about. James Briggs has great stuff on this: https://youtube.com/@jamesbriggs https://youtube.com/@jamesbriggs
- aeth0s 2y ago> crush the NYT’s recommender as a fun project for two reasons Could you share what recommender you're referring to here, and how you can evaluate "crushing" it? Sounds fun!
- adamgordonbell 2y agoMy problem with this is that it doesn't explain a lot. You can manually make a vector of a word and then step wise get up to word2vec approach and then document embedding. My post[1] does some of the first part and this great word2vec post[2] dives into it in more detail. [1] https://earthly.dev/blog/cosine_similarity_text_embeddings/ https://earthly.dev/blog/cosine_similarity_text_embeddings/ [2] https://jalammar.github.io/illustrated-word2vec/ https://jalammar.github.io/illustrated-word2vec/
- thomasfromcdnjs 2y agoI've been adding embeddings to every project I work on for the purpose of vector similarity searches. I was just trying to order uber eats and wondering why they don't have a better search based off embeddings. Almost finished building a feature on JSON Resume, that takes your hosted resume and WhoIsHiring job posts and uses embeddings to return relevant results -> https://registry.jsonresume.org/thomasdavis/jobs https://registry.jsonresume.org/thomasdavis/jobs
- voxelc4L 2y agoIt begs the question though, doesn't it...? Embeddings require a neural network or some reasonable facsimile to produce the embedding in the first place. Compression to a vector (a semantic space of some sort) still needs to happen – and that's the crux of the understanding/meaning. To just say "embeddings are cool let's use them" is ignoring the core problem of semantics/meaning/information-in-context etc. Knowing where an embedding came from is pretty damn important. Embeddings live a very biased existence. They are the product of a network (or some algorithm) that was trained (or built) with specific data (and/or code) and assume particular biases intrinsically (network structure/algorithm) or extrinsically (e.g., data used to train a network) which they impose on the translation of data into some n-dimensional space. Any engineered solution always lives with such limitations, but with the advent of more and more sophisticated methods for the generation of them, I feel like it's becoming more about the result than the process. This strikes me as problematic on a global scale... might be fine for local problems but could be not-so-great in an ever changing world.
- suprgeek 2y agoGreat project and excellent initiative to learn about embeddings. Two possible avenues to explore more. Your system backend could be thought of as being composed of two parts: |Icons->Embedder->|PGVector|->Retriever->Display Result| 1. In the embedder part trying out different embedding models and/or vector dimensions to explore if the Recall@K & Precision@K for your data set (icons) improves. Models make a surprising amount of difference to the quality of the results. Try the MTEB Leaderboard for ideas on which models to explore. 2. In the Information Retriever part you can try a couple of approaches: a.after you retrieve from PGVector see if you can use a reranker like Cohere to get better results https://cohere.com/blog/rerank https://cohere.com/blog/rerank b.You could try a "fusion ranking" similar to the one you do but structured such that 50% of the weight is for a plain old keyword search in the metadata and 50% is for the embedding based search Finally something more interesting to noodle on - what if the embeddings were based on the icon images and the model knew how to search for a textual descriptions in the latent space?
- Nuella19 2y ago[flagged]
- primitivesuave 2y agoI learned how to use embeddings by building semantic search for the Bhagavad Gita. I simply saved the embeddings for all 700 verses into a big file which is stored in a Lambda function, and compared against incoming queries with a single query to OpenAI's embedding endpoint. Shameless plug in case anyone wants to test it out - https://gita.pub https://gita.pub
- forgingahead 2y agoReally nice and beautiful site!
- primitivesuave 2y agoThank you! :)
- pantulis 2y agoI strongly agree with the title of the article. RAG is very interesting right now just as an example of how technology moves from being just fresh out of academia to being engineered and commoditized into regular out of the shelf tools. On the other hand I don't think it's that important to understand how embeddings are calculated, for the beginner it's more important to showcase why they work and why they enable simple reasoning like "queen = woman + (king - men)" and the possible use cases.
- deleted 2y ago[deleted]
- KasianFranks 2y agoThey are named 'feature' vectors with scored attributes, similar to associative arrays.Just ask MI. Jordan, D. Blie, S. Mian or A. Ng.
- jerrygenser 2y agoThey are embedded into a particular semantic vector space that is learned based on a model. Another feature vector could be hand rolled based on feature engineering, tidf ngrams etc. Embedding is typically distinct from feature engineering that is manual.
- deleted 2y ago[deleted]
- willcodeforfoo 2y agoOne thing I'm not sure of is how much of a larger bit of text should go into an embedding? I assume it's a trade off of context and recall, with one word not meaning much semantically, and the whole document being too much to represent with just numbers. Is there a sweet spot (e.g. split by sentence) or am I missing something here?
- mistermann 2y ago> You can even try dog breeds like ‘hound,’ ‘poodle,’ or my favorite ‘samoyed.’ It pretty much just works. But that’s not all; it also works for other languages. Try ‘chien’ and even ‘犬’1! Can anyone explain how this language translation works? The magic is in the embeddings of course, but how does it work, how does it translate ~all words across all languages?
- tapatio 2y agoTangential question: how are people using GenAI for financial datasets for insights and recommendations? Assume tens of desparate databases with financial data. Does NL2SQL work well for this? Or OpenAI Tools (formerly OpenAI Functions)? What have you found that is consistently accurate?