16 ms·
Embeddings are underrated
- kaycebasques 2y agoCool, first time I've seen one of my posts trend without me submitting it myself. Hopefully it's clear from the domain name and intro that I'm suggesting technical writers are underrating how useful embeddings can be in our work. I know ML practitioners do not underrate them.
- dartos 2y agoYeah embeddings are the unsung killer feature of LLMs
- donavanm 2y agoYou might want to highlight chunking and how embeddings can/should represent subsections of your document as well. It seems relevant to me for cases like similarity or semantics search, getting the reader to the relevant portion of the document or page. Theres probably some interesting ideas around tokenization and metadata as well. For example, if you’re processing the raw file I expect you want to strip out a lot of markup before tokenization of the content. Conversely, some markup like code blocks or examples would be meaningful for tokenization and embedding anyways. I wonder if both of those ideas can be combined for something like automated footnotes and annotations. Linking or mouseover relevant content from elsewhere in the documentation.
- MrGreenTea 2y agoDo you have any resources you recommend for representing sub sections? I'm currently prototyping a note/thoughts editor where one feature is suggesting related documents/thoughts (think linked notes in Obsidian) for which I would like to suggest sub sections and not only full documents.
- donavanm 2y agoSorry, no good references off hand. I’ve had to help write & generate public docs in DocBook in the past. But no expert on either editors, nlp, or embeddings besides hacking around some tools for my own note taking. My assumption is youll want to use your existing markup structure, if you have it. Or naively split on paragraphs with a tool like spacy. Or get real fancy and use dynamic ranges; something like an accumulation window that aggregates adjacent sentences based on individual similarity, break on total size or dissimilarity, and then treat that aggregate as the range to “chunk.”
- MrGreenTea 2y agoThanks for the elaborate and helpful response. I'm also hacking on this as a personal note taking project and already started playing around with your ideas. Thanks!
- enjeyw 2y agoHaha yeah I was about to comment that I recall a period just after Word2Vec came out where embeddings were most definitely not underrated but rather the most hyped ML thing out there!
- rahimnathwani 2y agoI'm not sure why the voyage-3 models aren't on the MTEB leaderboard. The code for the leaderboard suggests they should be there: https://huggingface.co/spaces/mteb/leaderboard/commit/b7faae9e2db6d721cacc15cb923f29b8bb9115a4 https://huggingface.co/spaces/mteb/leaderboard/commit/b7faae... But I don't see them when I filter the list for 'voyage'.
- newrotik 2y agoIt is unclear this model should be on that leaderboard because we don't know whether it has been trained on mteb test data. It is worth noting that their own published material [0] does not entail any score from any dataset from the mteb benchmark. This may sound nit picky, but considering transformers' parroting capabilities, having seen test data during training should be expected to completely invalidate those scores. [0] see excel spreadsheet linked here https://blog.voyageai.com/2024/09/18/voyage-3/ https://blog.voyageai.com/2024/09/18/voyage-3/
- jdthedisciple 2y agoI'm critical of the low number of embedding dims. Could hurt performance in niche applications, in my estimation. Looking forward to try the announced large models though.
- fzliu 2y ago(I work at Voyage) Many of the top-performing models that you see on the MTEB retrieval for English and Chinese tend to overfit to the benchmark nowadays. voyage-3 and voyage-3-lite are also pretty small in size compared to a lot of the 7B models that take the top spots, and we don't want to hurt performance on other real-world tasks just to do well on MTEB.
- jdthedisciple 2y agoIt would still be great to know how it compares? Why should I pick voyage-3 if for all I know it sucks when it comes to retrieval accuracy (my personally most important metric)?
- quantadev 2y agoThat was a good post. Vector Embeddings are in some sense a summary of a doc that's unique similar to a hashcode of a doc. It makes me think it would be cool if there were some universal standard for generating embeddings, but I guess they'll be different for each AI model, so they can't have the same kind of "permanence" hash codes have. It definitely also seems like there should be lots of ways to utilize "Cosine Similarity" (or other closeness algos) in databases and other information processing apps that we haven't really exploited yet. For example you could almost build a new kind of Job Search Service that matches job descriptions to job candidates based on nothing but a vector similarity between resume and job description. That's probably so obvious it's being done, already.
- kqr 2y agoFor one point of inspiration, see https://entropicthoughts.com/determining-tag-quality https://entropicthoughts.com/determining-tag-quality I really like the picture you are drawing with "semantic hashes"!
- quantadev 2y agoYeah for "Semantic Hashes" (that's a good word for them!) we'd need some sort of "Canonical LLM" model that isn't necessarily used for inference, nor does it need to even be all that smart, but it just needs to be public for the world. It would need to be updated like every 2 to 5 years tho to account for new words or words changing meaning? ...but maybe could be updated in such a way as to not "invalidate" prior vectors, if that makes sense? For example "ride a bicycle" would still point in the same direction even after a refresh of the canonical model? It seems like feeding the same training set could replicate the same model values, but there are nonlinear instabilities which could make it disintegrate.
- kqr 2y agoMaybe the embedding could be paired up with a set of words that embed to somewhere close to the original embedding? Then the embedding can be updated for new models by re-embedding those words. (And it would be more interpretible by a human.)
- Aeolun 2y agoIs there some way to compare different embeddings for different use cases?
- jdthedisciple 2y agoSearch for MTEB Leaderboard on huggingface
- fzliu 2y agoGreat post! One quick minor note is that the resulting embeddings for the same text string could be different, depending on what you specify the input type as for retrieval tasks (i.e. query or document) -- check out the `input_type` parameter here: https://docs.voyageai.com/reference/embeddings-api https://docs.voyageai.com/reference/embeddings-api.
- thund 2y agoDoesn’t OpenAI embedding model support 8191/8192 tokens? That aside, declaring a winner by token size is misleading. There are more important factors like cross language support and precision for example
- jdthedisciple 2y agoYep, voyage-3 is not even anywhere in the top of the MTEB leaderboard if you order by `retrieval score` desc. stella_en_1.5B_v5 seems to be an unsung hero model in that regard plus you may not even want such large token sizes if you just need accurate retrieval of snippets of text (like 1-2 sentences)
- kaycebasques 2y agoThanks thund and jdthedisciple for these points and corrections. I'll update the section today.
- kaycebasques 2y agoUpdated the section to refer to the "Retrieval Average" column of the MTEB leaderboard. Is that the right column to refer to? Can someone link me to an explanation of how that benchmark works? Couldn't find a good link on it
- OutOfHere 2y agoAnd that's not all because token encodings of different models can be very different.
- nerdright 2y agoGreat post indeed! I totally agree that embeddings are underrated. I feel like the "information retrieval/discovery" world is stuck using spears (i.e., term/keyword-based discovery) instead of embracing the modern tools (i.e., semantic-based discovery). The other day I found myself trying to figure out some common themes across a bunch of comments I was looking at. I felt lazy to go through all of them so I turned my attention to the "Sentence Transformers" lib. I converted each comment into a vector embedding, applied k-means clustering on these embeddings, then gave each cluster to ChatGPT to summarize the corresponding comments. I have to admit, it was fun doing this and saved me lots of time!
- Gooblebrai 2y agoInteresting approach. Did you tell GPT to summarise the comments of each cluster after grouping them?
- matusp 2y agoMy hot take: embeddings are overrated. They are overfitted on word overlap, leading to both many false positives and false negatives. If you identify a specific problem with them ("I really want to match items like these, but it does not work"), it is almost impossible to fix them. I often see them being used inappropriately, by people who read about their magical properties, but didn't really care about evaluating their results.
- cheevly 2y agoYou can easily fix this using embedding arithmetic to build embedding classifiers.
- mbanerjeepalmer 2y agoAre there good examples of this working in the wild? Before I comb through all ten blue links... https://www.google.com/search?q=embedding%20arithmetic%20embedding%20classifier https://www.google.com/search?q=embedding%20arithmetic%20emb...
- nostrebored 2y ago"I really want to match items like these, but it does not work" is just a fine tuning problem.
- matusp 2y agoYes, in a sense that if you have infinite appropriate dataset and compute. No, in a sense what is practically achievable.
- nostrebored 2y agoYou don't need infinite data. You need ~100k samples. It's also not particularly expensive.
- matusp 2y ago
- mrob 2y agoEmbeddings are the only aspect of modern AI I'm excited about because they're the only one that gives more power to humans instead of taking it away. They're the "bicycle for our minds" of Steve Jobs fame; intelligence amplification not intelligence replacement. IMO, the biggest improvement in computer usability in my lifetime was the introduction of fast and ubiquitous local search. I use Firefox's "Find in Page" feature probably 10 or more times per day. I use find and grep probably every day. When I read man pages or logs, I navigate by search. Git would be vastly less useful without git grep. Embeddings have the potential to solve the biggest weakness of search by giving us fuzzy search that's actually useful.
- gwervc 2y agoI agee with this view. Generative AI robs us of something (thinking, practicing) which is the long term ability to practice a skill and improve oneself in exchange of an immediate (often crappy) result. Embeddings is a tech that can help us solve problem, ut we still have to do most of the work.
- wussboy 2y agoI’m not sure it robs us. It makes it possible, but many people including myself find the artistic products of AI to be utterly without value for the reasons you list. I will always cherish the product of lifelong dedication and human skill
- jacobr1 2y agoIt doesn't diminish - but I do find it interesting how it influences. Realism became less important, less interesting, though still valued to a lesser degree, with the ubiquity of photography. Where will human creativity move towards when certain task become trivially machine replicable? Where will human ingenuity _enabled_ by new technology make new art possible?
- larve 2y agoI ask LLMs to give me exercises, tutorials then write up my experience into "course notes", along with flashcards. I ask it to simulate a teacher, I ask it to simulate students that I have to teach, etc... I haven't found a tool that is more effective in helping me learn.
- imgabe 2y agoIs there any benefit to fine-tuning a model on your corpus before using it to generate embeddings? Would that improve the quality of the matches?
- gunalx 2y agoYes. Especially if you work in a not well supported language and/or have specific datapairs you want to match that might be out of ordinary text. Training your own fine tune takes a really short time and GPU resources, and you can easily outperform even sota models on your specific problem with a smaller model/vector space Then again on general English text and doing a basic fuzzy search. I would not really expect high performance gains.
- tomthe 2y agoNice introduction, but I think that ranking the models purely by their input token limits is not a useful exercise. Looking at the MTEB leaderboard is better (although a lot of the models are probably overfitting to their test set). This is a good time to chill for my visualization of 5 Millionembeddings of HN posts, users and comments: https://tomthe.github.io/hackmap/ https://tomthe.github.io/hackmap/
- kaycebasques 2y agoThanks, a couple other people gave me this same feedback in another comment thread and it definitely makes sense not to overindex on input token size. Will update that section in a bit.
- l5870uoo9y 2y agoAre there any visualization libraries that visualize embeddings in a vector space?
- f_devd 2y agoUMAP: https://umap-learn.readthedocs.io/en/latest/ https://umap-learn.readthedocs.io/en/latest/ scikit-learn also has options: https://scikit-learn.org/stable/auto_examples/manifold/plot_compare_methods.html#sphx-glr-auto-examples-manifold-plot-compare-methods-py https://scikit-learn.org/stable/auto_examples/manifold/plot_...
- sk11001 2y agoThere's attempts but you can only do so much in hundreds/thousands of dimensions. Most of the time the visualization doesn't really provide anything meaningful.
- beejiu 2y agoMy instinct would be a principal component analysis (which someone has demonstrated here: https://www.youtube.com/watch?app=desktop&v=brt88wwoZtI https://www.youtube.com/watch?app=desktop&v=brt88wwoZtI). Not sure it would tell you much though, but it looks nice.
- OutOfHere 2y agoIf you need them visualized, you're already on the wrong track.
- lmcinnes 2y agoAssuming you have a dimension-reduction or manifold learning tool of choice (UMAP,PacMAP,t-SNE,PyMDE,etc.) then DataMapPlot (https://datamapplot.readthedocs.io/en/latest/ https://datamapplot.readthedocs.io/en/latest/) is a library specifically designed to make visualizations of the outputs of your dimension reduction.
- adamgordonbell 2y agoI was using embeddings to group articles by topic, and hit a specific issue. Say I had 10 articles about 3 topics, and articles are either dry or very casual in tone. I found clustering by topic was hard, because tone dimensions ( whatever they were ) seemed to dominate. How can you pull apart the embeddings? Maybe use an LLM to extract a topic, and then cluster by extracted topic? In the end I found it easier to just ask an LLM to group articles by topic.
- eamag 2y agoI agree, I tried several methods during my pet project [1], and all of them have their pros and cons. Looks like creating topics first and predicting them using LLM works the best [1] https://eamag.me/2024/Automated-Paper-Classification https://eamag.me/2024/Automated-Paper-Classification
- coredog64 2y agoAllegedly, the new hotness in RAG is exactly that. Use a smaller LLM to summarize the article and include that summary alongside the article when generating the embedding. Potentially solves your issue, but it is also handy when you have to chunk a larger document and would lose context from calculating the embedding just on the chunk.
- joerick 2y agoThe thing that puzzles me about embeddings is that they're so untargeted, they represent everything about the input string. Is there a method for dimensionality reduction of embeddings for different applications? Let's say I'm building a system to find similar tech support conversations and I am only interested in the content of the discussion, not the tone of it. How could I derive an embedding that represents only content and not tone?
- adamgordonbell 2y agoAgreed.. biggest problem with off the shelf embeddings I hit. Need a way to decompose embeddings.
- johndough 2y agoYou can do math with word embeddings. A famous example (which I now see has also been mentioned in the article) is to compute the "woman vector" by subtracting "man" from "woman". You can then add the "woman vector" to e.g. the "king" vector to obtain a vector which is somewhat close to "queen". To adapt this to your problem of ignoring writing style in queries, you could collect a few text samples with different writing styles but same content to compute a "style direction". Then when you do a query for some specific content, subtract the projection of your query embedding onto the style direction to eliminate the style: query_without_style = query - dot(query, style_direction) * style_direction I suspect this also works with text embeddings, but you might have to train the embedding network in some special way to maximize the effectiveness of embedding arithmetic. Vector normalization might also be important, or maybe not. Probably depends on the training. Another approach would be to compute a "content direction" instead of a "style direction" and eliminate every aspect of a query that is not content. Depending on what kind of texts you are working with, data collection for one or the other direction might be easier or have more/fewer biases. And if you feel especially lazy when collecting data to compute embedding directions, you can generate texts with different styles using e.g. ChatGPT. This will probably not work as well as carefully handpicked texts, but you can make up for it with volume to some degree.
- joerick 2y ago
- NameError 2y agoThis article really resonates with me - I've heard people (and vector database companies) describe transformer embeddings + vector databases as primarily a solution for "memory/context for your chatbot, to mitigate hallucinations", which seems like a really specific (and kinda dubious, in my experience) use case for a really general tool. I've found all of the RAG applications I've tried to be pretty underwhelming, but semantic search itself (especially combined with full-text search) is very cool.
- moffkalast 2y agoI dare say RAG with vector DBs is underwhelming because embeddings are not underrated but appropriately rated, and will not give you relevant info in every case. In fact, the way LLMs retrieve info internally [0] already works along the same principle and is a large factor in their unreliability. [0] https://nonint.com/2023/10/18/is-the-reversal-curse-a-generalization-problem/ https://nonint.com/2023/10/18/is-the-reversal-curse-a-genera...
- dmezzetti 2y agoAuthor of txtai (https://github.com/neuml/txtai https://github.com/neuml/txtai) here. I've been in the embeddings space since 2020 before the world of LLMs/GenAI. In principle, I agree with much of the sentiment here. Embeddings can get you pretty far. If the goal is to find information and citations/links, you can accomplish most of that with a simple embeddings/vector search. GenAI does have an upside in that it can distill and process those results into something more refined. One of the main production use cases is retrieval augmented generation (RAG). The "R" is usually a vector search but doesn't have to be. As we see with things like ChatGPT search and Perplexity, there is a push towards using LLMs to summarize the results but also linking to the results to increase user confidence. Even Google Search now has that GenAI section at the top. In general, users just aren't going to accept LLM responses without source citations at this point. The question is if the summary provides value or if the citations really provide the most value. If it's the later, then Embeddings will get the job done.
- esafak 2y agoUnderrated by people are unfamiliar with machine learning, maybe.
- vindex10 2y agoI actually tend to agree. In the article, I didn't see the strong argument highlighting what powerful feature exactly people were missing in relation to embeddings. Those who work in ML they probably know these basics. It is a nice read though - explaining the basics of vector spaces, similarity and how it is used in modern ML applications.
- kaycebasques 2y ago> Hopefully it's clear from the domain name and intro that I'm suggesting technical writers are underrating how useful embeddings can be in our work. I know ML practitioners do not underrate them. https://news.ycombinator.com/item?id=42014036 https://news.ycombinator.com/item?id=42014036 > I didn't see the strong argument highlighting what powerful feature exactly people were missing in relation to embeddings I had to leave out specific applications as "an exercise for the reader" for various reasons. Long story short, embeddings provide a path to make progress on some of the fundamental problems of technical writing.
- vindex10 2y agothank you for explanation, yes I later encountered your answer and upvoted it. > I had to leave out specific applications as "an exercise for the reader" this is very unfortunate. would be very interesting to hear some intel :)
- lokar 2y agoEven by ML people from 25 years ago. It’s a black box function that maps from a ~30k space to a ~1k space. It’s a better function then things like PCA, but does the same thing.
- kkielhofner 2y agoLLMs have nearly completely sucked the oxygen out of the room when it comes to machine learning or "AI". I'm shocked at the number of startups, etc you see trying to do RAG, etc that basically have no idea what they are, how they actually work, etc. The "R" in RAG stands for retrieval - as in the entire field of information retrieval. But let's ignore that and skip right to the "G" (generative)... Garbage in, garbage out people!
- jonathanrmumm 2y agoEmbeddings are a new jump to universality, like the alphabet or numbers. https://thebeginningofinfinity.xyz/Jump%20to%20Universality https://thebeginningofinfinity.xyz/Jump%20to%20Universality
- OutOfHere 2y agoMind-blowing. In effect, among humans, what separates the civilized from the crude is the quest for universality among the civilized. To say it differently, thinking in terms of attaining universality is the mark of a civilized mind. I made an episode to appreciate the book: https://podcasters.spotify.com/pod/show/podgenai/episodes/The-Beginning-of-Infinity-Explanations-That-Transform-the-World-book-e2qeesi https://podcasters.spotify.com/pod/show/podgenai/episodes/Th...
- freediver 2y agoWhat would be really cool if somebody figured out how to do embeddings -> text.
- kabla 2y agoIs it not possible? I'm not that familiar with the topic. Doing some sort of averaging over a large corpus of separate texts could be interesting and probably would also have a lot of applications. Let's say that you are gathering feedback from a large group of people and want to summarize it in an anonymized way. I imagine you'd need embeddings with a somewhat large dimensionality though?
- cubefox 2y agoI wonder if someone has already tried to do that. Though this might go in a similar direction: https://arxiv.org/abs/1711.00043 https://arxiv.org/abs/1711.00043
- 0x1ceb00da 2y agoThat's chatgpt
- kaibee 2y agoHmm as a very stupid first pass... 0. Generate an embedding of some text, so that you have a known good embedding, this will be your target. 1. Generate an array of random tokens the length of the response you want. 2. Compute the embedding of this response. 3. Pick a random sub-section of the response and randomize the tokens in it again. 4. Compute the embedding of your new response. 5. If the embeddings are closer together, keep your random changes, otherwise discard them, go back to step 2. 6. Repeat this process until going back to step 2 stops improving your score. Also you'll probably want to shrink the size of the sub-section you're randomizing the closer your computed embedding is to your target embedding. Also you might be able to be cleverer by doing some kind of masking strategy? Like let's say the first half of your response text already was actually the true text of the target embedding. An ideal randomizer would see that randomizing the first half almost always makes the result worse, and so would target the 2nd half more often (I'm hoping that embeddings work like this?). 7. Do this N times and use an LLM to score and discard the worst N-1 results. I expect that 99.9% of the time you're basically producing adversarial examples w/ this strategy. 8. Feed this last result into an LLM and ask it to clean it up.
- ericholscher 2y agoThis is a great post. I’ve also been having a lot of fun working with embeddings, with lots of those pages being documentation. We write up a quick post on how are using them in prod, if you want to go from having an embedding to actually using them in a web app: https://www.ethicalads.io/blog/2024/04/using-embeddings-in-production-with-postgres-django-for-niche-ad-targeting/ https://www.ethicalads.io/blog/2024/04/using-embeddings-in-p...
- kaycebasques 2y agoThanks, Eric. So what you're really telling me is that you might make an exception to the "no tools talks" general policy for Write The Docs conference talks and let me nerd out on embeddings for 30 mins?? ;P
- ericholscher 2y agoHaha. I think they are definitely relevant, and I’d call them a technology more than a tool. That is mostly just that we don’t want folks going up and doing a 30 minute demo of Sphinx or something :-)
- huijzer 2y ago> Is it terrible for the environment? > I don’t know. After the model has been created (trained), I’m pretty sure that generating embeddings is much less computationally intensive than generating text. But it also seems to be the case that embedding models are trained in similar ways as text generation models2, with all the energy usage that implies. I’ll update this section when I find out more. Although I do care about the environment, this question is completely the wrong one if you ask me. There is the public opinion (mainstream media?) some kind of idea that we should use less AI and somehow this would solve our climate problems. As a counterexample, let's go to the extreme. Let's ban Google Maps because it does take computational resources from the phone. As a result more people will take wrong routes, and thus use more petrol. Say you use one gallon of petrol extra, that then wastes 34 kWh. This is of course the equivalent of running 34 powerful vacuum cleaners on full power for an hour. In contrast, say you downloaded your map, then the total "cost" is only the power used by the phone. A mobile phone has a battery of about 4 mAh, so 0,004 Ah * 4.2 V = 0.168 W, or 0.000168 kW. This means that the phone is about 200 000 times as efficient! And then we didn't even consider the time-saving for the human. It's the same with running embeddings for doc generation. An Nvidia H100 consumes about 700 W, so say 1 kWh after an hour of running. 1 kWh should be enough to do a bunch of embedding runs. If this then saves, for example, one workday including the driving back and forth to the office, then again the tradeoff is highly in favor of the compute.
- archerx 2y agoIf people really cared about the environment then they would ban residential air conditioning, it’s a luxury. I lived somewhere that would get to 40c in the summers and an oscillating fan was good enough to keep cool, the AC was nice to have but it wasn’t necessary. I find it very hypocritical when people tell you to change your lifestyle for climate change but they have the A/C blasting all day long, everyday.
- wussboy 2y ago“Reduce” was never going to work. Only the deep electrification of our economy will save us.
- ggnore7452 2y agoif anything i would consider embeddings bit overrated, or it is safer to underrate them. They're not the silver bullet many initially hoped for, they're not a complete replacement for simpler methods like BM25. They only have very limited "semantic understanding" (and as people throw increasingly large chunks into embedding models, the meanings can get even fuzzier) Overly high expectations lets people believe that embeddings will retrieve exactly what they mean, and With larger top-k values and LLMs that are exceptionally good at rationalizing responses, it can be difficult to notice mismatches unless you examine the results closely.
- nostrebored 2y agoOff the shelf embedding models definitely underpromise and overdeliver. In ten years I'd be very surprised if companies weren't fine-tuning embedding models for search based on their data in any competitive domains.
- kkielhofner 2y agoMy startup (Atomic Canyon) developed embedding models for the nuclear energy space[0]. Let's just say that if you think off-the-shelf embedding models are going to work well with this kind of highly specialized content you're going to have a rough time. [0] - https://huggingface.co/atomic-canyon/fermi-1024 https://huggingface.co/atomic-canyon/fermi-1024
- kkielhofner 2y ago> they're not a complete replacement for simpler methods like BM25 There are embedding approaches that balance "semantic understanding" with BM25-ish. They're still pretty obscure outside of the information retrieval space but sparse embeddings[0] are the "most" widely used. [0] - https://zilliz.com/learn/sparse-and-dense-embeddings https://zilliz.com/learn/sparse-and-dense-embeddings
- deepsquirrelnet 2y agoAbsolutely. Embeddings have been around a while and most people don’t realize it wasn’t until the e5 series of models from Microsoft that they even benchmarked as well as BM25 in retrieval scores, while being significantly more costly to compute. I think sparse retrieval with cross encoders doing reranking is still significantly better than embeddings. Embedding indexes are also difficult to scale since hnsw consumes too much memory above a few million vectors and ivfpq has issues with recall.
- mlinksva 2y agohttps://technicalwriting.dev/data/embeddings.html#let-a-thousand-embeddings-bloom https://technicalwriting.dev/data/embeddings.html#let-a-thou... > As docs site owners, I wonder if we should start freely providing embeddings for our content to anyone who wants them, via REST APIs or well-known URIs. Who knows what kinds of cool stuff our communities can build with this extra type of data about our docs? Interesting idea. You'd have to specify the exact embedding model used to generate an embedding, right? Is there a well understood convention for such identification like say model_name:model_version:model_hash or something? For technical docs, obviously very broad field, is there an embedding model (or small number) widely used or obviously highly suitable that a site ownwer could choose one and have some reasonable expectation that publishing embeddings for their docs generated using that model would be useful to others? (Naive questions, I am not embedded in the field.)
- skybrian 2y agoIt seems like sharing the text itself would be a better API, since it lets API users calculate their own embeddings easily. This is what the crawlers for search engines do. If they use embeddings internally, that’s up to them, and it doesn’t need to be baked into the protocol.
- treefarmer 2y agoYeah, this is the main issue with the suggestion. Embeddings can only be compared to each other if they are in the same space (e.g., generated by the same model). Providing embeddings of a specific kind would require users to use the same model, which can quickly become problematic if you're using a closed-source embedding model (like OpenAI's or Cohere's).
- nkko 2y agoCould we work toward standardization at some point? Obviously, there will always be a newer model. I just hate that all the embedding work I did was with now depreciated openai model. At least single providers should see interest in ensuring that for their own model releases. Some trick like matryoshka embedding could secure that embedding from newer models nest or work within the space of older model preserving some form of comparability or alignment
- OutOfHere 2y agoThis article shows the incorrect value for the OpenAI text-embedding-3-large Input Limit as 3072 which is actually its output limit [1]. The correct value is 8191 [2]. Edit: This value has now been fixed in the article. [1] https://platform.openai.com/docs/models/embeddings#embeddings https://platform.openai.com/docs/models/embeddings#embedding... [2] https://platform.openai.com/docs/guides/embeddings/#embedding-models https://platform.openai.com/docs/guides/embeddings/#embeddin... Also, what each model means by a token can be very different due to the use of different model-specific encodings, so ultimately one must compare the number of characters, not tokens.
- kaycebasques 2y agoA couple other issues with that section surfaced here: * https://news.ycombinator.com/item?id=42014683 https://news.ycombinator.com/item?id=42014683 * https://news.ycombinator.com/item?id=42015282 https://news.ycombinator.com/item?id=42015282 Updating that section now
- abound 2y ago"Reckless" seems a bit aggressive for what is likely an honest mistake in an otherwise very nice article.
- OutOfHere 2y agoEdited.
- tootie 2y agoIs it accurate to say that any data that can be tokenized can be turned into embeddings?
- eproxus 2y agoI wonder if this can be used to detect code similarity between e.g. function or files etc.? Or are the existing algorithms overly trained on written prose?
- OutOfHere 2y agoYes, of course it can be used in that way, but the quality of the result depends on whether the model was also trained on such code or not.
- hambandit 2y agoEmbeddings from things like one-hot, count vectorization, tf-idf, etc into dimensionality reduction techniques like SVD and PCA have been around for a long time and also provided the ability to compare any two pieces of text to each other. Yes, neural networks and LLMs have provided the ability for the context of each word to affect the whole document's embedding and capture more meaning, potentially that pesky "semantic" sort even; but they still are fundamentally a dimensionality reduction technique.
- ABraidotti 2y agoThis reminds me- I gotta go back and reread Borges's short stories with ML theory in mind.
- luizsantana 2y agoEmbeddings are indeed great. I have been using it a lot. Even wrote about it at: https://blog.dobror.com/2024/08/30/how-embeddings-make-your-email-client-better/ https://blog.dobror.com/2024/08/30/how-embeddings-make-your-...
- 0x20cowboy 2y agoI have made several successful products in the past few years using primarily embeddings and cosine similarity. Can recommend. It’s amazingly effective (compared to what most people are using today anyway).
- rgavuliak 2y agoThe title of the post says they are underrated, but doesn't provide any real justification beyond saying - they are good for x. I am not denying their usefulness, but it's misleading.
- _jonas 2y agoIt's fun to try and guess what semantic concepts might be captured within individual dimensions / pairs of dimensions of the embeddings space.