4 ms·
Lots of great findings --- I'm curious if anyone knows whether it is better to pass structured data or unstructured data to embedding api's? If I ask ChatGPT,
by thomasfromcdnjs 2y ago
Lots of great findings
---
I'm curious if anyone knows whether it is better to pass structured data or unstructured data to embedding api's? If I ask ChatGPT, it says it is better to send unstructured data. (looking at the authors github, it looks like he generated embeddings from json strings)
My use case is for jsonresume, I am creating embeddings by sending full json versions as strings, but I've been experimenting with using models to translate resume.json's into full text versions first before creating embeddings. The results seem to be better but I haven't seen any concrete opinions on this.
My understanding is that unstructured data is better because it contains textual/semantic meaning because of natural lanaguage aka
skills: ['Javascript', 'Python']
is worse than;
Thomas excels at Javascript and Python
Another question: What if the search was also a json embedding? JSON <> JSON embeddings could also be great?
- minimaxir 2y agoIn general I like to send structured data (see the input format here: https://github.com/minimaxir/mtg-embeddings https://github.com/minimaxir/mtg-embeddings), but the ModernBERT base for the embedding model used here specifically has better benefits implicitly for structured data compared to previous models. That's worth another blog post explaining why.
- notpublic 2y agoplease do explain why
- minimaxir 2y agotl;dr the base ModernBERT was trained with code in mind unlike most encoder-only models (therefore assuming it was also trained on JSON/YAML objects) and also includes a custom tokenizer to support that, which is why I mention that indentation is important since different levels of indentation have different single tokens. This is mostly theoetical and does require a deeper dive to confirm.
- vunderba 2y agoI'd say the more important consideration is "consistency" between incoming query input and stored vectors. I have a huge vector database that gets updated/regenerated from a personal knowledge store (markdown library). Since the user is most likely to input a comparison query in the form of a question "Where does X factor into the Y system?" - I use a small 7b parameter LLM to pregenerate a list of a dozen possible theoretical questions a user might pose to a given embedding chunk. These are saved as 1536 dimension sized embeddings into the vector database (Qdrant) and linked to the chunks. The real question you need to ask is - what's the input query that you'll be comparing to the embeddings? If it's incoming as structured, then store structured, etc. I've also seen (anecdotally) similarity degradation for smaller chunks as well - so keep that in mind as well.