4 ms·
Shapeshift: Semantically map JSON objects using key-level vector embeddings
- WhatsName 2y agoMaybe I'm not the target audience, but here are simple questions to the author or potential users: What about anything more complex like date of birth to age or the other way round? Also since we will inevitably incur costs, why not let a llm write a transformation rule for us?
- lukasb 2y agoWhat is this for? The examples given could be handled deterministically. Is this for situations where you don't know JSON schemas in advance? What situations are those?
- saltwatercowboy 2y agoThe lazy part of my brain screams “use this instead of dealing properly with nested objects!” In a production setting I’d be worried about consistency from the base to result layers if based on LLM transpositioning.
- eezing 2y agoData import via customer self-service onboarding.
- tbrownaw 2y agoAs is, it's not good for much beyond looking cool. (Maybe implementing Postel's Law for a json API, but I think that's considered bad taste these days.) If instead of transforming a single object it would output a table of src_field->dst_field, it could potentially be a useful first pass in some ETL development.
- henry700 2y agoKeep the bug generators going, we will need the jobs
- simonw 2y agoThis is the code that does the work: https://github.com/rectanglehq/Shapeshift/blob/d954dab2a866c750cd6862d1eb44830b825bcb06/index.ts#L132-L160 https://github.com/rectanglehq/Shapeshift/blob/d954dab2a866c... There are a few ways this could be made a less expensive to run: 1. Cache those embeddings somewhere. You're only embedding simple strings like "name" and "address" - no need to do that work more than once in an entire lifetime of running the tool. 2. As suggested here https://news.ycombinator.com/item?id=40973028 https://news.ycombinator.com/item?id=40973028 change the design of the tool so instead of doing the work it returns a reusable data structure mapping input keys to output keys, so you only have to run it once and can then use that generated data structure to apply the transformations on large amounts of data in the future. 3. Since so many of the keys are going to have predictable names ("name", "address" etc) you could even pre-calculate embeddings for the 1,000 most common keys across all three embedding providers and ship those as part of the package. Also: in https://github.com/rectanglehq/Shapeshift/blob/d954dab2a866c750cd6862d1eb44830b825bcb06/index.ts#L57-L64 https://github.com/rectanglehq/Shapeshift/blob/d954dab2a866c... you're using Promise.map() to run multiple embeddings through the OpenAI API at once, which risks tripping their rate-limit. You should be able to pass the text as an array in a single call instead, something like this: const response = await this.openai!.embeddings.create({ model: this.embeddingModel, input: texts, encoding_format: "float", }); return response.data.map(item => item.embedding); https://platform.openai.com/docs/api-reference/embeddings/create https://platform.openai.com/docs/api-reference/embeddings/cr... says input can be a string OR an array - that's reflected in the TypeScript library here too: https://github.com/openai/openai-node/blob/5873a017f0f2040ef97040a8df19c5b4dc2a66fd/src/resources/embeddings.ts#L80-L90 https://github.com/openai/openai-node/blob/5873a017f0f2040ef...
- marvinkennis 2y agoThanks for the suggestions! Will implement these. Caching is a great idea.
- slantedview 2y agoIn general, you might cross reference with other object mapping libraries (including in other languages) to get ideas on how they approach this problem. Caching mappings is just one common strategy.
- lordofmoria 2y agoSince LLMs are bad at the null hypothesis (in this case, when a key does not exist in the source JSON), how does this prevent hallucinating transformations for missing keys?
- deleted 2y ago[deleted]
- mpeg 2y agoThis isn't using an LLM, it simply checks for similarity between keys using vector embeddings
- explaininjs 2y agoWhat’d be really great is a codegen aspect. A non-negligible part of any data munching operation is “this input object has fields X, Y, Z and we need an output object with fields X, f(X), Y, f(Y,Z)”. This is something and LLM has a decent chance at being really quite good at.
- visarga 2y agoThis task in the most general form is better done with question answering prompt than embeds. How do you solve "Full Name" -> "First Name", "Last Name" with embeds? QA is the right level of abstraction for schema conversion tasks. And it's simple, just put the source JSON + target JSON schema in the prompt and ask for value extraction.
- yetanotherjosh 2y agoSo this identifies keys from source and target objects that are fuzzy synonyms and copies the values over. What is a real world use case for this? Add the fact that it's fuzzy and won't always work, so would require a great deal of extra effort in QA/testing (harder than just mapping the keys programmatically), and I'm puzzled.
- anamexis 2y agoWe do something very similar with embeddings in our product. Users import files that they have to match to a dynamically-defined target schema. The embedding matching provides suggested matches to the user that are generally very accurate, so they don't have to go through and manually match up "telephone" to "phone number" etc. It even works across languages.
- momojo 2y agoHow much time dos this save your users? Is this QOL? Or more of a "our product wouldn't work without this feature" kind of thing?
- anamexis 2y agoQuite a bit of time. The product would still work without the feature, but it is a major feature. It bypasses lots of wading through dropdowns (potentially dozens for a single session)
- magicalhippo 2y agoI've got some similar use-cases. So, do I understand correctly that you take the source keyword and generate an embedding vector of it, then compare it using dot-product similarity or something to the embedded vectors of the target keywords?
- anamexis 2y agoExactly, although we use cosine similarity.
- hendler 2y agoCreated a Rust version using devin.ai. (untested) https://github.com/HumanAssisted/shapeshift-rust https://github.com/HumanAssisted/shapeshift-rust
- benzguo 2y agoPut together a quick version with an LLM, using Substrate: https://www.val.town/v/substrate/shapeshift https://www.val.town/v/substrate/shapeshift I've turned the target object into a JSON schema, but you could probably generate that JSON schema pretty reliably using a codegen LLM.
- happy_bacon 2y agoHere is an another DSL for implementing object model mappings: https://github.com/patleahy/lir https://github.com/patleahy/lir
- leobg 2y agoThe example could be handled with no machine learning at all. Just use a bag of words comparison with a subword tokenizer. And if you do need embeddings (to map synonyms/topics), fastText is faster, cheaper and runs locally. For hard cases, you can feed the source/target schemas to gpt-4o once to create a map - and then apply that one map to all instances.
- riku_iki 2y ago> fastText is faster, cheaper and runs locally the question is if quality will be acceptable
- flysand7 2y agoThe question if machine learning algorithm's produced embeddings will have the acceptable quality too. With a library I presume that the quality is at least predictable. I personally have less trust in machine learning though
- riku_iki 2y ago> The question if machine learning algorithm's produced embeddings will have the acceptable quality too there are tons of benchmarks and results which demonstrated that embeddings from language models are superior to word2vec in (almost) all scenarios.
- srean 2y agoBTW Bag of words models were once considered ML not too long ago.