4 ms·
> relatively common I'd love to see some links, all I see used in practice (including in the OP blog post) is semantic search and a bit of clustering > adding
by ozb 3y ago
> relatively common
I'd love to see some links, all I see used in practice (including in the OP blog post) is semantic search and a bit of clustering
> adding a pair of glasses
Actually that (and generally all the SD/VAE stuff) is a great example of the kind of thing I was thinking of, though I have yet to see that concept being used together with a vector database; generally all the user-facing stuff I've seen fits it into the standard "train a model, then do inference" workflow, in contrast to something like semantic search which more obviously focuses on the embeddings themselves
> First and third being identical
Definitely related, but I make the distinction between projection/sorting along an axis vs constructing a new vector by addition/subtraction
> Manipulate within gaussian space, then return to target space
This is definitely along the lines of what I had in mind, any example of this being used in practice?
> Embedding is an overloaded word
Yeah I'm using the term somewhat loosely and broadly here, as basically "a vector in a real vector space where distance represents some notion of semantic similarity"
> People do things like SVMs
Who?
- godelski 3y agoHere's an example from a Normalizing Flow. Good at density, not great at sampling. https://openai.com/research/glow https://openai.com/research/glow Here's a video of moving around in the latent space of a diffusion model https://www.youtube.com/watch?v=vEnetcj_728 https://www.youtube.com/watch?v=vEnetcj_728 Here's a stylegan one https://www.youtube.com/watch?v=bRrS74RXsSM https://www.youtube.com/watch?v=bRrS74RXsSM Or a VAE on mnist https://www.youtube.com/shorts/pgmnCU_DxzM https://www.youtube.com/shorts/pgmnCU_DxzM I mean it is a bit hard to answer your other questions because like I was pointing out, embeddings and latent spaces are pretty vague terms. For the mathy side, normalizing flows are a great choice since you can parameterize whatever data you want into whatever distribution you want. You then work in that distribution you created, which is approximately isomorphic to the data. But other models do similar things, just more lossy but better at things like sample generation. That's the tradeoff, interpretability/density vs expressitivity/sample quality. But diffusion and NODEs/Score models are closing that gap. But you're going to need to look at applied papers to view more people using them in ways like operational vector spaces. For example, there's VITs TTS uses a NF to parameterize parts of the model or controller networks tend to use similar things. It's more about thinking how your network works and communicates. I think a lot of people are just not thinking to hack away and operate on networks as if they're mathematical models instead of a locked box.