6 ms·
Hey all, I did a lot of the ML work for this. Let me know if you have any questions. The title might be a little ambitious since we only have two embeddings
by mlucy 8y ago
Hey all,
I did a lot of the ML work for this. Let me know if you have any questions.
The title might be a little ambitious since we only have two embeddings right now, but it really is our goal to have embeddings for everything. You can see some of our upcoming embeddings at https://www.basilica.ai/available-embeddings/ https://www.basilica.ai/available-embeddings/.
We basically want to do for these other datatypes what word2vec did for NLP. We want to turn getting good results with images, audio, etc. from a hard research problem into something you can do on your laptop with scikit.
- xapata 8y agoWhat's different between an "embedding" and a projection, which I believe is the more standard term for this kind of transformation?
- thanatropism 8y agoProjection is a type of embedding. But you can't really describe what LTSA, UMAP, etc. do as projection. LTSA "unrolls" data rather than projecting it.
- xapata 8y agoOn the contrary, I'd say an "unrolling" is a non-linear projection. Wikipedia suggests that embedding is a type of projection. I hadn't heard the term before today. https://en.wikipedia.org/wiki/Nonlinear_dimensionality_reduction#Modified_Locally-Linear_Embedding_(MLLE) https://en.wikipedia.org/wiki/Nonlinear_dimensionality_reduc...
- mlucy 8y ago"Embedding" is the term I've heard used for this most often. It's definitely the term that seems to dominate in the literature. (Just to pick a random paper off my reading list: https://arxiv.org/abs/1709.03856 https://arxiv.org/abs/1709.03856 .) In my mind "embedding" carries the connotation that you're moving into a smaller space that's easier to work with, and where things which are similar in some way are near each other.
- xapata 8y agoThe TensorFlow docs agree with you. This field has such jargon proliferation, it's hard to keep up. https://developers.google.com/machine-learning/crash-course/embeddings/video-lecture https://developers.google.com/machine-learning/crash-course/...
- mlucy 8y agoYeah, I agree. The number of new nouns per year is kind of ridiculous.
- sjg007 8y agoEmbedding is the ML term for a non-linear projection.
- xapata 8y agoAmusingly, it seems you and the other reply are contradicting each other, saying both that an embedding is a specific type of projection and that a projection is a specific type of embedding. When I was learning about SVMs way back when, we said "non-linear projection" instead of "embedding".
- sjg007 8y agoI mean both are correct given the literature at this point. Maybe the use of embedding captures the notion that you are mapping through a NN which is a variable function of sorts instead of a single or set of linear or orthonormal functions.
- FridgeSeal 8y agoThe mathematical term "Embedding" refers to finding a lower dimensional representation of some high(er) dimensional data, whether this is done with single function (maybe hand chosen) or a neural network doesn't matter.
- sjg007 8y agoI don't think that's quite right.. According to wikipedia, embedding is: https://en.wikipedia.org/wiki/Embedding https://en.wikipedia.org/wiki/Embedding So Projection sounds closer to what you describe. https://en.wikipedia.org/wiki/Projection_(mathematics) https://en.wikipedia.org/wiki/Projection_(mathematics) So my original comment is not quite right either. The embedding is a structure preserving map. A projection isn't necessarily structure preserving. ''' To be an embedding, such a mapping must preserve order "both ways": ''' http://mathworld.wolfram.com/Embedding.html http://mathworld.wolfram.com/Embedding.html I think that a projection from p dimensions to q<p dimensions isn't structure preserving. At this point a mathematician may be able to explain more.
- thanatropism 8y agoIsn't the point of word2vec that embeddings are semantically meaningful vectors?
- mlucy 8y agoDefinitely! In particular, semantically similar words are close to each other after embedding, so the space ends up with semantically meaningful clusters. Our embeddings have the same property. If you embed two similar images, they'll end up closer to each other than two dissimilar images. (Where "similar" depends on the training details, but that's true for word2vec as well.)
- deleted 8y ago[deleted]
- Fireflite 8y agoMost other word embeddings have hundreds of dimensions, not thousands. Are you able to hint at what causes this difference? Do you see better downstream task performance?
- mlucy 8y agoIt depends on the task. If you're doing clustering or instance retrieval, you probably want to PCA the number of dimensions down to 200 or so. (In fact, we do this in the tutorial at https://www.basilica.ai/tutorials/how-to-train-an-image-model/ https://www.basilica.ai/tutorials/how-to-train-an-image-mode... .) If you're training a big regression, you'll probably get better results with the larger embedding. We decided to err on the side of making the embeddings too big, because it's very easy to reduce the number of dimensions on the user's end, and impossible to increase it.
- farza 8y agoHi there Lucy! So, I don't know a ton about Word2Vec which probably doesn't help, but I do understand that it makes tasks, like NLP, much easier since you're going from this massive space (the english language) into a smaller embedding that you learn. That being said, how are you embedding images? Is it based on how similar they are, if so what does "similarity" mean? Also, what dataset was leveraged? Any more info on how you do this task would be awesome :).
- mlucy 8y agoHey! We're embedding images by feeding them through a deep neural net and using the activations of an intermediate layer as an embedding. You can read https://arxiv.org/abs/1403.6382 https://arxiv.org/abs/1403.6382 to learn more about this technique if you're interested. Our launch model is trained on ImageNet, which has enough variety that it usually generalizes well even when your dataset is very dissimilar to the input distribution. We're planning to train on a wider variety of image data in the future, but we wanted to get something into people's hands quickly.
- ru999gol 8y agoIn the case of images I can just take an off-the-shelf pre-trained model like ResNet as a feature extractor, why should I use a cloud service for that? I don't quite get what the benefit of it is. Are your embeddings better? Well prove it then? In terms of transfer learning, fine-tuning convolutional layers would perform way better anyways?
- PeterisP 8y agoDo you plan to make these (many in the future) embeddings to refer to a single 'semantic vector space' or have each of them be separate? I.e. do you plan to do the work to align the embeddings of different types of media so that the contents are somewhat similar and e.g. an audio recording of a snippet gets a similar embedding to the equivalent text and a somewhat similar embedding to a picture that's being described?
- mlucy 8y agoWe aren't currently doing this. In the future I think we'll try to embed into a single space on a best-effort basis, assuming we can find the engineering resources. It will be really hard for some data types, but for the big ones like text/image/audio it isn't that hard, and will probably be valuable to people.
- itronitron 8y agocan you give a brief history on the use of the word 'embedding' ?