4 ms·
I don't quite understand how this works in an unsupervised setting. The only thing that comes to mind is embedding that preserves distance, such as MDS (https:
by maxs 5y ago
I don't quite understand how this works in an unsupervised setting.
The only thing that comes to mind is embedding that preserves distance, such as MDS (https://en.wikipedia.org/wiki/Multidimensional_scaling#Metric_multidimensional_scaling_(mMDS) https://en.wikipedia.org/wiki/Multidimensional_scaling#Metri...)
- adw 5y agoOne intuition is that you can generate pairs which you know to be the “same thing” (a single example under heavy augmentation) and ensure they’re close in representation space whereas mismatched pairs are maximized in distance. That’s a label-free approach which should give you a space with nice properties for eg nearest-neighbor approaches, and there’s, it follows, some reason to believe then that it’d be a generally useful feature space for downstream problems.
- randcraw 5y agoIf you're pairing samples that you have decided share a sameness, then implicitly, you're labeling. I would not call that unsupervised.
- m3at 5y agoYes this is more often called self-supervised. Note that most sample pairings, especially for images, is done through augmentations currently, so the implicit labeling you're doing is still weak on priors. Of the methods mentioned in the article, BYOL (and even more the follow-up SimSiam [1]), have the weakest assumptions and work surprisingly well despite their simplicity. [1] https://arxiv.org/abs/2011.10566 https://arxiv.org/abs/2011.10566
- zwaps 5y agoI agree with Op that this is still essentially learning on labeled data. I say this, since there are also cases of constrastive sampling like ideas with truly unsupervised data. For example, Graph Embedding, where a graph implies structural features of similarity and distance that the representations should capture.