5 ms·
I believe that your interpretation is not correct. Based on my brief reading of the paper, the model contains 1) some known architecture for embedding the modal
by dfgfhjkjlkhgjfh 5y ago
I believe that your interpretation is not correct. Based on my brief reading of the paper, the model contains 1) some known architecture for embedding the modality, and 2) the feature reconstruction transformer network, the two being trained at the same time.
If I am not mistaken, the masking occurs in the input modality, not the feature-space, even though it is the feature-space that is used for the reconstruction task.
Regarding how the feature space is kept uncollapsed, it seems like a hyperparameter-tweaking (ie unsolved?) problem; quoting the paper:
"
Representation collapse.
A common issue with algorithms which create and predict their own targets is representation collapse. This occurs when the model produces very similar representations for all masked segments making the problem trivial to solve. Different strategies have been proposed to address this issue, e.g., contrastive models such as wav2vec 2.0 (Baevski et al., 2020b) use the same target representation both as a positive and a negative example, preventing collapse. Algorithms such as BYOL (Grill et al., 2020) do not optimize the teacher parameters to minimize the loss. VicReg (Bardes et al., 2021) adds an explicit loss encouraging variance among different representations.
In our experiments we found that collapse is most likely to happen in the following scenarios:
First, the learning rate is too large or the learning rate warmup is too short which can often be solved by tuning the respective hyper-parameters.
Second, the EMA decay rate is too low which leads to student model collapse which is propagated to the teacher due to parameter tracking. This can be addressed by carefully tuning τ0, τe and τn.
Third, we found collapse to be more likely for modalities where adjacent targets are very correlated and where longer spans need to be masked, such as for speech. We address this by either explicitly penalizing the lack of variance (Bardes et al., 2021), or by promoting variance through normalizing target representations over the current sequence or batch (Grill et al., 2020). The former worked well for small models but is less reliable for larger models and it also requires tuning additional hyper-parameters. In contrast, we found applying instance or batch normalization before or after averaging targets to
work well while being simpler. For models where targets are less correlated such as for vision and NLP, momentum tracking is sufficient to prevent representation collapse.
"
- WithinReason 5y agoYou're right. That seems to explain why representations don't collapse into a constant, but not why they don't collapse to the same feature...
- algo_trader 5y agoAre there papers that show similar results on varied structured/relational/graphed data modalities? Even with large/labeled/cleaned dataset, it seems that each domain change or even formatting/encoding forces you change the the architecture.