2 ms·Transformer Multimodal Self-Supervised Learning Video, Audio and Text (2021)2 points by reqo 3y ago