2 ms·Video Representation Learning with Joint-Embedding Predictive Architectures2 points by fofoz 2y ago