4 ms·
> We train our models on a combination of an internal dataset consisting of 14 million video-text pairs The paper is sorely lacking evaluation; one thing I'd l
by evouga 4y ago
> We train our models on a combination of an internal dataset consisting of 14 million video-text pairs
The paper is sorely lacking evaluation; one thing I'd like to see for instance (any time a generative model is trained on such a vast corpus of data) is a baseline comparison to nearest-neighbor retrieval from the training data set.
- Garlef 4y agoIn the end: Who cares? Sure, from an academic perspective this might be interesting. But in the end it's the users who will pick the tool. And they will figure out the things the tool can not express in the first week after release. So: Focusing on increasing expressiveness and ergonomics should beat academic rigour.