5 ms·
> the size of a vector is over 2M Do you mean the dimension of the vector or the number of vectors?
by shri_krishna 3y ago
> the size of a vector is over 2M
Do you mean the dimension of the vector or the number of vectors?
- nl 3y agoThe dimension of the vector. It's the hidden state from a video vision transformer.
- ashvardanian 3y agoOuch! That’s fat! Which model is that? We have built a few video-search system by now, using USearch and UForm for embedding. They are only 256 dims and you can concatenate a few from different parts of the video. Any chance it would help? https://github.com/unum-cloud/uform https://github.com/unum-cloud/uform
- nl 3y agoIt's https://huggingface.co/docs/transformers/main/model_doc/vivit https://huggingface.co/docs/transformers/main/model_doc/vivi... I'm doing the most naive implementation possible at the moment though so it's likely I could improve it. > UForm Looks interesting. I'll have a play, thanks. I'm surprised there aren't more options in this space actually
- shri_krishna 3y agoDamn! 2M dimension dense vector is huge. Maybe you need to do some dimensionality reduction before attempting ANN. Something like PCA should help.
- henrydark 3y agoHow do you do pca on dimension 2M?
- woadwarrior01 3y agoHave you considered training a shallow MLP autoencoder, perhaps with tied weights between the encoder and the decoder to reduce the dimensionality of the embeddings? Another (IMO, better) approach I can think of off the top of my head, would be to use a semi-supervised contrastive learning approach, with labelled similar and dissimilar video pairs, like in this notebook[1] from OpenAI. [1]: https://github.com/openai/openai-cookbook/blob/main/examples/Customizing_embeddings.ipynb https://github.com/openai/openai-cookbook/blob/main/examples...
- nl 3y agoI originally started doing a triplet loss for video similarity similar to https://keras.io/examples/vision/siamese_network/ https://keras.io/examples/vision/siamese_network/ except for video instead of images. But I'm sort of hoping to avoid training a model.