3 ms·
Not quite. I meant something that models pure binary sequences, not higher level tokens. That way, it could learn from any source that can be represented as bin
by optimalsolver 4y ago
Not quite. I meant something that models pure binary sequences, not higher level tokens. That way, it could learn from any source that can be represented as binary data. Could be video, text, audio, or all three at once.
It wouldn't be "video model", it would be an "anything that can be expressed in binary" model.
- visarga 4y agoMaybe you are interested in this paper: > Perceiver: General Perception with Iterative Attention Biological systems perceive the world by simultaneously processing high dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. Perceiver is a deep learning model that can process multiple modalities, such as images, point clouds, audio, and video, simultaneously. It is based on the transformer architecture and uses an asymmetric attention mechanism to distill a large number of inputs into a smaller latent bottleneck. This allows it to scale to handle very large inputs and outperform specialized models on classification tasks across various modalities. https://arxiv.org/abs/2103.03206 https://arxiv.org/abs/2103.03206
- optimalsolver 4y agoThanks! This looks really interesting.