3 ms·
Compared to Mirasol3B[1] this is not supporting audio as a modality. What Google has done with Mirasol3B made the demo of "Astro" in Google I/O possible. They d
by msoad 2y ago
Compared to Mirasol3B[1] this is not supporting audio as a modality. What Google has done with Mirasol3B made the demo of "Astro" in Google I/O possible. They do a little of cheating by converting audio to images(spectrogram) and video to 25 photo frames with some sort of attention system to things that change during those frames. So the tokenizer is basically the same for audio and video and images.
I believe Meta is going to this direction with multimodality as well. The new GPT voice mode is probably using the same architecture.
What's mind boggling is that models perform better at the same parameter size with new modality added to them!
It seems obvious that 3D is the next modality.
[1] https://arxiv.org/pdf/2311.05698 https://arxiv.org/pdf/2311.05698