4 ms·
Fully multimodal on the input side: The same model can take in text, images, audio tokens, reason about them, and generate text as outputs. Compact yet powerfu
by vykthur 2y ago
Fully multimodal on the input side: The same model can take in text, images, audio tokens, reason about them, and generate text as outputs.
Compact yet powerful: At just 5.8 billion parameters, this model can be optimized for deployment worldwide. As the official blog post mentions, it "delivers highly efficient low latency inference, all while optimizing for on-device execution and reduced computational overhead."