2 ms·
Gemini multimodal live docs here: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/multimodal-live https://cloud.google.com/vertex-ai/gener
by dandiep 2y ago
Gemini multimodal live docs here: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/multimodal-live https://cloud.google.com/vertex-ai/generative-ai/docs/model-...
A little thin...
Also no pricing is live yet. OpenAI's audio inputs/outputs are too expensive to really put in production, so hopefully Gemini will be cheaper. (Not to mention, OAI's doesn't follow instructions very well.)
- kwindla 2y agoThe Multimodal Live API is free while the model/API is in preview. My guess is that they will be pretty aggressive with pricing when it's in GA, given the 1.5 Flash multimodal pricing. If you're interested in this stuff, here's a full chat app for the new Gemini 2 API's with text, audio, image, camera video and screen video. This shows how to use both the WebSocket API and to route through WebRTC infrastructure. https://github.com/pipecat-ai/gemini-multimodal-live-demo https://github.com/pipecat-ai/gemini-multimodal-live-demo
- dandiep 2y agoThanks, this is great!
- spencerchubb 2y agoI am eager to learn the pricing as well. It works sooo well but the pricing will make or break whether it's viable for apps