3 ms·
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
- vov_or 3y agoGuys trained a multi-modal chatbot with visual and language instructions based on the open-source multi-modal model OpenFlamingo! Paper link: https://arxiv.org/abs/2305.04790 https://arxiv.org/abs/2305.04790