3 ms·
Does anyone know how they linked image recognition with an LLM to give such specific instructions as shown in the bike video on the website?
by bkfh 3y ago
Does anyone know how they linked image recognition with an LLM to give such specific instructions as shown in the bike video on the website?
- HerculePoirot 3y agoI don't know but GPT4 was multimodal from the beginning. They just delayed the release of its image processing abilities. > We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks. > March 14, 2023 https://openai.com/research/gpt-4 https://openai.com/research/gpt-4