4 ms·
Show HN: NExT-GPT – First LLM working with multimodal input and output
NExT-GPT is the first end-to-end MM-LLM that perceives input and generates output in arbitrary combinations (any-to-any) of text, image, video, and audio and beyond.
- jdwg 3y agoI gave it a photo of my face, and it said that I look approachable and friendly. I asked it to improve my hair, and it replaced me with a completely different person. Oh well! Still very impressive tech.
- Anil1331 3y agoInitial stages of Multi Modal AI but should get much better results soon. Believe it is going to come soon in GPT-5 and Google's Gemini