3 ms·
Wow, this is pretty good. My sort of benchmark is a photo of someone holding my sweet little Ollie. Well, he's not so little and he's not mine anymore, but he'l
by devinprater 3y ago
Wow, this is pretty good. My sort of benchmark is a photo of someone holding my sweet little Ollie. Well, he's not so little and he's not mine anymore, but he'll always be my widdle Ollie!
Anyways, GPT4-vision wasn't always able to tell me that he doesn't really look the most comfortable being held, cause that's a lot of gravity pulling him down. Neither was Llava in the past. But with Llava 1.6 34B, it can, with no further questions asked besides "Please describe this image" as the first user message along with the image. So yeah this is really amazing! Its OCR has also definitely improved. Before, it'd just say text is in another language, but now it just shows the text. Can't wait to have a good enough computer to quickly run this locally.
- dog436zkj3p7 3y agoWhat?
- washadjeffmad 3y agoCheck their profile.
- deleted 3y ago[deleted]
- thelastparadise 3y agoI did check it...
- kromem 3y agoIt must be a pretty exciting time for you. I can't imagine how much of a difference it would make when vision models like this can be run at the edge in real time for blind users, including proper triage of relevant scene information narration. Wild to think about the long tail of accessibility and how that's going to be wildly different with the improvements to generative AI.