4 ms·
Using Multimodal LLMs to Understand UI Elements on Websites
- patricklef 2y agoMLLMs are surprisingly bad at this out of the box and to some extent even with fine tuning. https://jina.ai/news/the-what-and-why-of-text-image-modality-gap-in-clip-models/ https://jina.ai/news/the-what-and-why-of-text-image-modality...
- while1 2y agoLoving this! Very surprising that the LLMs of today are so bad at understanding interfaces but it also makes it a very interesting case for finetuning!
- antonoo 2y agoLove that you made an interactive app to visualize how the model performs similar to how Meta AI usually releases their models (direct link: https://qa-tech-minicpm-demo.gptengineer.run https://qa-tech-minicpm-demo.gptengineer.run)
- albinekb 2y ago[flagged]