2 ms·Yes, it could just make one call to a multimodal llm to describe the sceneby killingtime74 4mo agoYes, it could just make one call to a multimodal llm to describe the scene