3 ms·Grasp Any Region: Precise, Contextual Pixel Understanding for Multimodal LLMs1 points by badmonster 1y ago