5 ms·
This isn't surprising at all - most VLMs today are quite poor on localization even though they've been explicitly post-trained on object detection tasks. One i
by fzysingularity 1y ago
This isn't surprising at all - most VLMs today are quite poor on localization even though they've been explicitly post-trained on object detection tasks.
One insight that the author calls out is the inconsistencies in coordinate systems used in post-training these - you can't just swap models and get similar results. Gemini uses (ymin, xmin, ymax, xmax) integers b/w 0-1000. Qwen uses (xmin, ymin, xmax, ymax) floats b/w 0-1. We've been evaluating most of the frontier models for bounding boxes / segmentation masks, and this is quite a footgun to new users.
One of the reasons we chose to delegate object-detection to specialized tools is essentially due to the poor performance (~0.34 mAP w/ Gemini to 0.6 mAP w/ DETR like architectures). Check out this cookbook [1] we recently released, we use any LLM to delegate tasks like object-detection, face-detection and other classical CV tasks to a specialized model while still giving the user the dev-ex of a VLM.
[1] https://colab.research.google.com/github/vlm-run/vlmrun-cookbook/blob/main/notebooks/10_mcp_showcase.ipynb https://colab.research.google.com/github/vlm-run/vlmrun-cook...
- joshvm 1y agoBox format degeneracy has been a footgun for computer vision developers since forever. You can define a rectangle as two corner coordinates, a coordinate + width + height. Since "one coordinate" can be a corner or the center, there are normally 6 variations and every single one exists somewhere. This also causes havoc for validation because it's easy to forget and wonder why your metrics are all practically zero because you didn't specify the right one. There's a neat table here: https://dragoneye.ai/blog/a-guide-to-bounding-box-formats https://dragoneye.ai/blog/a-guide-to-bounding-box-formats Picking yxyx was certainly a decision.