4 ms·
I wonder what the speed of this approach vs traditional ocr techniques. Also, curious if this could be used for text detection (find a bounding box containing t
by gfiorav 2y ago
I wonder what the speed of this approach vs traditional ocr techniques. Also, curious if this could be used for text detection (find a bounding box containing text within an image).
- vunderba 2y agoWas just coming here to say this, there does not yet exist a multimodal vision LLM approach that is capable of identifying bounding boxes of where the text occurs. I suppose you could manually cut the image up and send each part separately to the LLM but that feels like an kludge and it's still in-exact.
- EarlyOom 2y agoWe can do bounding boxes too :) we just call it visual grounding https://github.com/vlm-run/vlmrun-cookbook/blob/main/notebooks/04_visual_grounding.ipynb https://github.com/vlm-run/vlmrun-cookbook/blob/main/noteboo...
- vunderba 2y agoWait what? That's pretty neat. I'm on my phone right now, so I can't really view the notebook very easily. How does this work? Are you using some kind of continual partitioning of the image and refeeding that back into the LLM to sort of pseudo-zoom in/out on the parts that contain non-cut off text until you can resolve that into rough coordinates?
- deleted 2y ago[deleted]
- what 2y agoKind of skeptical since you also provide a “confidence” value, which has to be entirely made up. Do you have an example that isn’t a sample drivers license? Something that is unlikely to have appeared in an LLM’s training data?
- chpatrick 2y agoqwen 2.5 vl was specifically trained to produce bounding boxes I believe.